Unsupervised keyword extraction using the GoW model and centrality scores

Batziou, Elissavet; Gialampoukidis, Ilias; Vrochidis, Stefanos; Kompatsiaris, Ioannis

Unsupervised keyword extraction using the GoW model and centrality scores

Mostra el registre complet Registre parcial de l'ítem

dc.contributor.author Batziou, Elissavetca
dc.contributor.author Gialampoukidis, Iliasca
dc.contributor.author Vrochidis, Stefanosca
dc.contributor.author Kompatsiaris, Ioannisca
dc.date.accessioned 2018-03-02T17:44:12Z
dc.date.available 2018-03-02T17:44:12Z
dc.date.issued 2017
dc.description Comunicació presentada a: The 4th International Conference, INSCI 2017, celebrat a Thessaloniki, Grècia, del 22 al 24 de novembre de 2017.ca
dc.description.abstract Nowadays, a large amount of text documents are produced on a daily basis, so we need e cient and e ective access to their con- tent. News articles, blogs and technical reports are often lengthy, so the reader needs a quick overview of the underlying content. To that end we present graph-based models for keyword extraction, in order to compare the Bag of Words model with the Graph of Words model in the key- word extraction problem. We compare their performance in two publicly available datasets using the evaluation measures Precision@10, mean Av- erage Precision and Jaccard coe cient. The methods we have selected for comparison are grouped into two main categories. On the one hand, centrality measures on the formulated Graph-of-Words (GoW) are able to rank all words in a document from the most central to the less central, according to their score in the GoW representation. On the other hand, community detection algorithms on the GoW provide the largest commu- nity that contains the key nodes (words) in the GoW. We selected these methods as the most prominent methods to identify central nodes in a GoW model. We conclude that term-frequency scores (BoW model) are useful only in the case of less structured text, while in more structured text documents, the order of words plays a key role and graph-based models are superior to the term-frequency scores per document.en
dc.description.sponsorship This work was supported by the projects H2020-645012 (KRISTINA) and H2020-700024 (TENSOR), funded by the European Commission.en
dc.format.mimetype application/pdf
dc.identifier.citation Batziou E, Gialampoukidis I, Vrochidis S, Kompatsiaris I. Unsupervised keyword extraction using the GoW model and centrality scores. In: Kompatsiaris I, Cave J, Satsiou A, Carle G, Passani A, Kontopoulos E, Diplaris S, McMillan D. Internet Science. 4th International Conference, INSCI 2017 Proceedings. 2017 Nov 22-24; Thessaloniki, Greece. [Cham]: Springer, 2017. (LNCS; no. 10673). DOI: 10.1007/978-3-319-70284-1
dc.identifier.doi http://dx.doi.org/10.1007/978-3-319-70284-1
dc.identifier.issn 0302-9743
dc.identifier.uri http://hdl.handle.net/10230/34045
dc.language.iso eng
dc.publisher Springerca
dc.relation.ispartof Kompatsiaris I, Cave J, Satsiou A, Carle G, Passani A, Kontopoulos E, Diplaris S, McMillan D. Internet Science. 4th International Conference, INSCI 2017 Proceedings. 2017 Nov 22-24; Thessaloniki, Greece. [Cham]: Springer, 2017. (LNCS; no. 10673).
dc.relation.projectID info:eu-repo/grantAgreement/EC/H2020/645012
dc.relation.projectID info:eu-repo/grantAgreement/EC/H2020/700024
dc.rights.accessRights info:eu-repo/semantics/openAccess
dc.subject.keyword Keyword-based searchen
dc.subject.keyword Topic-based filteringen
dc.subject.keyword Graph-based modelsen
dc.subject.keyword Graph of Wordsen
dc.subject.keyword Centrality measuresen
dc.subject.keyword Community detectionen
dc.title Unsupervised keyword extraction using the GoW model and centrality scoresca
dc.type info:eu-repo/semantics/conferenceObject
dc.type.version info:eu-repo/semantics/acceptedVersion

Col·leccions

Congressos (Departament de Tecnologies de la Informació i les Comunicacions)
Documents OpenAIRE (Open Access Infrastructure for Research in Europe)