Updating controlled vocabularies by analysing query logs
Full item page Simple item page
- dc.contributor.author Vàllez Letrado, Marica
- dc.contributor.author Pedraza, Rafaelca
- dc.contributor.author Codina, Lluísca
- dc.contributor.author Blanco, Saúlca
- dc.contributor.author Rovira, Cristòfolca
- dc.date.accessioned 2017-04-25T07:52:34Z
- dc.date.available 2017-04-25T07:52:34Z
- dc.date.issued 2015
- dc.description.abstract Purpose: Controlled vocabularies play an important role in information retrieval. Numerous studies have shown that conceptual searches based on vocabularies are more effective than keyword searches, at least in certain contexts. Consequently, new ways must be found to improve controlled vocabularies. The purpose of this paper is to present a semi-automatic model for updating controlled vocabularies through the use of a text corpus and the analysis of query logs. Design/methodology/approach: An experimental development is presented in which, first, the suitability of a controlled vocabulary to a text corpus is examined. The keywords entered by users to access the text corpus are then compared with the descriptors used to index it. Finally, both the query logs and text corpus are processed to obtain a set of candidate terms to update the controlled vocabulary. Findings: This paper describes a model applicable both in the context of the text corpus of an online academic journal and to repositories and intranets. The model is able to: first, identify the queries that led users from a search engine to a relevant document; and second, process these queries to identify candidate terms for inclusion in a controlled vocabulary. Research limitations/implications: Ideally, the model should be used in controlled web environments, such as repositories, intranets or academic journals. Social implications: The proposed model directly improves the indexing process by facilitating the maintenance and updating of controlled vocabularies. It so doing, it helps to optimise access to information. Originality/value: The proposed model takes into account the perspective of users by mining queries in order to propose candidate terms for inclusion in a controlled vocabulary.
- dc.description.sponsorship This article is part of the projects ‘Audiencias activas y periodismo’ (Active audiences and journalism; CSO2012-39518-C04-02) and ‘El turista en la Web: hábitos informacionales y selección de destino turístico’ (Tourists online: information habits and choice of tourism destination; CSO2014-59896-P) funded through grants under the Spanish National Research, Development and Innovation Plan of the Spanish Ministry of Economics and Competitiveness.
- dc.format.mimetype application/pdfca
- dc.identifier.citation Vàllez M, Pedraza-Jiménez R, Blanco S, Codina L, Rovira C. Updating controlled vocabularies by analysing query logs. Online Information Review. 2015;39(7):870-84. DOI: 10.1108/OIR-06-2015-0180
- dc.identifier.doi http://dx.doi.org/10.1108/OIR-06-2015-0180
- dc.identifier.issn 1468-4527
- dc.identifier.uri http://hdl.handle.net/10230/30893
- dc.language.iso eng
- dc.publisher Emeraldca
- dc.relation.ispartof Online Information Review. 2015;39(7):870-84.
- dc.relation.projectID info:eu-repo/grantAgreement/ES/3PN/CSO2012-39518-C04-02
- dc.relation.projectID info:eu-repo/grantAgreement/ES/1PE/CSO2014-59896-P
- dc.rights © Emerald Group Publishing Limited
- dc.rights.accessRights info:eu-repo/semantics/openAccess
- dc.subject.keyword Information retrieval
- dc.subject.keyword Controlled vocabularies
- dc.subject.keyword Indexing languages
- dc.subject.keyword Keywords
- dc.subject.keyword Query logs
- dc.title Updating controlled vocabularies by analysing query logsca
- dc.type info:eu-repo/semantics/article
- dc.type.version info:eu-repo/semantics/acceptedVersion