Large linguistically-processed web corpora for multiple languages

Citació

  • Baroni M, Kilgarriff A. Large linguistically-processed web corpora for multiple languages. In: EACL '06: Proceedings of the Eleventh Conference of the European Chapter of the Association for Computational Linguistics: Posters & Demonstrations; 2006 Apr 5-6; Trento, Italy. Stroudsburg (PA): Association for Computational Linguistics; 2006. p. 87-90.

Enllaç permanent

Descripció

  • Resum

    The Web contains vast amounts of linguistic data. One key issue for linguists and language technologists is how to access it. Commercial search engines give highly compromised access. An alternative is to crawl the Web ourselves, which also allows us to remove duplicates and near-duplicates, navigational material, and a range of other kinds of non-linguistic matter. We can also tokenize, lemmatise and part-of-speech tag the corpus, and load the data into a corpus query tool which supports sophisticated linguistic queries. We have now done this for German and Italian, with corpus sizes of over 1 billion words in each case. We provide Web access to the corpora in our query tool, the Sketch Engine.
  • Descripció

    Comunicació presentada a: EACL '06: Eleventh Conference of the European Chapter of the Association for Computational Linguistics: Posters & Demonstrations celebrat del 5 al 6 d'abril de 2006 a Trento, Itàlia.
  • Mostra el registre complet