Bridging information-theoretic and geometric compression in language models

Mostra el registre complet Registre parcial de l'ítem

  • dc.contributor.author Cheng, Emily
  • dc.contributor.author Kervadec, Corentin
  • dc.contributor.author Baroni, Marco
  • dc.date.accessioned 2023-12-18T07:03:06Z
  • dc.date.available 2023-12-18T07:03:06Z
  • dc.date.issued 2023
  • dc.description.abstract For a language model (LM) to faithfully model human language, it must compress vast, potentially infinite information into relatively few dimensions. We propose analyzing compression in (pre-trained) LMs from two points of view: geometric and information-theoretic. We demonstrate that the two views are highly correlated, such that the intrinsic geometric dimension of linguistic data predicts their coding length under the LM. We then show that, in turn, high compression of a linguistic dataset predicts rapid adaptation to that dataset, confirming that being able to compress linguistic information is an important part of successful LM performance. As a practical byproduct of our analysis, we evaluate a battery of intrinsic dimension estimators for the first time on linguistic data, showing that only some encapsulate the relationship between informationtheoretic compression, geometric compression, and ease-of-adaptation.
  • dc.format.mimetype application/pdf
  • dc.identifier.citation Cheng E, Kervadec C, Baroni M. Bridging information-theoretic and geometric compression in language models. In: Bouamor H, Pino J, Bali K (Editors). Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; 2023 Dec 6-10; Singapore. East Stroudsburg PA: ACL; 2023. p. 12397-420.
  • dc.identifier.uri http://hdl.handle.net/10230/58559
  • dc.language.iso eng
  • dc.publisher ACL (Association for Computational Linguistics)
  • dc.relation.ispartof Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; 2023 Dec 6-10; Singapore. East Stroudsburg PA: ACL; 2023. p. 12397-420.
  • dc.rights © ACL, Creative Commons Attribution 4.0 License
  • dc.rights.accessRights info:eu-repo/semantics/openAccess
  • dc.rights.uri https://creativecommons.org/licenses/by/4.0/
  • dc.subject.other Models lingüístics
  • dc.title Bridging information-theoretic and geometric compression in language models
  • dc.type info:eu-repo/semantics/conferenceObject
  • dc.type.version info:eu-repo/semantics/publishedVersion