Distribution of content words and phrases in text and language modelling

SLAVA M. KATZ

doi:10.1017/S1351324996001246

Abstract

This paper addresses the problem of distribution of words and phrases in text, a problem of great general interest and of importance for many practical applications. The existing models for word distribution present observed sequences of words in text documents as an outcome of some stochastic processes; the corresponding distributions of numbers of word occurrences in the documents are modelled as mixtures of Poisson distributions whose parameter values are fitted to the data. We pursue a linguistically motivated approach to statistical language modelling and use observable text characteristics as model parameters. Multi-word technical terms, intrinsically content entities, are chosen for experimentation. Their occurrence and the occurrence dynamics are investigated using a 100-million word data collection consisting of a variety of about 13,000 technical documents. The derivation of models describing word distribution in text is based on a linguistic interpretation of the process of text formation, with the probabilities of word occurrence being functions of observable and linguistically meaningful text characteristics. The adequacy of the proposed models for the description of actually observed distributions of words and phrases in text is confirmed experimentally. The paper has two focuses: one is modelling of the distributions of content words and phrases among different documents; and another is word occurrence dynamics within documents and estimation of corresponding probabilities. Accordingly, among the application areas for the new modelling paradigm are information retrieval and speech recognition.

Crossref Citations

This article has been cited by the following publications. This list is generated based on data provided by Crossref.

Lewis, David D. 1998. Machine Learning: ECML-98. Vol. 1398, Issue. , p. 4.

Franz, Martin and McCarley, J. Scott 2000. Word document density and relevance scoring (poster session). p. 345.

Kazi, Z. and Ravin, Y. 2000. Who's who? Identifying concepts and entities across multiple documents. Vol. vol.1, Issue. , p. 7.

McDonald, Scott Turcato, Davide McFetridge, Paul Popowich, Fred and Toole, Janine 2000. Advances in Artificial Intelligence. Vol. 1822, Issue. , p. 126.

Nomoto, Tadashi and Matsumoto, Yuji 2001. A new approach to unsupervised text summarization. p. 26.

Nomoto, Tadashi and Matsumoto, Yuji 2003. The diversity-based approach to open-domain text summarization. Information Processing & Management, Vol. 39, Issue. 3, p. 363.

Steiner, Petra 2003. Sprache zwischen Theorie und Technologie / Language between Theory and Technology. p. 299.

Teevan, Jaime and Karger, David R. 2003. Empirical development of an exponential probabilistic model for text retrieval. p. 18.

Ming Lin Nunamaker, J.F. Chau, M. and Hsinchun Chen 2004. Segmentation of lecture videos based on text: a method combining multiple linguistic features. p. 9 pp..

Schneider, Karl-Michael 2004. Advances in Natural Language Processing. Vol. 3230, Issue. , p. 474.

Kobayashi, Mei and Aono, Masaki 2004. Survey of Text Mining. p. 103.

Federico, Marcello and Bertoldi, Nicola 2004. Broadcast news LM adaptation over time. Computer Speech & Language, Vol. 18, Issue. 4, p. 417.

Lean Yu Kin Keung Lai and Yue Wu 2005. A Framework of Web-Based Text Mining on the Grid. p. 97.

Saravanan, M. Raman, S. and Ravindran, B. 2005. A Probabilistic Approach to Multi-document Summarization for Generating a Tiled Summary. p. 167.

Schneider, Jesper W. and Borlund, Pia 2005. Context: Nature, Impact, and Role. Vol. 3507, Issue. , p. 226.

Gao, Jianfeng Li, Mu Huang, Chang-Ning and Wu, Andi 2005. Chinese Word Segmentation and Named Entity Recognition: A Pragmatic Approach. Computational Linguistics, Vol. 31, Issue. 4, p. 531.

Ding, Chris H.Q. 2005. A probabilistic model for Latent Semantic Indexing. Journal of the American Society for Information Science and Technology, Vol. 56, Issue. 6, p. 597.

Damle, Dileep and Uren, Victoria 2005. Extracting significant words from corpora for ontology extraction. p. 187.

Schneider, Karl-Michael 2005. Computational Linguistics and Intelligent Text Processing. Vol. 3406, Issue. , p. 682.

Chirita, Paul - Alexandru Firan, Claudiu S. and Nejdl, Wolfgang 2006. Pushing task relevant web links down to the desktop. p. 59.

Download full list

Article contents

Distribution of content words and phrases in text and language modelling

Abstract

Access options

Article purchase

Temporarily unavailable

This article has been cited by the following publications. This list is generated based on data provided by Crossref.

Article contents

Distribution of content words and phrases in text and language modelling

Abstract

Access options

Article purchase

Temporarily unavailable

Save article to Kindle

Save article to Dropbox

Save article to Google Drive

Reply to: Submit a response

Your details

You have entered the maximum number of contributors

Conflicting interests