2004
Conference article  Unknown

An experimental comparison of term representations for term management applications

Lavelli A., Sebastiani F., Zanoli R.

Extensional Term Representations  Term co-occurrence 

A number of content management tasks, including term clustering, term categorization, and automated thesaurus generation, see natural language terms (e.g. words, noun phrases) as first-class objects, i.e. as objects endowed with an internal representation which makes them suitable for being explicitly manipulated by the corresponding algorithms. The information retrieval (IR) literature has traditionally used an extensional representation for terms according to which a term is represented by the 'bag of documents' in which the term occurs. The computational linguistics (CL) literature has independently developed an alternative extensional representation for terms, according to which a term is represented by the 'bag of terms' that co-occur with it in some document. This paper aims at discovering which of the two representations is most effective, i.e. brings about higher effectiveness once used in tasks that require terms to be explicitly represented and manipulated. In order to discover this we carry out experiments on a term categorization task, which allows us to compare the two different representations in closely controlled experimental conditions. We report the results of a large scale experimentation carried out by classifying under 42 different classes the terms extracted from a corpus of more than 60,000 documents. Our results show a substantial difference in effectiveness between the two representation styles; we give both an intuitive explanation and an information-theoretic justification for these different behaviours.

Source: SEBD 2004. 12° Convegno Nazionale su Sistemi Evoluti per Basi di Dati, pp. 190–201, S.Margherita di Pula, Cagliari, 21-23 June 2004



Back to previous page
BibTeX entry
@inproceedings{oai:it.cnr:prodotti:91040,
	title = {An experimental comparison of term representations for term management applications},
	author = {Lavelli A. and Sebastiani F. and Zanoli R.},
	booktitle = {SEBD 2004. 12° Convegno Nazionale su Sistemi Evoluti per Basi di Dati, pp. 190–201, S.Margherita di Pula, Cagliari, 21-23 June 2004},
	year = {2004}
}