Document - An Analysis of the Relative Hardness of Reuters-21578 Subsets

2003

Other Open Access

An Analysis of the Relative Hardness of Reuters-21578 Subsets

Debole F, Sebastiani F

Text categorization Standard benchmarks Experimental evaluation

The existence, public availability, and widespread acceptance of a standard benchmark for a given information retrieval (IR) task are beneficial to research on this task, since they allow different researchers to experimentally compare their own systems by comparing the results they have obtained on this benchmark. The Reuters-21578 test collection, together with its earlier variants, has been such a standard benchmark for the text categorization (TC) task throughout the last ten years. However, the benefits that this has brought about have somehow been limited by the fact that different researchers have 'carved' different subsets out of this collection, and tested their systems on one of these subsets only; systems that have been tested on different Reuters-21578 subsets are thus not readily comparable. In this paper we present a systematic, comparative experimental study of the three subsets of Reuters-21578 that have been most popular among TC researchers. The results we obtain allow us to determine the relative hardness of these subsets, thus establishing an indirect means for comparing TC systems that have, or will be, tested on these different subsets.

Back to previous page

Cite as

BibTeX entry

@misc{oai:it.cnr:prodotti:160110,
	title = {An Analysis of the Relative Hardness of Reuters-21578 Subsets},
	author = {Debole F and Sebastiani F},
	year = {2003}
}

CNR authors and affiliations

CNR authors

Debole, Franca
0000-0002-0369-6045
Sebastiani, Fabrizio
0000-0003-4221-6427

Laboratories

Networked Multimedia Information System (2002-2020)

Download

CNR IRIS

Bibliographic record
Deposited version

ISTI Repository

Deposited version

An Analysis of the Relative Hardness of Reuters-21578 Subsets

Share

Cite as

CNR authors and affiliations

Download