Document - Mining@home: public resource computing for distributed data mining

2008

Contribution to book Restricted

Mining@home: public resource computing for distributed data mining

Lucchese C, Orlando S, Talia D, Matroianni C, Barbalace D

Distributed Data mining

Several kinds of scientific and commercial applications require the execution of a large number of independent tasks. One highly successful and low cost mechanism for acquiring the necessary compute power for these applications is the "public-resource computing", or "desktop Grid" paradigm, which exploits the computational power of private computers. So far, this paradigmhas not been applied to data mining applications for two main reasons. First, it is not trivial to decompose a data mining algorithm into truly independent sub-tasks. Second, the large volume of data involved makes it difficult to handle the communication costs of a parallel paradigm. In this paper, we focus on one of the main data mining problem: the extraction of closed frequent itemsets from transactional databases. We show that is possible to decompose this problem into independent tasks, which however need to share a large volume of data. We thus introduce a data-intensive computing network, which adopts a P2P topology based on super peers with caching capabilities, aiming to support the dissemination of large amounts of information. Finally, we evaluate the execution of our data mining job on such network.

Metrics

Back to previous page

Cite as

BibTeX entry

@inbook{oai:it.cnr:prodotti:139029,
	title = {Mining@home: public resource computing for distributed data mining},
	author = {Lucchese C and Orlando S and Talia D and Matroianni C and Barbalace D},
	doi = {10.1007/978-0-387-09455-7_16},
	year = {2008}
}

CNR authors and affiliations

CNR authors

Lucchese, Claudio
0000-0002-2545-0425
Orlando, Salvatore
0000-0002-4155-9797

Laboratories

High Performance Computing (2002-ongoing)

Download

CNR IRIS

Bibliographic record
Deposited version

DOI

10.1007/978-0-387-09455-7_16

Also available from

Mining@home: public resource computing for distributed data mining

Metrics

Share

Cite as

CNR authors and affiliations

Download