Skip to main navigation Skip to search Skip to main content

Structural classification of XML documents using multisets

  • University of Massachusetts Boston

Research output: Contribution to journalArticlepeer-review

Abstract

In this paper, we investigate the problem of clustering XML documents based on their structure. We represent the paths in an XML document as a multiset and use the symmetric difference operation on multisets to define certain metrics. These metrics are then used to obtain a measure of similarity between any two documents in a collection. Our technique was successfully applied to real and synthesized XML documents yielding high-quality clusterings.

Original languageEnglish
Pages (from-to)1003-1022
Number of pages20
JournalInternational Journal on Artificial Intelligence Tools
Volume17
Issue number5
DOIs
StatePublished - Oct 2008

ASJC Scopus Subject Areas

  • Artificial Intelligence

Keywords

  • Clustering
  • XML

Fingerprint

Dive into the research topics of 'Structural classification of XML documents using multisets'. Together they form a unique fingerprint.

Cite this