Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Tagged Documents Co-Clustering

Oct 14, 2021

Gaëlle Candel, David Naccache

Figure 1 for Tagged Documents Co-Clustering

Figure 2 for Tagged Documents Co-Clustering

Figure 3 for Tagged Documents Co-Clustering

Figure 4 for Tagged Documents Co-Clustering

Share this with someone who'll enjoy it:

Abstract:Tags are short sequences of words allowing to describe textual and non-texual resources such as as music, image or book. Tags could be used by machine information retrieval systems to access quickly a document. These tags can be used to build recommender systems to suggest similar items to a user. However, the number of tags per document is limited, and often distributed according to a Zipf law. In this paper, we propose a methodology to cluster tags into conceptual groups. Data are preprocessed to remove power-law effects and enhance the context of low-frequency words. Then, a hierarchical agglomerative co-clustering algorithm is proposed to group together the most related tags into clusters. The capabilities were evaluated on a sparse synthetic dataset and a real-world tag collection associated with scientific papers. The task being unsupervised, we propose some stopping criterion for selectecting an optimal partitioning.

* 15 pages, submitted and accepted to the 2021 World Congress in Computer Science, Computer Engineering, & Applied Computing (CSCE'21) - track ICAI21

View paper on

Share this with someone who'll enjoy it:

Title:Tagged Documents Co-Clustering

Paper and Code