COMMUNITY

BOARD

News

Regular Paper Accepted at the Top International Conference ACM International Conference on Information and Knowledge Management (CIKM) 2026

Author
College of Software Convergence
Date
2026-09-01
View
119

Files



The paper 'CLARE: Cluster-Number-Independent Latent Representation Learning with Graph Autoencoders for Deep Clustering', written by Kim Yong-dam (doctoral, first author) and Professor Jung Sung-won (corresponding author) of the Bigdata Processing & DB Lab, has been accepted for publication as a short research paper at the top international conference ACM International Conference on Information and Knowledge Management (CIKM) 2026 (acceptance rate: 236/764 = 30.9%).

Deep clustering is used to find structure in unlabelled data — organizing documents, segmenting customers, grouping images. In real analysis, however, the number of clusters is almost never known in advance. How many groups customers should be divided into, or how many topics a document collection contains, can only be determined by trying several candidate counts and comparing. The problem is that most recent high-performing deep clustering models build the cluster count into the loss function. Changing the count by one means retraining the model from scratch, and comparing several candidates means redoing the computation each time. On AgNews, the state-of-the-art model COTC takes more than 2,700 seconds for a single training run.

CLARE, proposed by the team, first computes a hybrid similarity using keyword-based TF-IDF (a sparse representation) together with bge-m3 (a dense representation) produced by a pretrained model. The design is intended to capture both keyword matching and semantic context. A kNN graph is built from this similarity, but rather than always connecting k neighbours, selective filtering keeps only sufficiently similar candidates as edges. The resulting kNN graph is then embedded using an improved graph autoencoder.

On top of this the team designed a new cluster loss function independent of the number of clusters. Extending the concept of the global clustering coefficient from network science, it exploits the property that if A is connected to B and B to C, then A and C are likely to belong to the same cluster. Since not every two-hop pair is in the same cluster, however, an IQR filter selects and reinforces only the upper outliers of the similarity distribution. Because neither the graph reconstruction loss nor the newly proposed clustering loss refers to the number of clusters, an embedding obtained from a single training run can be used for clustering comparisons across many specified cluster counts.

Validated on eight text datasets and three image datasets, CLARE recorded the best performance on two of four long-text and short-text settings without using the cluster count in training at all. On 20Newsgroups it achieved 80.8% accuracy, well ahead of COTC (67.2%). For images it was applied as is, changing only the encoder and the similarity computation, and recorded 94.5% accuracy on ImageNet-Dogs, where fine-grained classification is difficult — raising the previous best (72.6%) by about 22 points. In efficiency, sweeping the cluster count from 2 to 6 on AgNews took 38.1 times less time than retraining each time, and was more than 89.7 times faster on 20Newsgroups.

The research is significant in removing the premise of recent deep clustering that the number of clusters must be fixed in advance, and in presenting a representation learning approach that allows several candidates to be compared cheaply during exploratory data analysis.

Kim Yong-dam, the paper's first author and a doctoral student, said: "When analysing real data one almost never knows the number of clusters in advance, so it always felt awkward that the better-performing methods required that number to be fixed first. Taking the count out of the training objective and drawing the cluster signal from the graph structure itself let us keep accuracy while greatly reducing the cost of sweeping candidate counts. Going forward I want to extend this to subgraph learning for larger corpora and to multimodal deep clustering handling text and images together. I am grateful to Professor Jung Sung-won for supervising this research."

The ACM International Conference on Information and Knowledge Management (CIKM) is a globally prestigious international conference covering information retrieval, knowledge management, data mining, databases and artificial intelligence. CIKM is listed at a recognized IF of 3 among the outstanding international conferences in computer science under BK21, and is classified as a top conference in the Korean Institute of Information Scientists and Engineers' list of outstanding conferences in the software field. This year it will be held in Rome, Italy, from 7 to 11 November.

References:

-       35th ACM International Conference on Information and Knowledge Management (CIKM 2026)

-       Website: https://cikm2026.diag.uniroma1.it/