Clustering methods for categorical time series and sequences : a scoping review
Abstract
<h4>Objective</h4>To provide an overview of clustering methods for categorical time series (CTS), a data structure common in epidemiology, sociology, biology, and marketing, and to support method selection according to data characteristics.<h4>Materials and methods</h4>We searched PubMed (via MEDLINE), Web of Science, and Google Scholar up to November 2024 for articles proposing and evaluating CTS clustering techniques. Methods were classified into three families-distance-based, feature-based, and model-based-and assessed for their ability to address challenges such as variable sequence length, multivariate data, continuous time, missing data, covariates, and large data volumes.<h4>Results</h4>Of 14,607 records retrieved, 124 articles describing 129 methods were included. Distance-based approaches, especially those using Optimal Matching, were most common, with 56 methods. We found 28 model-based methods, which covered a broader range of complex data structures such as multivariate data, continuous time and time-invariant covariates. We recorded 45 feature-based approaches, which were on average more scalable but less flexible. Fewer than half of the methods provided public implementations. A searchable Web application ( https://cts-clustering-scoping-review-7sxqj3sameqvmwkvnzfynz.streamlit.app/ ) was developed to support method selection.<h4>Discussion</h4>CTS clustering methods are highly heterogeneous in assumptions, capabilities, and scalability. Distance-based approaches dominate, but model-based methods offer richer modeling potential, while feature-based ones emphasize performance at the cost of flexibility.<h4>Conclusion</h4>This review highlights methodological diversity and gaps in CTS clustering. The proposed typology and Web application aim to help researchers choose appropriate methods to choose appropriate methods for their data.