Modelling and Data Analysis
2026. Vol. 16, no. 3, 7–29
https://doi.org/10.17759/mda.2026160301
ISSN: 2219-3758 / 2311-9454 (online)
Methods for discovering patterns based on cluster analysis and big data of secondary education institutions
Abstract
Context and relevance. The digitalization of education generates large volumes of learning-activity data containing patterns inaccessible to traditional statistical methods limited to aggregated performance indicators. Deep learning can extract latent patterns from data structure. Objective. To develop and compare student’s vector representation approaches from examination results for cluster analysis and latent-pattern discovery. Hypothesis. Big data on examination results contain patterns not reducible to aggregated performance indicators, and neural-network methods can capture them as vector representations suitable for clustering. Methods and materials. Data from the GIA-9 and GIA-11 examinations for 2023–2025 were used (11 subjects, approximately 5,000 students and 500,000 interactions). Three approaches were compared: a baseline on aggregated scores (Score stats), a variational autoencoder (VAE), and a masked autoencoder (MAE), including a version trained with a contrastive loss. Clustering was performed with HDBSCAN after dimensionality reduction (PCA); quality was assessed with internal metrics (silhouette coefficient, DBCV) and t-SNE visualization. Results. Score stats yields a stable partition reflecting examination type, subject set, and performance. VAE achieves high internal-metric values under aggressive dimensionality reduction, but its clusters are determined by the interaction-sequence length. MAE without a contrastive loss does not produce meaningful clusters; with it, the metrics reach the Score stats level, and the clusters group students by characteristic combinations of jointly chosen examinations. Conclusions. MAE with a contrastive loss extracts latent student characteristics, not reducible to aggregated indicators, from big educational data. The approaches are applicable to state examination data for categorizing students and supporting decision-making.
General Information
Keywords: educational data mining, educational data, clustering, deep learning, representation learning, autoencoder
Journal rubric: Data Analysis
Article type: scientific article
DOI: https://doi.org/10.17759/mda.2026160301
Received 11.06.2026
Revised 10.08.2026
Accepted
Published
For citation: Chaikin, G.A., Blekanov, I.S. (2026). Methods for discovering patterns based on cluster analysis and big data of secondary education institutions. Modelling and Data Analysis, 16(3), 7–29. (In Russ.). https://doi.org/10.17759/mda.2026160301
© Chaikin G.A., Blekanov I.S., 2026
License: CC BY-NC 4.0
References
- Айназаров, Р.Р., Вострокнутов, И.Е. (2025). Математическое моделирование кластеризации по результатам мониторинга деятельности образовательных организаций высшего образования. Computational Nanotechnology, 12(5), 118–128. https://doi.org/10.33693/2313-223X-2025-12-5-118-128
Ainazarov, R.R., Vostroknutov, I.E. (2025). Mathematical modeling of clustering based on results of monitoring the activities of higher education institutions. Computational Nanotechnology, 12(5), 118–128. (In Russ.). https://doi.org/10.33693/2313-223X-2025-12-5-118-128 - Пак, Н.И., Клунникова, М.М. (2022). Кластерный подход к критериальному оцениванию качества образовательного результата обучаемого. Вестник РУДН. Серия: Информатизация образования, 19(3), 196–207. https://doi.org/10.22363/2312-8631-2022-19-3-196-207
Pak, N.I., Klunnikova, M.M. (2022). Cluster approach to criteria-based assessment of the quality of educational outcomes of learners. RUDN Journal of Informatization in Education, 19(3), 196–207. (In Russ.). https://doi.org/10.22363/2312-8631-2022-19-3-196-207 - Юрьева, Н.Е. (2025). Искусственный интеллект в психодиагностике: когнитивные состояния в цифровой образовательной среде. Моделирование и анализ данных, 15(3), 47–55. https://doi.org/10.17759/mda.2025150303
Yuryeva, N.E. (2025). Artificial intelligence in psychodiagnostics: cognitive states in a digital educational environment. Modelling and Data Analysis, 15(3), 47–55. (In Russ.). https://doi.org/10.17759/mda.2025150303 - Ansari, A., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S.S., Pineda Arango, S., Kapoor, S., Zschiegner, J., Maddix, D.C., Wang, H., Mahoney, M.W., Torkkola, K., Wilson, A.G., Bohlke-Schneider, M., Flunkert, V. (2024). Chronos: Learning the language of time series. Transactions on Machine Learning Research. https://doi.org/10.48550/arXiv.2403.07815
- Barbeiro, L., Gomes, A., Correia, F., Bernardino, J. (2024). A review of educational data mining trends. Procedia Computer Science, 237, 88–95. https://doi.org/10.1016/j.procs.2024.05.083
- Choi, W., Lam, C.T., Mendes, A.J. (2025). Comparison of data imputation performance in deep generative models for educational tabular missing data. In: Proceedings of the 18th International Conference on Educational Data Mining (pp. 133–142). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.15870169
- Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
- Dutt, A., Aghabozrgi, S., Ismail, M.B., Mahroeian, H. (2015). Clustering algorithms applied in educational data mining. International Journal of Information and Electronics Engineering, 5(2), 112–116. https://doi.org/10.7763/IJIEE.2015.V5.513
- Dutt, A., Ismail, M., Herawan, T. (2017). A systematic review on educational data mining. IEEE Access, 5, 15991–16005. https://doi.org/10.1109/ACCESS.2017.2654247
- Freire, G., Curi, M. (2024). Masked autoencoder transformer for missing data imputation of PISA. In Artificial intelligence in education. Posters and late breaking results, workshops and tutorials, industry and innovation tracks, practitioners, doctoral consortium and blue sky (pp. 364–372). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-64315-6_33
- Guo, J., Fan, W., Amayri, M., Bouguila, N. (2025). Deep clustering analysis via variational autoencoder with Gamma mixture latent embeddings. Neural Networks, 183, 106979. https://doi.org/10.1016/j.neunet.2024.106979
- Hernández-Blanco, A., Herrera-Flores, B., Tomás, D., Navarro-Colorado, B. (2019). A systematic review of deep learning approaches to educational data mining. Complexity, 2019, 1306039. https://doi.org/10.1155/2019/1306039
- Huang, C., He, G. (2025). Text clustering as classification with LLMs. In: Proceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (pp. 374–384). ACM. https://doi.org/10.1145/3767695.3769519
- Kostopoulos, G., Fazakis, N., Kotsiantis, S., Dimakopoulos, Y. (2025). Enhancing semi-supervised learning in educational data mining through synthetic data generation using tabular variational autoencoder. Algorithms, 18(10), 663. https://doi.org/10.3390/a18100663
- Li, Z., Rao, Z., Pan, L., Wang, P., Xu, Z. (2023). Ti-MAE: Self-supervised masked time series autoencoders. arXiv. https://doi.org/10.48550/arXiv.2301.08871
- Lin, Y., Chen, H., Xia, W., Lin, F., Wang, Z., Liu, Y. (2025). A comprehensive survey on deep learning techniques in educational data mining. Data Science and Engineering. https://doi.org/10.1007/s41019-025-00303-z
- Lu, Y., Li, H., Li, Y., Lin, Y., Peng, X. (2024). A survey on deep clustering: From the prior perspective. Vicinagearth, 1, 4. https://doi.org/10.1007/s44336-024-00001-w
- Lu, Y., Yeom, S., Maktoubian, J., Rahman, M., Kim, S.-H. (2025). Improve student risk prediction with clustering techniques: A systematic review in education data mining. Education Sciences, 15(12), 1695. https://doi.org/10.3390/educsci15121695
- Paaßen, B., Dywel, M., Fleckenstein, M., Pinkwart, N. (2022). Sparse factor autoencoders for item response theory. In: Proceedings of the 15th International Conference on Educational Data Mining (pp. 17–26). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.6853067
- Saïdi, I., Durand, N., Flouvat, F. (2025). Analysis of students' attempts trajectories in learning programming. In: Proceedings of the 18th International Conference on Educational Data Mining (pp. 66–76). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.15870155
- Scarlatos, A., Brinton, C., Lan, A. (2022). Process-BERT: A framework for representation learning on educational process data. In: Proceedings of the 15th International Conference on Educational Data Mining (pp. 715–719). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.6853006
- Shrivastava, A., Rameshan, R., Agnihotri, S. (2026). Robust representation learning in masked autoencoders. arXiv. https://doi.org/10.48550/arXiv.2602.03531
- Viswanathan, V., Gashteovski, K., Lawrence, C., Wu, T., Neubig, G. (2024). Large language models enable few-shot clustering. Transactions of the Association for Computational Linguistics, 12, 321–333. https://doi.org/10.1162/tacl_a_00648
- Wei, Y., Carvalho, P., Stamper, J. (2025). KCluster: An LLM-based clustering approach to knowledge component discovery. In: Proceedings of the 18th International Conference on Educational Data Mining (pp. 228–240). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.15870196
- Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., Sahoo, D. (2024). Unified training of universal time series forecasting transformers. In: Proceedings of the 41st International Conference on Machine Learning (ICML 2024). https://doi.org/10.48550/arXiv.2402.02592
- Zhao, M., Dong, X. (2024). Evaluation of deep clustering for assessing undergraduate understanding in ideological and political education: Data-driven analytics. In: Genetic and evolutionary computing (pp. 103–111). Springer Nature Singapore. https://doi.org/10.1007/978-981-97-0068-4_10
Information About the Authors
Contribution of the authors
Georgii A. Chaikin — collection, processing and analysis of data; conduct of the study; visualization of the research results; writing and formatting the manuscript.
Ivan S. Blekanov — study design; data collection; supervision of the research process.
All authors participated in the discussion of the results and approved the final text of the manuscript.
Conflict of interest
The authors declare no conflict of interest.
Metrics
Web Views
Whole time: 0
Previous month: 0
Current month: 0
PDF Downloads
Whole time: 0
Previous month: 0
Current month: 0
Total
Whole time: 0
Previous month: 0
Current month: 0