Methods for discovering patterns based on cluster analysis and big data of secondary education institutions

 
Audio is AI-generated
0

Abstract

Context and relevance. The digitalization of education generates large volumes of learning-activity data containing patterns inaccessible to traditional statistical methods limited to aggregated performance indicators. Deep learning can extract latent patterns from data structure. Objective. To develop and compare student’s vector representation approaches from examination results for cluster analysis and latent-pattern discovery. Hypothesis. Big data on examination results contain patterns not reducible to aggregated performance indicators, and neural-network methods can capture them as vector representations suitable for clustering. Methods and materials. Data from the GIA-9 and GIA-11 examinations for 2023–2025 were used (11 subjects, approximately 5,000 students and 500,000 interactions). Three approaches were compared: a baseline on aggregated scores (Score stats), a variational autoencoder (VAE), and a masked autoencoder (MAE), including a version trained with a contrastive loss. Clustering was performed with HDBSCAN after dimensionality reduction (PCA); quality was assessed with internal metrics (silhouette coefficient, DBCV) and t-SNE visualization. Results. Score stats yields a stable partition reflecting examination type, subject set, and performance. VAE achieves high internal-metric values under aggressive dimensionality reduction, but its clusters are determined by the interaction-sequence length. MAE without a contrastive loss does not produce meaningful clusters; with it, the metrics reach the Score stats level, and the clusters group students by characteristic combinations of jointly chosen examinations. Conclusions. MAE with a contrastive loss extracts latent student characteristics, not reducible to aggregated indicators, from big educational data. The approaches are applicable to state examination data for categorizing students and supporting decision-making.

General Information

Keywords: educational data mining, educational data, clustering, deep learning, representation learning, autoencoder

Journal rubric: Data Analysis

Article type: scientific article

DOI: https://doi.org/10.17759/mda.2026160301

Received 11.06.2026

Revised 10.08.2026

Accepted

Published

For citation: Chaikin, G.A., Blekanov, I.S. (2026). Methods for discovering patterns based on cluster analysis and big data of secondary education institutions. Modelling and Data Analysis, 16(3), 7–29. (In Russ.). https://doi.org/10.17759/mda.2026160301

© Chaikin G.A., Blekanov I.S., 2026

License: CC BY-NC 4.0

References

  1. Айназаров, Р.Р., Вострокнутов, И.Е. (2025). Математическое моделирование кластеризации по результатам мониторинга деятельности образовательных организаций высшего образования. Computational Nanotechnology, 12(5), 118–128. https://doi.org/10.33693/2313-223X-2025-12-5-118-128
    Ainazarov, R.R., Vostroknutov, I.E. (2025). Mathematical modeling of clustering based on results of monitoring the activities of higher education institutions. Computational Nanotechnology, 12(5), 118–128. (In Russ.). https://doi.org/10.33693/2313-223X-2025-12-5-118-128
  2. Пак, Н.И., Клунникова, М.М. (2022). Кластерный подход к критериальному оцениванию качества образовательного результата обучаемого. Вестник РУДН. Серия: Информатизация образования, 19(3), 196–207. https://doi.org/10.22363/2312-8631-2022-19-3-196-207
    Pak, N.I., Klunnikova, M.M. (2022). Cluster approach to criteria-based assessment of the quality of educational outcomes of learners. RUDN Journal of Informatization in Education, 19(3), 196–207. (In Russ.). https://doi.org/10.22363/2312-8631-2022-19-3-196-207
  3. Юрьева, Н.Е. (2025). Искусственный интеллект в психодиагностике: когнитивные состояния в цифровой образовательной среде. Моделирование и анализ данных, 15(3), 47–55. https://doi.org/10.17759/mda.2025150303 
    Yuryeva, N.E. (2025). Artificial intelligence in psychodiagnostics: cognitive states in a digital educational environment. Modelling and Data Analysis, 15(3), 47–55. (In Russ.). https://doi.org/10.17759/mda.2025150303
  4. Ansari, A., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S.S., Pineda Arango, S., Kapoor, S., Zschiegner, J., Maddix, D.C., Wang, H., Mahoney, M.W., Torkkola, K., Wilson, A.G., Bohlke-Schneider, M., Flunkert, V. (2024). Chronos: Learning the language of time series. Transactions on Machine Learning Research. https://doi.org/10.48550/arXiv.2403.07815
  5. Barbeiro, L., Gomes, A., Correia, F., Bernardino, J. (2024). A review of educational data mining trends. Procedia Computer Science, 237, 88–95. https://doi.org/10.1016/j.procs.2024.05.083
  6. Choi, W., Lam, C.T., Mendes, A.J. (2025). Comparison of data imputation performance in deep generative models for educational tabular missing data. In: Proceedings of the 18th International Conference on Educational Data Mining (pp. 133–142). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.15870169
  7. Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
  8. Dutt, A., Aghabozrgi, S., Ismail, M.B., Mahroeian, H. (2015). Clustering algorithms applied in educational data mining. International Journal of Information and Electronics Engineering, 5(2), 112–116. https://doi.org/10.7763/IJIEE.2015.V5.513
  9. Dutt, A., Ismail, M., Herawan, T. (2017). A systematic review on educational data mining. IEEE Access, 5, 15991–16005. https://doi.org/10.1109/ACCESS.2017.2654247
  10. Freire, G., Curi, M. (2024). Masked autoencoder transformer for missing data imputation of PISA. In Artificial intelligence in education. Posters and late breaking results, workshops and tutorials, industry and innovation tracks, practitioners, doctoral consortium and blue sky (pp. 364–372). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-64315-6_33
  11. Guo, J., Fan, W., Amayri, M., Bouguila, N. (2025). Deep clustering analysis via variational autoencoder with Gamma mixture latent embeddings. Neural Networks, 183, 106979. https://doi.org/10.1016/j.neunet.2024.106979
  12. Hernández-Blanco, A., Herrera-Flores, B., Tomás, D., Navarro-Colorado, B. (2019). A systematic review of deep learning approaches to educational data mining. Complexity, 2019, 1306039. https://doi.org/10.1155/2019/1306039
  13. Huang, C., He, G. (2025). Text clustering as classification with LLMs. In: Proceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region (pp. 374–384). ACM. https://doi.org/10.1145/3767695.3769519
  14. Kostopoulos, G., Fazakis, N., Kotsiantis, S., Dimakopoulos, Y. (2025). Enhancing semi-supervised learning in educational data mining through synthetic data generation using tabular variational autoencoder. Algorithms, 18(10), 663. https://doi.org/10.3390/a18100663
  15. Li, Z., Rao, Z., Pan, L., Wang, P., Xu, Z. (2023). Ti-MAE: Self-supervised masked time series autoencoders. arXiv. https://doi.org/10.48550/arXiv.2301.08871
  16. Lin, Y., Chen, H., Xia, W., Lin, F., Wang, Z., Liu, Y. (2025). A comprehensive survey on deep learning techniques in educational data mining. Data Science and Engineering. https://doi.org/10.1007/s41019-025-00303-z
  17. Lu, Y., Li, H., Li, Y., Lin, Y., Peng, X. (2024). A survey on deep clustering: From the prior perspective. Vicinagearth, 1, 4. https://doi.org/10.1007/s44336-024-00001-w
  18. Lu, Y., Yeom, S., Maktoubian, J., Rahman, M., Kim, S.-H. (2025). Improve student risk prediction with clustering techniques: A systematic review in education data mining. Education Sciences, 15(12), 1695. https://doi.org/10.3390/educsci15121695
  19. Paaßen, B., Dywel, M., Fleckenstein, M., Pinkwart, N. (2022). Sparse factor autoencoders for item response theory. In: Proceedings of the 15th International Conference on Educational Data Mining (pp. 17–26). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.6853067
  20. Saïdi, I., Durand, N., Flouvat, F. (2025). Analysis of students' attempts trajectories in learning programming. In: Proceedings of the 18th International Conference on Educational Data Mining (pp. 66–76). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.15870155
  21. Scarlatos, A., Brinton, C., Lan, A. (2022). Process-BERT: A framework for representation learning on educational process data. In: Proceedings of the 15th International Conference on Educational Data Mining (pp. 715–719). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.6853006
  22. Shrivastava, A., Rameshan, R., Agnihotri, S. (2026). Robust representation learning in masked autoencoders. arXiv. https://doi.org/10.48550/arXiv.2602.03531
  23. Viswanathan, V., Gashteovski, K., Lawrence, C., Wu, T., Neubig, G. (2024). Large language models enable few-shot clustering. Transactions of the Association for Computational Linguistics, 12, 321–333. https://doi.org/10.1162/tacl_a_00648
  24. Wei, Y., Carvalho, P., Stamper, J. (2025). KCluster: An LLM-based clustering approach to knowledge component discovery. In: Proceedings of the 18th International Conference on Educational Data Mining (pp. 228–240). International Educational Data Mining Society. https://doi.org/10.5281/zenodo.15870196
  25. Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., Sahoo, D. (2024). Unified training of universal time series forecasting transformers. In: Proceedings of the 41st International Conference on Machine Learning (ICML 2024). https://doi.org/10.48550/arXiv.2402.02592
  26. Zhao, M., Dong, X. (2024). Evaluation of deep clustering for assessing undergraduate understanding in ideological and political education: Data-driven analytics. In: Genetic and evolutionary computing (pp. 103–111). Springer Nature Singapore. https://doi.org/10.1007/978-981-97-0068-4_10

Information About the Authors

Georgii A. Chaikin, Postgraduate Student, Saint Petersburg State University, St.Petersburg, Russian Federation, ORCID: https://orcid.org/0009-0007-9956-9457, e-mail: st061320@student.spbu.ru

Ivan S. Blekanov, Candidate of Science (Engineering), Associate Professor, Head of the Department of Programming Technology, Saint Petersburg State University, St.Petersburg, Russian Federation, ORCID: https://orcid.org/0000-0002-7305-1429, e-mail: i.blekanov@spbu.ru

Contribution of the authors

Georgii A. Chaikin — collection, processing and analysis of data; conduct of the study; visualization of the research results; writing and formatting the manuscript.

Ivan S. Blekanov — study design; data collection; supervision of the research process.

All authors participated in the discussion of the results and approved the final text of the manuscript.

Conflict of interest

The authors declare no conflict of interest.

Metrics

 Web Views

Whole time: 0
Previous month: 0
Current month: 0

 PDF Downloads

Whole time: 0
Previous month: 0
Current month: 0

 Total

Whole time: 0
Previous month: 0
Current month: 0