Linguistic Foundations for Automatic Simplification of Kazakh Texts
DOI:
https://doi.org/10.17507/tpls.1608.13Keywords:
text simplification, lexical simplification, syntactic simplification, ATSAbstract
Universal and multilingual large language models cannot correctly simplify Kazakh texts in accordance with internal linguistic patterns. Using LLM without a linguistically sound simplification model leads to uncontrolled generation and increases the risk of lexical and grammatical errors. Therefore, there is a need to create formalized linguistic patterns for Kazakh text simplification, which will facilitate the creation of automatic simplification systems within LLM. This can be achieved through the analysis of manually adapted materials. A corpus of 20 Kazakh texts was annotated by the authors for the analysis, each presented in three versions such as the original (O) and two successive adaptation levels (A1 and A2). At the A1 stage, syntactic transformations predominate, while A2 demonstrates lexical stabilization. The text size is comparable at all levels, which ensures control and reliability of the analysis. The process is formalized through a two-level tagging system: lexical (LS) and syntactic simplification (SS) that records each operation. Frequency analysis revealed a limited core of dominant techniques, supplemented by peripheral layers, forming structured polyoperational architecture. Transitions from A1 to A2 demonstrate a process shift toward lexical changes while maintaining discursive coherence. The results emphasize the importance of linguistic expertise, operation typologies, and manual annotation for the interpretability and quality of systems. Hybrid approaches with humans in the loop improve the accuracy and reliability of ATS.
References
Agrawal, S., & Carpuat, M. (2024). Do Text Simplification Systems Preserve Meaning? A Framework for Human Evaluation Using Reading Comprehension. Transactions of the Association for Computational Linguistics, 12, 432–448. https://doi.org/10.1162/tacl_a_00653
Al-Thanyyan, S. S., & Azmi, A. M. (2021). Automated text simplification: A survey. ACM Computing Surveys (CSUR), 54(2), 1–36. https://doi.org/10.1145/3442695
Alva-Manchego, F., Bingel, J., Paetzold, G. H., Scarton, C., & Specia, L. (2017). Learning how to simplify from explicit labeling of complex–simplified text pairs. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Vol. 1, pp. 295–305). Asian Federation of Natural Language Processing. Retrieved January 16, 2026, from https://aclanthology.org/I17-1030/
Alva-Manchego, F., Scarton, C., & Specia, L. (2020). Data-driven sentence simplification: Survey and benchmark. Computational Linguistics, 46(1), 135–187. https://doi.org/10.1162/coli_a_00370
Bott, S., & Saggion, H. (2014). Text Simplification Resources for Spanish. Language Resources and Evaluation, 48(1), 93–120. https://doi.org/10.1007/s10579-014-9265-4
Brunato, D., Dell'Orletta, F., & Venturi, G. (2022). Linguistically-Based Comparison of Different Approaches to Building Corpora for Text Simplification: A Case Study on Italian. Frontiers in Psychology, 13, 1–19. https://doi.org/10.3389/fpsyg.2022.707630
Cardon, R., & Grabar, N. (2020). French biomedical text simplification: When small and precise helps. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 710–716). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.62
Carroll, J., Minnen, G., Canning, Y., Devlin, S., & Tait, J. (1998). Practical simplification of English newspaper text to assist aphasic readers. In Proceedings of the AAAI-98 Workshop on Integrating Artificial Intelligence and Assistive Technology (pp. 7–10). Association for the Advancement of Artificial Intelligence. Retrieved January 16, 2026, from https://www.researchgate.net/publication/2740075_Practical_Simplification_of_English_Newspaper_Text_to_Assist_Aphasic_Readers
Chen, P., Rochford, J., Kennedy, D. N., Djamasbi, S., Fay, P., & Scott, W. (2017). Automatic text simplification for people with intellectual disabilities. Artificial Intelligence Science and Technology, 725–731. https://doi.org/10.1142/9789813206823_0091
Crossley, S. A., Allen, D., & McNamara, D. (2012). Text simplification and comprehensible input: A case for an intuitive approach. Language Teaching Research, 16(1), 89–108. https://doi.org/10.1177/1362168811423456
Coxhead, A., & Boutorwick, T. J. (2018). Longitudinal vocabulary development in an EMI international school context: Learners and texts in EAL, maths, and science. TESOL Quarterly, 52(4), 1000–1023. https://doi.org/10.1002/tesq.450
De Belder, J., & Moens, M. (2010). Text simplification for children. In Proceedings of the SIGIR Workshop on Accessible Search Systems (pp. 19–26). Retrieved January 16, 2026, from https://www.researchgate.net/publication/228736550_Text_simplification_for_children
Devaraj, A., Marshall, I., Wallace, B., & Li, J. (2021). Paragraph-Level Simplification of Medical Texts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4972–4984). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.naacl-main.395
Dras, M. (1999). Tree adjoining grammar and the reluctant paraphrasing of text [Doctoral thesis, Sydney: Macquarie University].
Ermakova, L., SanJuan, E., Kamps, J., Huet, S., Ovchinnikova, I., Nurbakova, D., Ara´ujo, S., Hannachi, R., Mathurin, E., Bellot, P. (2022). Overview of the CLEF 2022 SimpleText Lab: Automatic simplification of scientific texts. Lecture Notes in Computer Science, 13390, 470–494. https://doi.org/10.1007/978-3-031-13643-6_28
Evans, R., Orasan, C., & Dornescu, I. (2014). An evaluation of syntactic simplification rules for people with autism. In Proceedings of the 3rd Workshop on Predicting and Improving Text Readability for Target Reader Populations (pp. 131–140). Association for Computational Linguistics. https://doi.org/10.3115/v1/W14-1215
Flores, L., Huang, H., Shi, K., Chheang, S., & Cohan, A. (2023). Medical text simplification: optimizing for readability with unlikelihood training and reranked beam search decoding. Computer science, 1. https://doi.org/10.48550/arXiv.2310.11191
Gonzalez-Dios, I., Aranzabe, M. J., & Diaz de Ilarraza, A. (2018). The corpus of Basque simplified texts (CBST). Language Resources and Evaluation, 52(1), 217–247. https://doi.org/10.1007/s10579-017-9407-6
Grabar, N., & Saggion, H. (2022). Evaluation of Automatic Text Simplification: Where Are We Now, Where Should We Go From Here. Actes de la 29e Conférence sur le Traitement Automatique des Langues Naturelles, 1, 453–463. Retrieved January 16, 2026, from https://aclanthology.org/2022.jeptalnrecital-taln.47/
Kajiwara, T., Matsumoto, H., & Yamamoto, K. (2013). Selecting proper lexical paraphrase for children. In Proceedings of the 25th Conference on Computational Linguistics and Speech Processing (ROCLING 2013) (pp. 59–73). ACLCLP. Retrieved January 16, 2026, from https://aclanthology.org/O13-1007/
Kassenkhan, A. M., Mukazhanov, N. K., Nuralykyzy, S., & Kalpeyeva, Z. B. (2024). Text generation models for paraphrase on Kazakh language. Vestnik KazUTB, 1(22). https://doi.org/10.58805/kazutb.v.1.22-249
Kover, S. T., Haebig, E., Oakes, A., McDuffie, A., Hagerman, R. J., & Abbeduto, L. (2014). Sentence comprehension in boys with autism spectrum disorder. American journal of speech-language pathology, 23(3), 385–394. https://doi.org/10.1044/2014_AJSLP-13-0073
Koptient, A., Cardon, R., & Grabar, N. (2019). Simplification-induced transformations: Typology and some characteristics. In Proceedings of the 18th BioNLP Workshop and Shared Task (pp. 309–318). Association for Computational Linguistics. https://doi.org/10.18653/v1/W19-5033
Kumar, D., Mou, L., Golab, Ł., & Vechtomova, O. (2020). Iterative edit-based unsupervised sentence simplification. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7918–7928). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.707
Lai, A., & Tetreault, J. (2018). Discourse Coherence in the Wild: A Dataset, Evaluation and Methods. In Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue (pp. 214–223). Association for Computational Linguistics. https://doi.org/10.18653/v1/W18-5023
Martin, L., Fan, A., de la Clergerie, É., Bordes, A., & Sagot, B. (2020). MUSS: Multilingual unsupervised sentence simplification by mining paraphrases. Computation and Language. https://doi.org/10.48550/arXiv.2005.00352
Maddela, M., & Alva-Manchego, F. (2025). Adapting Sentence-Level Automatic Metrics for Document-Level Simplification. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 6444–6459). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-long.327
Mambetova, M., Karymkhan, A., & Abdikadyrova, А. (2024). Adapting kazakh texts for language learners: application of adaptation techniques. Eurasian Journal of Philology. Science and Education, 195(3), 58–67. https://doi.org/10.26577/EJPh.2024.v195.i3.ph06
Rello, L., Bayarri, C., Górriz, A., Baeza-Yates, R., Gupta, S., Kanvinde, G., Saggion, H., Bott, S., Carlini, R., & Topac, V. (2013). DysWebxia 2.0! More accessible text for people with dyslexia. In Proceedings of the 10th International Cross-Disciplinary Conference on Web Accessibility (pp. 1–2). Association for Computing Machinery. https://doi.org/10.1145/2461121.2461150
Ryan, M. J., Naous, T., & Xu, W. (2023). Revisiting non-English text simplification: A unified multilingual benchmark. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 4898–4927). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.269
Scholz, K., & Wenzel, M. (2025). Evaluating readability metrics for German medical text simplification. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 6049–6062). Association for Computational Linguistics. Retrieved January 16, 2026, from https://aclanthology.org/2025.coling-main.405/
Shardlow, M. (2014). A Survey of Automated Text Simplification. International Journal of Advanced Computer Science and Applications, 4, 58–70. https://doi.org/10.14569/SpecialIssue.2014.040109
Sun, R., Jin, H., & Wan, X. (2021). Document-level text simplification: Dataset, criteria and baseline. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 7997–8013). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.630
Štajner, S., & Saggion, H. (2018). Data-Driven Text Simplification. In Proceedings of the 27th International Conference on Computational Linguistics: Tutorial Abstracts (pp. 19–23). Association for Computational Linguistics. Retrieved January 16, 2026, from https://aclanthology.org/C18-3005/
Toleu, A., Tolegen, G., & Ualiyeva, I. (2025). Fine-tuning large language models for Kazakh text simplification. Applied Sciences, 15(15), 8344. https://doi.org/10.3390/app15158344
Tanprasert, T., & Kauchak, D. (2021). Flesch-Kincaid is not a text simplification evaluation metric. In Proceedings of the First Workshop on Natural Language Generation, Evaluation, and Metrics (GEM) (pp. 1–14). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.gem-1.1
Vásquez-Rodríguez, L., Shardlow, M., Przybyła, P., & Ananiadou, S. (2023). document-level text simplification with coherence evaluation. In Proceedings of the Second Workshop on Text Simplification, Accessibility and Readability (pp. 85–101). Retrieved January 16, 2026, from https://aclanthology.org/2023.tsar-1.9/
Watanabe, W. M., Junior, A. C., Uzêda, V. R., Fortes, R. P. de M., Pardo, T. A. S., & Aluísio, S. M. (2009). Facilita: Reading assistance for low-literacy readers. In Proceedings of the 27th ACM International Conference on Design of Communication (pp. 29–36). Association for Computing Machinery. https://doi.org/10.1145/1621995.1622002
Xu, W., Callison-Burch, C., & Napoles, C. (2015). Problems in current text simplification research: New data can help. Transactions of the Association for Computational Linguistics, 3, 283–297. https://doi.org/10.1162/tacl_a_00139
Xu, W., Napoles, C., Pavlick, E., Chen, Q., & Callison-Burch, C. (2016). Optimizing statistical machine translation for text simplification. Transactions of the Association for Computational Linguistics, 4, 401–415. https://doi.org/10.1162/tacl_a_00107
Yamaguchi, D., Miyata, R., Shimada, S., & Sato, S. (2023). Gauging the gap between human and machine text simplification through analytical evaluation of simplification strategies and errors. In Findings of the Association for Computational Linguistics: EACL 2023 (pp. 359–375). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-eacl.27
Yeshpanov, R., Polonskaya, A., & Varol, H. A. (2024). KazParC: Kazakh parallel corpus for machine translation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 9633–9644). ELRA and ICCL. https://doi.org/10.48550/arXiv.2403.19399