The Use Of Generative Ai Tools In Corpus Linguistics: A Comparative Analysis Based On Uzbek And Russian Language Materials
Main Article Content
Abstract
As Corpus linguistics has not been studied fully yet, the purpose of the article is to find out how accurately generative AI tools can assist in corpus-based linguistic analysis of Uzbek and Russian language materials, and what methodological limitations emerge in low-resource and relatively high-resource language contexts. Moreover, the study is based on a qualitative review and comparative synthesis of recent research on large language models, corpus based AI systems, Uzbek computational linguistics, semantic tagging, tense annotation, web-corpus neologism extraction, and Russian corpus traditions. The findings show that Russian, as a relatively high-resource language with longer-standing corpus infrastructure, benefits from more stable morphological, syntactic, and semantic processing tools. Uzbek, by contrast, is undergoing active corpus development, with growing research on semantic analyzers, author corpora, grammatical annotation, and web-based lexical databases. Generative AI can support Uzbek corpus linguistics by assisting in query generation, preliminary annotation, concordance interpretation, translation alignment, metadata enrichment, and educational material generation. Nevertheless, the study emphasizes that generative AI cannot replace linguistically verified corpus methods. Problems such as hallucination, inconsistent corpus grounding, bias toward high-resource languages, inaccurate treatment of agglutinative morphology, and unreliable semantic interpretation require human expert validation. The article argues for a hybrid model in which generative AI functions as an assistant to corpus linguists rather than as an autonomous analytical authority. Such a model is especially important for Uzbek, where corpus resources are still developing and where careful integration of linguistic expertise, corpus design, and AI-based automation can accelerate the creation of reliable digital language resources.