Assessing the Multilingual Corpus Extractor for Urdu Environmental Corpora: Overcoming Data Scarcity

Assessing the Multilingual Corpus Extractor for Urdu Environmental Corpora: Overcoming Data Scarcity

Authors

  • Muhammad Ahmad Faculty of Humanities, Higher School of Economics, Russia https://orcid.org/0009-0000-3207-3499
  • Amna Mehtab Faculty of Humanities, Higher School of Economics, Russia
  • Fatima Mehtab Department of Education, Government College University Faisalabad, Pakistan.

DOI:

https://doi.org/10.58932/MULK0013

Keywords:

Urdu NLP, Corpus Linguistics, Low-Resource Languages, Nastaliq Script Preservation, Environmental Discourse, Ecolinguistics

Abstract

Computational linguistics for Urdu currently faces a significant barrier where standard extraction tools fracture the native Nastaliq script and render meaningful text into dissociated characters. This study validates the Multilingual Corpus Extractor or MCE as a solution to this infrastructural deficit. We tested this tool by constructing a corpus of the "2024 to 2025 Pakistan Smog Crisis". The results provide two primary contributions to the field. First the study establishes a technical benchmark regarding data hygiene as 49.3% of raw data extracted from vernacular news sites consists of non-linguistic noise that requires removal. Second the MCE successfully preserved 100% of the script ligatures whereas standard tools failed to maintain semantic unity. This technical success enabled a sociolinguistic analysis which revealed that Urdu media attributes environmental damage to specific human actors while English media frames the smog as a passive meteorological event. We conclude that accurate script preservation is not merely a technical repair but a necessary step to accurately record the history of the Global South.

Downloads

Published

2026-06-29

How to Cite

Ahmad, M., Mehtab, A., & Mehtab, F. . (2026). Assessing the Multilingual Corpus Extractor for Urdu Environmental Corpora: Overcoming Data Scarcity. Journal of Advanced Corpus-Oriented Research, 2(1), 77–94. https://doi.org/10.58932/MULK0013

Issue

Section

Articles
Loading...