Assessing the Multilingual Corpus Extractor for Urdu Environmental Corpora: Overcoming Data Scarcity
DOI:
https://doi.org/10.58932/MULK0013Keywords:
Urdu NLP, Corpus Linguistics, Low-Resource Languages, Nastaliq Script Preservation, Environmental Discourse, EcolinguisticsAbstract
Computational linguistics for Urdu currently faces a significant barrier where standard extraction tools fracture the native Nastaliq script and render meaningful text into dissociated characters. This study validates the Multilingual Corpus Extractor or MCE as a solution to this infrastructural deficit. We tested this tool by constructing a corpus of the "2024 to 2025 Pakistan Smog Crisis". The results provide two primary contributions to the field. First the study establishes a technical benchmark regarding data hygiene as 49.3% of raw data extracted from vernacular news sites consists of non-linguistic noise that requires removal. Second the MCE successfully preserved 100% of the script ligatures whereas standard tools failed to maintain semantic unity. This technical success enabled a sociolinguistic analysis which revealed that Urdu media attributes environmental damage to specific human actors while English media frames the smog as a passive meteorological event. We conclude that accurate script preservation is not merely a technical repair but a necessary step to accurately record the history of the Global South.