Exploring Cross-Lingual Knowledge Transfer via Transliteration-Based MLM Fine-Tuning for Critically Low-resource Chakma Language

Published in Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025), pp. 280–291, 2025

This paper introduces a native-speaker-validated corpus of Bangla-transliterated Chakma and studies masked-language-model fine-tuning for this critically low-resource language. Fine-tuned multilingual models outperform their pretrained counterparts, reaching up to 73.54% token accuracy and a perplexity of 2.90. The study also examines the effects of data quality and the limitations of OCR for morphologically rich Indic scripts.

View Publication