Researcher and academic mentor working on low-resource, multilingual, and dialectal language technologies.
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi
Research interests: computational social science, multilingual and dialectal NLP, low-resource language technology, conversational datasets, cultural and linguistic evaluation, and speech technologies.
Thesis: IndoDialect–A Multi-Dialect Conversational Dataset for Indonesian Local Languages.
Universitas Islam Indonesia, Yogyakarta
Final Project: Customer Satisfaction Sentiment Analysis on Courier Services Using BiLSTM and BiGRU.
Built the first large-scale conversational dialect benchmark for Indonesian local languages: 14.2K multi-turn dialogues across 11 dialects and four languages spanning three major islands.
Designed a dual-stream human curation pipeline with perturbation validation scoring 99.6% quality and native-speaker evaluations exceeding 95% cultural authenticity and coherence. Benchmarked 10+ models across three tasks.
Published: EMNLP 2026.
Evaluated 10 Indonesian-capable LLMs across four persona conditions on 7,524 IndoDiscourse posts with 24,655 gender-disaggregated human judgments.
Used crossed mixed-effects logistic regression, permutation tests, semantic clustering, and sensitivity analysis to show how persona prompts change calibration without selectively increasing demographic label agreement.
Published: PANDORA 2026 workshop, co-located with EMNLP 2026.
Benchmarked four Whisper architectures across 120+ hours of speech, 10 SNR levels, and 24 AudioSet environmental noise classes.
NoiseTrain and SpecAugment reduced low-SNR WER, including Sundanese Medium from 199.1% to 56.2% and Javanese Large-v3 to 41.4%.
Published: WiNLP 2025 at EMNLP 2025. Best Paper Award.
Constructed a culturally grounded StoryCloze benchmark with 3.3K narratives across 12 cultural domains in Javanese and Sundanese.
Built an XLM-R filtering pipeline over 12K synthetic stories and fine-tuned six regional and open LLMs with QLoRA to improve cultural commonsense reasoning.
Published: MRL 2025, co-located with EMNLP 2025.
MBZUAI Department of NLP
Awarded a $1,500 internal grant supporting thesis research on multilingual and dialectal language modeling.
Cohere
Awarded a $1,000 grant supporting thesis research on multilingual and dialectal language modeling.
Danish Data Science Academy
Awarded a 15,000 DKK grant for collaborative research at the University of Copenhagen under Dr. Daniel Hershcovich.
Recognized for research on robust automatic speech recognition for Sundanese and Javanese.
Powered by Jekyll and Minimal Light theme.