avatar

Salsabila Zahirah Pranida

M.Sc @ NLP
MBZUAI
irasalsabila (at) gmail (dot) com


Curriculum Vitae

Researcher and academic mentor working on low-resource, multilingual, and dialectal language technologies.

Education

2024–2026

M.Sc. in Natural Language Processing

Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi

Research interests: computational social science, multilingual and dialectal NLP, low-resource language technology, conversational datasets, cultural and linguistic evaluation, and speech technologies.

Thesis: IndoDialect–A Multi-Dialect Conversational Dataset for Indonesian Local Languages.

2018–2022

S.Kom in Informatics

Universitas Islam Indonesia, Yogyakarta

Final Project: Customer Satisfaction Sentiment Analysis on Courier Services Using BiLSTM and BiGRU.

Research Experience

2025–2026

IndoDialect — A Multi-Dialect Conversational Dataset

Built the first large-scale conversational dialect benchmark for Indonesian local languages: 14.2K multi-turn dialogues across 11 dialects and four languages spanning three major islands.

Designed a dual-stream human curation pipeline with perturbation validation scoring 99.6% quality and native-speaker evaluations exceeding 95% cultural authenticity and coherence. Benchmarked 10+ models across three tasks.

Published: EMNLP 2026.

2025–2026

Persona and Gender Effects in Indonesian Toxicity Classification

Evaluated 10 Indonesian-capable LLMs across four persona conditions on 7,524 IndoDiscourse posts with 24,655 gender-disaggregated human judgments.

Used crossed mixed-effects logistic regression, permutation tests, semantic clustering, and sensitivity analysis to show how persona prompts change calibration without selectively increasing demographic label agreement.

Published: PANDORA 2026 workshop, co-located with EMNLP 2026.

2025

ASR Under Noise: Exploring Robustness for Sundanese and Javanese

Benchmarked four Whisper architectures across 120+ hours of speech, 10 SNR levels, and 24 AudioSet environmental noise classes.

NoiseTrain and SpecAugment reduced low-SNR WER, including Sundanese Medium from 199.1% to 56.2% and Javanese Large-v3 to 41.4%.

Published: WiNLP 2025 at EMNLP 2025. Best Paper Award.

2024

Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages

Constructed a culturally grounded StoryCloze benchmark with 3.3K narratives across 12 cultural domains in Javanese and Sundanese.

Built an XLM-R filtering pipeline over 12K synthetic stories and fine-tuned six regional and open LLMs with QLoRA to improve cultural commonsense reasoning.

Published: MRL 2025, co-located with EMNLP 2025.

Awards & Grants

2025

OpenAI Mini-Grant Recipient

MBZUAI Department of NLP

Awarded a $1,500 internal grant supporting thesis research on multilingual and dialectal language modeling.

2025

Cohere Labs Catalyst Grant Recipient

Cohere

Awarded a $1,000 grant supporting thesis research on multilingual and dialectal language modeling.

Jun–Aug
2025

DDSA Visit Grant Recipient

Danish Data Science Academy

Awarded a 15,000 DKK grant for collaborative research at the University of Copenhagen under Dr. Daniel Hershcovich.

2025

WiNLP Best Paper Award

Recognized for research on robust automatic speech recognition for Sundanese and Javanese.


Powered by Jekyll and Minimal Light theme.