CS SEMINAR

The Making of Multilingual LLMs: Data & Evaluation

Speaker
Alham Fikri Aji, Associate Professor, Mohamed bin Zayed University of Artificial Intelligence
Chaired by
Dr NG Hwee Tou, Provost's Chair Professor, School of Computing
nght@comp.nus.edu.sg

14 Oct 2026 Wednesday, 03:30 PM to 04:30 PM

MR25, COM3 02-70

Abstract:

Data availability remains one of the central bottlenecks in building multilingual LLMs, particularly for languages underrepresented on the web and in existing NLP resources. Over the years, the community has addressed this challenge through translation, synthetic generation, task adaptation, and large-scale participatory efforts. Multilingual evaluation has evolved in parallel, moving from translated English benchmarks toward locally authored, culturally grounded, and continuously updated community resources. In this talk, we trace these developments and examine how different construction choices shape what models learn and what benchmarks make visible. We conclude by discussing persistent gaps and directions for building multilingual data and evaluations that remain meaningful in the frontier-model era.

Biodata:

Alham Fikri Aji is an Associate Professor of Natural Language Processing at MBZUAI (Mohamed bin Zayed University of Artificial Intelligence) and a Visiting Research Scientist at Google Research, where he works on Gemini’s multilingual capabilities. His research aims to make language technology more accessible to speakers of underrepresented languages. He has contributed to open multilingual models such as BLOOMZ, Jais, or Cendol, and to methods that make language models more efficient. His work also includes benchmarks such as NusaX, SEACrowd, and CVQA, which examine how well AI serves different languages and cultures. Beyond individual projects, he works with the Indonesian and Southeast Asian NLP communities to build shared resources and support researchers across the region.