Patrick Amadeus Irawan
I am a Ph.D. student at MBZUAI, advised by Alham Fikri Aji. Previously, I was a Research Engineer at Singapore Management University with Chong-Wah Ngo, and received my B.S. in Computer Science from Institut Teknologi Bandung, advised by Ayu Purwarianti.
I build evaluation and post-training methods for vision-language and video models, with a current focus on interactive environments. My overarching goal is to build robust, self-improving multimodal agents.
My work has appeared at CVPR, NAACL, EMNLP, COLING, and AACL, and received the Best Theme Paper Award at NAACL 2025.
I am looking for 2027 Research Scientist, Applied Scientist, and Research Engineer internships. Feel free to reach out!
[first.last]@mbzuai.ac.ae
Research Interests
- Multimodal reasoning and evaluation: finding where models fail, across tasks, languages, and cultures.
- Multimodal post-training: adding capabilities efficiently without forgetting existing ones.
- Interactive environments and world models: whether models can simulate goal-directed actions in games, sandboxes, and open worlds.
News
- 2026: Regional Adaptation is accepted to AACL 2026.
- Oct. 2026: Ego2Act, a benchmark for goal-directed egocentric video generation, is on arXiv. [Project]
- Apr. 2026: LinguDistill is out on arXiv.
- Feb. 2026: M4-RAG (CVPR 2026) and Confused Tourists (CVPR 2026 Findings) are accepted.
- Oct. 2025: Entropy2Vec is accepted to the MRL Workshop at EMNLP 2025.
- Jul. 2025: Seeing Culture is accepted to EMNLP 2025.
- Apr. 2025: WorldCuisines receives the Best Theme Paper Award at NAACL 2025.
- Mar. 2025: Admitted to the MBZUAI Ph.D. program in NLP.
- Nov. 2024: VQA-NLE data generation is accepted to COLING 2025 (Oral).
Selected Publications (Full Publications)
* equal contribution
Interactive Environments and World Models
-
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Preprint, 2026. [Project]
Multimodal Post-Training
-
LinguDistill: Recovering Linguistic Ability in Vision Language Models via Selective Cross-Modal Distillation
Preprint, 2026. [arXiv] -
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
International Conference on Computational Linguistics (COLING), 2025. Oral [arXiv] [Code]
Multimodal Reasoning and Evaluation
-
Counting to Four is still a Chore for VLMs
Computer Vision and Pattern Recognition Conference (CVPR) Workshop, 2026. [arXiv] [Code] -
Vision Language Models are Confused Tourists
Computer Vision and Pattern Recognition Conference (CVPR), 2026 Findings. [arXiv] -
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
Computer Vision and Pattern Recognition Conference (CVPR), 2026. [arXiv] -
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
North American Chapter of the Association for Computational Linguistics (NAACL), 2025. Best Theme Paper [arXiv] [Code]
Experience
- Ph.D. Student, MBZUAI, 2025 – present
- Research Engineer, Singapore Management University, 2025
- Software Engineer, IT Bauschmiede, 2024 – 2025
- Data Science and SWE Intern, Supertype, Blibli, Ruangguru, 2022 – 2023
Honors & Services
- Reviewer, CVPR 2026, ACL Rolling Review (ARR) 2025, NLPCC 2025
- Scholarship, Columbia Machine Learning Summer School, 2026
- Best Theme Paper Award, NAACL 2025
- OpenAI Micro Grant, 2025
- Winner, BCG SEA Emeralds Case Competition, 2024
- Finalist, Gemastik XV Data Mining Division, 2022