Patrick Amadeus Irawan
[first.last]@mbzuai.ac.ae · Google Scholar · Website · LinkedIn · Github
Research Interests
My research goal is to build robust, self-improving multimodal agents that perceive, reason, and act reliably across diverse environments. I develop evaluation and post-training methods for vision-language and video models, with a current focus on interactive environments. My prior work spans VLM post-training and knowledge distillation, multimodal reasoning and robustness, and multilingual and multicultural benchmarks.
Research Areas: multimodal reasoning and evaluation, multimodal post-training, interactive environments and world models, embodied AI
Education
Advisor: Alham Fikri Aji
computational neuroscience, tech management
visual question answering; synthetic data generation; reasoning LLM
Publications (* (Co-)first author † Second author)
-
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation*
video generation, world model, evaluation, multi agent system, embodied agentsPreprint 2026
-
LinguDistill: Recovering Linguistic Ability in Vision Language Models via Selective Cross-Modal Distillation*
VLM, SFT, post-training, knowledge distillation, architecturePreprint 2026
-
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
VLM, SFT, post-training, data curation, evaluation, multilingual, multimodalAACL 2026
-
Vision Language Models are Confused Tourists*
VLM, visual reasoning, robustness, hallucination, interpretabilityCVPR 2026
-
A Massive-Scale Multilingual Multi-Cultural Multimodal RAG†
RAG, multilingual, multimodal, evaluation, test-time scalingCVPR 2026
-
Counting to Four is Still a Chore for VLMs*
VLM, numerical reasoning, interpretability, RLCVPRW 2026
-
Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations*
LLM, multilingual, representation learning, unsupervised learningEMNLPW 2025
-
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural VQA on Global Cuisines*
VQA, multilingual, multicultural, benchmark, evaluation, data curationNAACL 2025 (Best Theme Paper)
-
Seeing Culture: A Benchmark for Visual Reasoning and Grounding†
visual grounding, evaluation, object detection, multiculturalEMNLP 2025
-
Predicting Language Model Performance on Multilingual Tasks via Proxy Models
multilingual, evaluation, efficient, regression, SFTNAACL 2025
-
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models*
VQA, VLM, NLP, synthetic data generation, visual reasoningCOLING 2025
-
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
multilingual, multicultural, video understanding, speech understanding, evaluationEMNLP 2024
Work Experience
- Discovered up to a 30% performance degradation in frontier VLMs on a multicultural visual grounding & VQA benchmark, resulting in an EMNLP benchmark publication and an analysis-focused preprint.
- Reduced data validation timeline by days for 50K+ spatial annotations with an automated VLM-based evaluation pipeline that automatically flagged poor polygon segmentation via foreground-color ratios.
- Minimized model guessing shortcuts and ensured benchmark robustness by engineering a heuristic-based question-generation engine to synthesize hard adversarial distractors.
- Automated manual enterprise operations by integrating real-time scheduling, invoicing, and geolocation systems serving 3+ scaleup companies.
- Cut system latency by ~20% while maximizing uptime and scalability by architecting a microservices migration from a legacy monolithic MVC infrastructure.
- Scaled command-line autograder software to support 1k+ concurrent users with 99% uptime.
- Contributed to an 8% increase in absolute approved test-case volume within 6 months at major corporation scale.
- Optimized a topic extraction pipeline, resulting in a 2x efficiency gain and 20% classification improvement.
- Led the development of real-time AR and RAG multimodal systems for a high-stakes demo of 20k+ audiences.