Google DeepMind as a Research Scientist in Oct 2026 to keep working on
Gemini RSI.
Shot in Tamarindo, Costa Rica
Building
RSI Agents @ Google DeepMind | PhD @ McGill University
I am an Incoming Research Scientist at
Google DeepMind, where I pioneer Agentic Recursive Self-Improvement (RSI). Previously, as a Founding Scientist at
Abaka AI, I directed evaluation research and architected large-scale agentic data & RL environment systems, delivering mission-critical datasets and infrastructure to several frontier AI labs.
My research focuses on the fundamental challenge of building self-evolving AI agents through the lens of automated evaluation, reward modeling, and verifier harnesses. I develop high-impact benchmarks and alignment suites across the agent stack: from sandboxed evaluation harnesses (Harbor-Index, VeriWeb) and human-aligned reward/world models (EditReward, WorldReasonBench) to tool-augmented multimodal reasoning (AgentThink, ChartNet). My mission is to engineer reliable intelligence capable of open-ended, autonomous self-evolution.
Google DeepMind as a Research Scientist in Oct 2026 to keep working on
Gemini RSI.
Google DeepMind (London) as a Research Intern to work on Self-Evolving Agents.
2077AI-Foundationβthrilled to contribute to the AI open-source community!
Abaka AI as a Founding Technical Member in Palo Alto, California.* indicates equal contribution. For the complete list, visit Google Scholar.
Self-Evolving Agents: Engineered an execution-grounded RL architecture for autonomous self-improvement, enabling Gemini Flash-tier models to outperform Pro-tier baselines on competitive coding benchmarks.
Test-Time Compute Distillation: Designed an execution-consistency verification pipeline, internalizing test-time scaling compute into permanent policy weight updates through continuous RL.
Harness-Data Co-Evolution: Architected an automated harness pipeline to actively harvest high-hardness RL data and edge cases, driving continuous model gains via the co-evolution of evaluation harnesses and trajectory data.
RL Diagnostics & Reward Hacking: Developed an automated pipeline to identify and mitigate reward hacking loops; enhanced training stability and policy performance by filtering anomalous agentic rollouts in the RL loop.
Research: As a founding member of the Research team, I lead benchmarking and evaluation for agentic and multimodal LLMs. I led the EditReward (ICLR'26) project and co-developed large-scale benchmarks including SuperGPQA (NeurIPS'25), ChartNet (CVPR'26), EgoTL (CVPR'26) and VeriWeb.
Advanced Dataset & Pipeline Design: Led several zero-to-one pipeline buildsβarchitecting and deploying high-difficulty dataset solutions and production pipelines from scratch across coding, IMO-level math, multimodal data, agentic trajectories, and RL environments. These datasets and pipelines are directly used for model training and evaluation for several frontier AI labs.
As a core contributor, conducting substantial research across benchmarks, datasets, and agent evaluation for the open-source community.
Agent Evaluation: Led research on agent evaluation and training datasets, focusing on long-horizon reasoning, tool use, and self-evolving agent capabilities.
Multimodal Image Datasets: Led multimodal dataset research for image generation, including preference data and evaluation frameworks for alignment and controllability.
Multimodal Data Pipelines: Built data pipelines and multi-stage QA systems for multimodal LLM projects, overseeing large-scale annotation workflows and label consistency.
Dataset Quality & Validation: Conducted analysis and validation to refine annotations and ensure robust datasets for LLM post-training.
AgentThink (Agent Reasoning): Led a collaboration with Xiaomi and Tsinghua on tool-augmented reasoning for vision-language models in autonomous driving, achieving +54% answer accuracy on open-source models.
Adversarial LLM4TS: Developed a black-box attack framework and public benchmarks for LLM-based time-series forecasting, in collaboration with the Amazon Chronos and Nixtla teams.
Multi-Agent RL Exploration: Developed a multi-agent search strategy combining MADDPG with frontier-based exploration, and built evaluation benchmarks for exploration efficiency.
I enjoy music by Tyler, the Creator, SZA and Chappell Roan.
Sometimes I also listen to Taylor Swift, Olivia Rodrigo and 9m88.
My favorite influencer is Allywoo on RedNote.
Cat: Bobo, a golden shaded British Shorthair who is good at programming with buttons.