Sicong Jiang

Shot in Tamarindo, Costa Rica

Sicong Jiang

Building RSI Agents @ Google DeepMind | PhD @ McGill University

I am an Incoming Research Scientist at Google DeepMind, where I pioneer Agentic Recursive Self-Improvement (RSI). Previously, as a Founding Scientist at Abaka AI, I directed evaluation research and architected large-scale agentic data & RL environment systems, delivering mission-critical datasets and infrastructure to several frontier AI labs.

My research focuses on the fundamental challenge of building self-evolving AI agents through the lens of automated evaluation, reward modeling, and verifier harnesses. I develop high-impact benchmarks and alignment suites across the agent stack: from sandboxed evaluation harnesses (Harbor-Index, VeriWeb) and human-aligned reward/world models (EditReward, WorldReasonBench) to tool-augmented multimodal reasoning (AgentThink, ChartNet). My mission is to engineer reliable intelligence capable of open-ended, autonomous self-evolution.

Aug 2026
πŸš€ Excited to join Google DeepMind as a Research Scientist in Oct 2026 to keep working on Gemini RSI.
Aug 2026
πŸŽ‰ One paper accepted by WACV 2027. Check EvaDrive.
Jul 2026
πŸŽ‰ Released Harbor-Index β€” one of the most challenging agentic benchmarks to date. Great teamwork!
Feb 2026
πŸŽ‰ Two papers accepted by CVPR 2026 (one Main + one Findings). Check ChartNet, EgoTL.
Jan 2026
πŸŽ‰ One paper accepted by ICLR 2026. Check EditReward.
Nov 2025
πŸŽ‰ One paper accepted (oral) by Bridge Program of AAAI 2026.
Aug 2025
πŸŽ‰ One paper accepted by EMNLP 2025. Check AgentThink.
Aug 2025
🀝 Joined 2077AI-Foundationβ€”thrilled to contribute to the AI open-source community!
Jul 2025
πŸš€ Joined Abaka AI as a Founding Technical Member in Palo Alto, California.
Jul 2025
πŸŽ‰ One paper accepted by ICCV 2025 Foundation Models for AD Workshop. Check VLA4AD Survey.
Feb 2025
πŸŽ‰ One paper accepted by ICLR 2025 Trustworthy LLM Workshop. Check SparseAttack-LLM4TS.
Jan 2025
πŸŽ‰ One paper accepted by AISTATS 2025. Check Attack-LLM4TS.

* indicates equal contribution. For the complete list, visit Google Scholar.

AI Agents, Benchmarks & Evaluation

Foundation Models: Robustness, Safety & Applications

Research Intern
Mar 2026 – Present Β· London, UK / Toronto, Canada

Self-Evolving Agents: Engineered an execution-grounded RL architecture for autonomous self-improvement, enabling Gemini Flash-tier models to outperform Pro-tier baselines on competitive coding benchmarks.

Test-Time Compute Distillation: Designed an execution-consistency verification pipeline, internalizing test-time scaling compute into permanent policy weight updates through continuous RL.

Harness-Data Co-Evolution: Architected an automated harness pipeline to actively harvest high-hardness RL data and edge cases, driving continuous model gains via the co-evolution of evaluation harnesses and trajectory data.

RL Diagnostics & Reward Hacking: Developed an automated pipeline to identify and mitigate reward hacking loops; enhanced training stability and policy performance by filtering anomalous agentic rollouts in the RL loop.

Founding Scientist
Aug 2025 – Feb 2026 Β· Palo Alto, CA, United States

Research: As a founding member of the Research team, I lead benchmarking and evaluation for agentic and multimodal LLMs. I led the EditReward (ICLR'26) project and co-developed large-scale benchmarks including SuperGPQA (NeurIPS'25), ChartNet (CVPR'26), EgoTL (CVPR'26) and VeriWeb.

Advanced Dataset & Pipeline Design: Led several zero-to-one pipeline buildsβ€”architecting and deploying high-difficulty dataset solutions and production pipelines from scratch across coding, IMO-level math, multimodal data, agentic trajectories, and RL environments. These datasets and pipelines are directly used for model training and evaluation for several frontier AI labs.

Core Contributor
Aug 2025 – Mar 2026

As a core contributor, conducting substantial research across benchmarks, datasets, and agent evaluation for the open-source community.

Agent Evaluation: Led research on agent evaluation and training datasets, focusing on long-horizon reasoning, tool use, and self-evolving agent capabilities.

Multimodal Image Datasets: Led multimodal dataset research for image generation, including preference data and evaluation frameworks for alignment and controllability.

Applied Scientist Intern
May 2025 – Aug 2025

Multimodal Data Pipelines: Built data pipelines and multi-stage QA systems for multimodal LLM projects, overseeing large-scale annotation workflows and label consistency.

Dataset Quality & Validation: Conducted analysis and validation to refine annotations and ensure robust datasets for LLM post-training.

Research Assistant
Jan 2022 – May 2025 Β· Montreal, QC, Canada

AgentThink (Agent Reasoning): Led a collaboration with Xiaomi and Tsinghua on tool-augmented reasoning for vision-language models in autonomous driving, achieving +54% answer accuracy on open-source models.

Adversarial LLM4TS: Developed a black-box attack framework and public benchmarks for LLM-based time-series forecasting, in collaboration with the Amazon Chronos and Nixtla teams.

Research Assistant
Aug 2019 – Dec 2020 Β· Atlanta, GA, United States

Multi-Agent RL Exploration: Developed a multi-agent search strategy combining MADDPG with frontier-based exploration, and built evaluation benchmarks for exploration efficiency.

Awards

2024 McGill Engineering Doctoral Award (MEDA)
2021 TISED Doctoral Recruitment Award (DRA), McGill University
2019 Outstanding Graduate of Liaoning Province; Most Influential Graduate, Northeastern University
2017 National 1st Prize, China Undergraduate Mathematical Contest in Modeling
2017 1st Class Academic Scholarship, Northeastern University

Academic Service

Workshops Organizer

Conferences Reviewer

  • Advances in Neural Information Processing Systems (NeurIPS)
  • International Conference on Learning Representations (ICLR)
  • International Conference on Artificial Intelligence and Statistics (AISTATS)
  • IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
  • International Conference on Computer Vision (ICCV)
  • Conference on Language Modeling (COLM)
  • Conference on Empirical Methods in Natural Language Processing (EMNLP)
  • Association for the Advancement of Artificial Intelligence (AAAI)
  • IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
  • IEEE International Conference on Robotics and Automation (ICRA)
  • IEEE Intelligent Transportation Systems Conference (ITSC)

Journals Reviewer

  • IEEE Robotics and Automation Letters (RA-L)
  • Transportation Research Part C: Emerging Technologies (TRC)
  • IEEE Transactions on Intelligent Transportation Systems (T-ITS)

I enjoy music by Tyler, the Creator, SZA and Chappell Roan.

Sometimes I also listen to Taylor Swift, Olivia Rodrigo and 9m88.

My favorite influencer is Allywoo on RedNote.

Cat: Bobo, a golden shaded British Shorthair who is good at programming with buttons.

Bobo