Shaofeng Yin (殷绍峰)

I'm currently working in stealth mode. I graduated from Peking University with a bachelor's degree in artificial intelligence, where I was a member of the Zhi Class. In spring 2025, I was a visiting student at Berkeley AI Research.

My earlier research focused on human-object interaction (HOI) detection. More recently, advances in vision-language models (VLMs) have led me to explore multimodal agents.

Outside of research, I enjoy art, literature, and philosophy. I'm also committed to volunteer teaching. I have a background in competitive programming and became a Codeforces Master at age 15.

Email  /  Scholar  /  Linkedin  /  Github

profile photo

📚 Selected Publications

[schema]: Frontier Models with the Right Harness Achieve ~99% on ARC-AGI-3 Public
Guanning Zeng*, Jiani Wang, Wenjie Ma, Shaofeng Yin, Chenyang Wang, Shichen Liu, Angjoo Kanazawa, Wode Ni, Xiuyu Li*, Andrea Zanette*, Haiwen Feng*
(*Project leads)
Project release, 2026
project page / traces / tweet

Schema equips frontier models with an editable symbolic world model for discovering object dynamics, testing hypotheses, and planning in ARC-AGI-3. With the right agent harness, it reaches approximately 99% RHAE on the public set.

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
Shaofeng Yin, Jiaxin Ge, Zora Zhiruo Wang, Xiuyu Li, Michael J. Black, Trevor Darrell, Angjoo Kanazawa, Haiwen Feng
ECCV, 2026
project page / arXiv / tweet

Vision-as-inverse-graphics, the concept of reconstructing an image as an editable graphics program is a long-standing goal of computer vision. Yet even strong VLMs aren't able to achieve this in one-shot as they lack fine-grained spatial and physical grounding capability. Our key insight is that closing this gap requires interleaved multimodal reasoning through iterative execution and verification. Stemming from this, we present VIGA (Vision-as-Inverse-Graphic Agent) that starts from an empty world and reconstructs or edits scenes through a closed-loop write-run-render-compare-revise procedure.

ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
Shaofeng Yin, Ting Lei, Yang Liu
ICCV, 2025
project page / arXiv

Recent benchmarks reveal significant gaps in real-world tool-use proficiency, particularly in functionally diverse multimodal settings requiring multi-step reasoning. To bridge this gap, we propose ToolEngine, a novel data generation pipeline that employs Depth-First Search (DFS) with a dynamic in-context example matching mechanism to simulate human-like tool-use reasoning.

🏆 Selected Awards

2025: First Prize in CVPR International CulturalVQA Benchmark Challenge
2025: SenseTime Scholarship (30/year in China)
2024: National Scholarship (Highest honor for undergraduates)

📸 Selected Photography

Photography 21
Photography 20
Photography 19
Photography 18
Photography 17
Photography 16
Photography 14
Photography 13
Photography 12
Photography 11

The template is stole from Jon Barron.