|
Shaofeng Yin (殷绍峰)
I'm currently working in stealth mode. I graduated from Peking University with a bachelor's degree in artificial intelligence, where I was a member of the Zhi Class. In spring 2025, I was a visiting student at Berkeley AI Research.
My earlier research focused on human-object interaction (HOI) detection. More recently, advances in vision-language models (VLMs) have led me to explore multimodal agents.
Outside of research, I enjoy art, literature, and philosophy. I'm also committed to volunteer teaching. I have a background in competitive programming and became a Codeforces Master at age 15.
Email /
Scholar /
Linkedin /
Github
|
|
|
|
[schema]: Frontier Models with the Right Harness Achieve ~99% on ARC-AGI-3 Public
Guanning Zeng*, Jiani Wang, Wenjie Ma, Shaofeng Yin, Chenyang Wang, Shichen Liu, Angjoo Kanazawa, Wode Ni, Xiuyu Li*, Andrea Zanette*, Haiwen Feng*
(*Project leads)
Project release, 2026
project page /
traces /
tweet
Schema equips frontier models with an editable symbolic world model for discovering object dynamics, testing hypotheses, and planning in ARC-AGI-3. With the right agent harness, it reaches approximately 99% RHAE on the public set.
|
|
|
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
Shaofeng Yin, Jiaxin Ge, Zora Zhiruo Wang, Xiuyu Li, Michael J. Black, Trevor Darrell, Angjoo Kanazawa, Haiwen Feng
ECCV, 2026
project page /
arXiv /
tweet
Vision-as-inverse-graphics, the concept of reconstructing an image as an editable graphics program is a long-standing goal of computer vision. Yet even strong VLMs aren't able to achieve this in one-shot as they lack fine-grained spatial and physical grounding capability. Our key insight is that closing this gap requires interleaved multimodal reasoning through iterative execution and verification. Stemming from this, we present VIGA (Vision-as-Inverse-Graphic Agent) that starts from an empty world and reconstructs or edits scenes through a closed-loop write-run-render-compare-revise procedure.
|
|
|
ToolVQA: A Dataset for Multi-step Reasoning VQA with External Tools
Shaofeng Yin,
Ting Lei,
Yang Liu
ICCV, 2025
project page /
arXiv
Recent benchmarks reveal significant gaps in real-world tool-use proficiency, particularly in functionally diverse multimodal settings requiring multi-step reasoning. To bridge this gap, we propose ToolEngine, a novel data generation pipeline that employs Depth-First Search (DFS) with a dynamic in-context example matching mechanism to simulate human-like tool-use reasoning.
|
|