๐ News
Introduction
SeePhys Pro is a fine-grained modality-transfer benchmark for multimodal physics reasoning. Each problem preserves the same physical semantics while progressively moving task-critical information from text into diagrams, revealing whether a model reasons over stable physics or over the surface form of the prompt.
Benchmark Design
The benchmark is built on the principle of same physics, different representation. Four aligned levels decompose the cost of structural transfer, visual variable grounding, and full image rendering.
Data Construction
Seed problems are curated from public datasets, textbooks, exam papers, olympiad archives, and physics problem books, then manually transformed into aligned multimodal variants with preserved answers and solution paths.
- Source-matched: benchmark and training corpora share a broad physics source pool but remain instance-disjoint.
- Manually aligned: annotators redraw structure and variable layers while preserving the physical system.
- Fine-grained metadata: discipline, field, domain, visual evidence type, and reasoning skill annotations support targeted analysis.
ICML 2026 ยท 3rd AI for Math Workshop ยท Challenge 3
Workshop Challenge Leaderboard
Final private evaluation results released on June 16, 2026. Rankings are determined by question-count weighted accuracy across 3,320 private test questions.
| Rank | Participant | Overall | Level 1 | Level 2 | Level 3 | Level 4 | Level 5 |
|---|---|---|---|---|---|---|---|
| ๐ฅ 1 | kaistaailab | 74.31 | 74.50 | 76.50 | 73.00 | 71.25 | 87.50 |
| ๐ฅ 2 | jasperdekoninck | 72.23 | 73.00 | 75.00 | 69.87 | 68.50 | 89.17 |
| ๐ฅ 3 | baibzihe1 | 71.69 | 73.88 | 74.25 | 69.37 | 66.87 | 87.50 |
| 4 | lh12345 | 71.51 | 73.25 | 74.75 | 68.25 | 67.50 | 86.67 |
| 5 | ctree4113 | 70.36 | 71.63 | 71.88 | 67.87 | 67.00 | 90.83 |
| 6 | huyttuan | 62.17 | 66.37 | 67.63 | 56.63 | 54.50 | 85.83 |
| 7 | bkhoi | 56.87 | 58.00 | 56.87 | 57.13 | 56.75 | 48.33 |
| 7 | mihena | 56.87 | 57.75 | 57.50 | 56.50 | 57.00 | 48.33 |
| 9 | kurone02 | 55.66 | 58.37 | 59.13 | 50.50 | 49.63 | 89.17 |
| 10 | cmxu7 | 42.44 | 50.12 | 44.37 | 36.88 | 33.50 | 75.00 |
| 11 | robertboy18 | 39.34 | 45.87 | 41.38 | 31.62 | 31.37 | 86.67 |
| 12 | usercrab | 31.42 | 39.87 | 33.25 | 25.50 | 23.13 | 57.50 |
| 13 | ykjung | 29.73 | 29.50 | 30.00 | 26.25 | 29.38 | 55.00 |
| 14 | deleted_user_70413 | 21.90 | 20.13 | 20.00 | 20.62 | 18.25 | 79.17 |
| 15 | gptrans5_5 | 21.87 | 20.13 | 20.00 | 20.50 | 18.25 | 79.17 |
Top 15 final private leaderboard entries are shown. Scores are percentages, and tied Overall scores share the same rank. Level 5 is an additional OOD test for the workshop challenge and is not included in the paper. Source: official CodaBench final private leaderboard.
Benchmark Leaderboard
Across evaluated models, average accuracy drops from 49.2% at Level 1 to 35.8% at Level 4. The largest average gap occurs at Level 2 to Level 3, where models must ground variables and labels from images.
| Model | Accuracy / consistency | Transfer gap | Avg(L1-L4) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| L1 | L2 | L3 | L4 | Cons4 | ΔS | ΔV | ΔR | ΔT | ||
| Human Performance | 54.0 | 58.5 | 59.5 | 56.0 | 49.0 | -4.5 | -1.0 | 3.5 | -2.0 | 57.00 |
| Gemini-3.1-Pro | 71.0 | 72.0 | 66.5 | 66.5 | 47.0 | -1.0 | 5.5 | 0.0 | 4.5 | 69.00๐ฅ |
| Claude-4.7-Opus | 74.0 | 67.0 | 56.5 | 46.5 | 33.5 | 7.0 | 10.5 | 10.0 | 27.5 | 61.00๐ฅ |
| GPT-5.4 | 67.4 | 64.1 | 55.8 | 53.0 | 32.6 | 3.3 | 8.3 | 2.8 | 14.4 | 60.08๐ฅ |
| Qwen-3.6-flash | 61.4 | 59.3 | 49.9 | 48.4 | 29.9 | 2.1 | 9.4 | 1.5 | 13.0 | 54.75 |
| Qwen3.5-27B | 45.0 | 34.8 | 28.0 | 25.6 | 9.9 | 10.3 | 6.8 | 2.4 | 19.4 | 33.35 |
| GPT-5 | 41.8 | 32.9 | 23.8 | 23.2 | 8.9 | 8.9 | 9.1 | 0.5 | 18.5 | 30.43 |
| Gemma-4-31B-it | 38.9 | 33.5 | 23.9 | 22.0 | 8.9 | 5.4 | 9.6 | 1.9 | 16.9 | 29.58 |
| Average | 49.2 | 46.1 | 38.7 | 35.8 | 21.4 | 3.0 | 7.4 | 2.9 | 13.4 | 42.45 |
Accuracy and consistency are percentages. Positive transfer gaps indicate degradation under a more visual representation; ฮS, ฮV, and ฮR measure the gaps in structural transfer (Level 1โ2), variable grounding (Level 2โ3), and rendering (Level 3โ4), respectively. ฮT measures the overall gap between Level 1 and Level 4, while Cons4 measures consistency across all four levels.
Training-Time Diagnostic
SeePhys Pro also tests whether multimodal RLVR improvements are visually grounded. A blind-training control masks all training images, yet still improves unmasked validation accuracy, showing that final-answer gains alone can overstate visual grounding.
Cross-Benchmark Controls
Blind gains are not unique to SeePhys Pro. Across external physics and math benchmarks, masked-image RL can recover a substantial fraction of normal RL gains, indicating sensitivity to residual textual and distributional cues.
Citation
@article{xiang2026seephyspro,
title={SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning},
author={Xiang, Kun and Zhang, Terry Jingchen and Liu, Zirong and Zhou, Bokai and Tang, Yueling and Yu, Junjie and Lu, Jiacong and Huang, Shangrui and Li, Heng and Zhang, Likui and Liu, Kunkun and Zhang, Changzheng and Fang, Yangle and Guo, Boqiang and Zhen, Hui-Ling and Tu, Dandan and Huang, Yinya and Liang, Xiaodan},
journal={arXiv preprint arXiv:2605.09266},
year={2026},
url={https://arxiv.org/abs/2605.09266}
}
@article{xiang2026seephys,
title={Seephys: Does seeing help thinking?--benchmarking vision-based physics reasoning},
author={Xiang, Kun and Li, Heng and Zhang, Terry Jingchen and Huang, Yinya and Liu, Zirong and Qu, Peixin and He, Jixi and Chen, Jiaqi and Yuan, Yu-Jie and Han, Jianhua and others},
journal={Advances in Neural Information Processing Systems},
volume={38},
year={2026}
}