-
Video models are zero-shot learners and reasoners.
Arxiv 25.09| [PDF] | [Project Page] -
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-COF Benchmark.
Arxiv 25.10| [PDF] [Project Page] | [Github] -
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models.
Arxiv 25.11| [PDF] [Project Page] | [Github] -
VMEvalKit: Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven’ Matrices.
25.10[Github] -
Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark.
Arxiv 25.11| [PDF] | [Github] -
Reasoning via Video: The First Evaluation of Video Models’ Reasoning Abilities through Maze-Solving Tasks.
Arxiv 25.11| [PDF] | [Page] | [Github] -
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm.
Arxiv 25.11| [PDF] [Project Page] | [Github]
-
MiniVeo3-Reasoner: Thinking with Videos from Open-Source Priors.
25.10[Github] -
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models.
Arxiv 26.01| [PDF] | [Project Page] | [Github]