Weaving Visual World Modeling into Multimodal LLMs
Can visual transition reasoning serve as a shared training primitive for multimodal LLMs?
Training subsets of only about 2,000 WOVEN items each collectively improve 22 of 26 external benchmarks, by up to 27.3 percentage points. The gains come from one shared capability, and controlled comparisons yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks.
Visual transition reasoning: inferring the missing part of (s, a, s′)
Given a transition (s, a, s′), where s and s′ are the visual states of a scene before and after an action a, which may be agent-driven or passive, visual transition reasoning is the process of inferring unobserved parts of the transition from its known or observed parts. Predicting outcomes, inferring actions from observed changes, and reasoning about alternative actions are different inferences over the same triplet. Each choice of what is given and what is asked defines a reasoning operation; WOVEN covers eight, in four families.
Notation follows the paper: aalt is an alternative action from the same initial state, aexo a passive event, aq a passive event with a cue naming the physical principle, s states in time order and s̃ the same states shuffled.
Controlled transitions from a video generation model
Studying this capability requires controlled comparisons of supervision: varying one axis such as scene, action, or reasoning operation while holding the others fixed. Real video offers neither controlled actions nor alternative outcomes from the same initial state, and simulators offer control but limited realism and scene diversity. WOVEN therefore generates rollouts with a video generation model (Wan2.2-I2V-A14B) from specified initial frames and action descriptions, following two principles: the resulting and intermediate states come from the rollout rather than from the conditioning prompt, and several rollouts under different actions start from the same initial state, so that alternative outcomes are available.
Every item is a four-option multiple-choice question labeled along three axes: 20 scene types, 5 action types (passive physical events, camera motion, object inspection, navigation, and object manipulation), and 8 reasoning types. The three wrong options are usually outcomes of other rollouts from the same initial state, and each is labeled with its error type, such as “reversed direction” or “violating gravity”.
Current MLLMs show a substantial and systematic deficit
To establish the need for training, we evaluate 38 frontier MLLMs on the in-distribution and state-perturbation test sets. The best model, GPT-5.4, reaches 65.8% in-distribution accuracy against 92.3% for humans (the mean of three annotators), and the next three fall between 57.8% and 63.2%. The gap persists across model scales, and it has a consistent structure across model families: for every model, temporal adjacency is harder than temporal ordering, counterfactual removal is harder than counterfactual substitution, and accuracy drops more under geometric than under appearance changes.
| Action type | Reasoning type | Perturbation | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Accuracy (%) | Overall | Exo | Perc | Insp | Navi | Mani | Fwd | Inv | Rmv | Sub | Otm | Cue | Ord | Adj | Geo | App |
| Human | 92.3 | 92.8 | 92.8 | 92.6 | 90.6 | 92.8 | 92.8 | 92.8 | 91.1 | 92.2 | 93.3 | 91.1 | 92.8 | 91.7 | 97.3 | 98.7 |
| GPT-5.4 | 65.8 | 62.4 | 61.2 | 58.3 | 68.8 | 81.2 | 83.8 | 81.7 | 64.5 | 76.0 | 81.9 | 85.8 | 57.1 | 27.0 | 62.7 | 95.5 |
| Intern-S1-Pro | 63.2 | 56.2 | 52.8 | 53.7 | 70.1 | 87.1 | 77.6 | 77.5 | 66.7 | 73.3 | 61.5 | 69.1 | 63.4 | 26.5 | 33.3 | 79.2 |
| Qwen2.5-VL-72B | 58.9 | 53.7 | 47.7 | 47.8 | 66.4 | 83.4 | 74.7 | 74.7 | 54.8 | 72.0 | 59.4 | 76.0 | 49.4 | 28.0 | 38.5 | 83.0 |
| Gemma-4-31B | 57.8 | 50.8 | 49.4 | 53.1 | 61.2 | 74.3 | 76.3 | 76.6 | 63.6 | 64.5 | 64.9 | 70.5 | 30.8 | 25.3 | 59.0 | 95.5 |
In-distribution accuracy overall, by action type (Exo: passive physical events; Perc: camera motion; Insp: object inspection; Navi: navigation; Mani: object manipulation) and by reasoning type (Fwd/Inv: forward/inverse dynamics; Rmv/Sub: counterfactual removal/substitution; Otm/Cue: outcome/cued prediction; Ord/Adj: temporal ordering/adjacency), and state-perturbation accuracy under geometric and appearance changes. Results for all 38 models are in the paper.
Visual transition reasoning is learnable and transfers broadly
Supervised fine-tuning of Qwen2.5-VL-3B-Instruct on the full WOVEN training set raises WOVEN accuracy as follows; the gains also hold at larger model scales and in a second model family.
Each controlled subset of the training set (one action type × one reasoning family, about 2,000 items) also raises in-distribution accuracy on its own, by 6.1 to 25.3 percentage points, including on action types it does not train on. The learning is data-efficient: with about 7 hours of generated video, WOVEN post-training yields a larger gain on a spatial-reasoning benchmark than mid-training the same backbone on 125,000 hours of video (Orca).
Transfer to 22 of 26 external benchmarks
Evaluated zero-shot on 26 external benchmarks, the full training set improves 12 of them, by up to 24.0 percentage points on SAT. The controlled subsets transfer as well: with the best subset chosen for each benchmark, the coverage rises to 22 of 26, that is, every benchmark that involves visual transitions, with gains of up to 27.3 percentage points. The 4 benchmarks of static perception and general video QA do not improve under any training, so the gains do not come from an overall stronger model or from better multiple-choice answering.
Visual transition reasoning serves as a shared training primitive
The gains so far could still be a sum of separate effects, each subset helping the tasks closest to it. For visual transition reasoning to serve as a shared training primitive, two requirements must hold.
(a) Different sources improve the same capability
Each subset raises WOVEN accuracy on action types it does not train on, in 43 of the 44 such cells below. The capability also spans reasoning operations: training on forward prediction alone raises counterfactual substitution and removal by 48 and 25 percentage points.
| Training subset | Passive physical events | Camera motion | Object inspection | Navigation | Object manipulation | Temporal adjacency |
|---|---|---|---|---|---|---|
| Camera motion · causal dynamics | +5.1 | +33.0 | +13.4 | +26.9 | +39.8 | +0.2 |
| Object inspection · causal dynamics | +6.8 | +8.9 | +39.9 | +27.2 | +37.7 | +0.9 |
| Navigation · causal dynamics | +7.6 | +10.1 | +12.7 | +30.0 | +50.9 | +4.9 |
| Object manipulation · causal dynamics | +8.2 | +3.9 | +5.7 | +26.2 | +61.4 | +2.3 |
| Object inspection · counterfactual | +6.6 | +7.3 | +38.5 | +24.8 | +42.7 | +1.3 |
| Object manipulation · counterfactual | +10.6 | +5.2 | +5.5 | +27.2 | +64.2 | +2.9 |
| Passive physical events · physical modeling | +27.4 | −1.2 | +5.7 | +6.7 | +9.6 | −1.5 |
| Camera motion · temporal | +8.4 | +10.9 | +4.9 | +8.0 | +1.3 | +29.9 |
| Object inspection · temporal | +9.2 | +6.4 | +3.4 | +11.2 | +1.2 | +33.1 |
| Navigation · temporal | +8.1 | +12.2 | +4.0 | +8.6 | +2.9 | +32.0 |
| Passive physical events · temporal | +8.4 | +7.1 | +5.5 | +10.1 | +2.3 | +30.6 |
Gain in WOVEN in-distribution accuracy (percentage points over the base model) after training Qwen2.5-VL-3B-Instruct on each controlled subset, by action type of the test items and on temporal adjacency. Outlined cells mark the action type the subset trains on; every other column is an action type the subset does not train on.
(b) The downstream gains result from that capability
If so, all subsets should improve largely the same benchmarks, and subsets that improve the capability more should transfer more. The transfer profiles of the 11 subsets are all positively correlated, and the more a subset improves WOVEN accuracy, the larger its average gain on the external benchmarks (r = 0.86 across the 11 subsets). The same WOVEN format with permuted or mirrored transitions does not reproduce the downstream gains, and the gains are specific to benchmarks that involve visual transitions.
We further test this by intervening on the training data. With the training set fixed at 2,000 items, we replace a growing share of task-specific items with a WOVEN subset on three benchmarks from three domains (SAT, ActionEQA, and CLEVRER), averaging over three seeds. Replacing 30–50% of the task-specific items leaves accuracy comparable to training on those items alone, and above a control that simply deletes the replaced items.
A training recipe for visual world modeling
From the transfer profiles, error-typed analyses, and perturbation robustness of the controlled subsets, we distill four recipe items, then test the recipe prospectively on held-out benchmarks.
(1) The reasoning operation decides where supervision transfers
Similarity in scenes, actions, or application domains is an intuitive basis for selecting transferable supervision. To test it, we compare subsets that share only an action type with subsets that share only a reasoning family. All transfer profiles contain the common component from Question 3, so we remove it by standardizing the gains within each benchmark and correlate the residual profiles pairwise. Subsets teaching the same reasoning family have similar residual profiles (mean pairwise correlation +0.32), whereas subsets sharing only an action type are no more similar than subsets sharing neither (−0.22 versus −0.23). The same ordering holds under three other normalizations of the gains.
(2) Larger state changes give more robustness
Among the causal-dynamics subsets, gains on the state-perturbation test set increase with how much the action changes the scene, and passive physical events give the largest gain of all subsets.
Gain on the state-perturbation test set (percentage points over the base model) after training on each subset; the first four are causal-dynamics subsets, the last is the passive-physical-event subset.
(3, 4) Where transfer breaks
Supervision on one reasoning operation often teaches others, but two kinds of transition reasoning are not learned this way, nor are capabilities beyond state transitions. Temporal localization is learned only from temporal supervision: the four temporal subsets raise temporal-adjacency accuracy from 25.7% to 55.6–58.7%, whereas the seven other subsets leave it near the baseline. Agent-driven transitions are not learned from passive physical events: that subset contains no agent action to learn from, and it lowers accuracy on the two benchmarks built on an agent’s action and its outcome, WorldPrediction (−10.2) and ActionEQA (−1.7). The 4 benchmarks of static perception and general video QA do not improve under any training.
The recipe predicts transfer on held-out benchmarks
We use the recipe to select training supervision before any training, and evaluate on external benchmarks that played no part in its derivation (questions from MMSI-Bench, VLM4D, PhysBench, SPAR-Bench, MVBench, IntPhys 2, and CameraBench). For questions about agent actions and their outcomes, a mixture of WOVEN items composed by the recipe is compared with mixtures of the same size composed by matching the action (camera-heavy), by a different reasoning operation (temporal-heavy), by passive physical events (exogenous-heavy), or with no selection (uniform). Each mixture trains Qwen2.5-VL-3B-Instruct with 8 seeds. All five comparisons come out as the recipe predicts.
Data, code, and models
Citation
If you find WOVEN useful in your research, please consider citing:
@article{fan2026woven,
title={WOVEN: Weaving Visual World Modeling into Multimodal LLMs},
author={Fan, Zheyu and Zhang, Yue and Deng, Mingkai and Wang, Kangrui and Wang, Qineng and Chen, Canyu and Hao, Jie and Fan, Xing and Guo, Chenlei and Xing, Eric P. and Bansal, Mohit and Li, Manling},
journal={arXiv preprint arXiv:2610.12417},
year={2026}
}