WOVEN

WOVEN

Weaving Visual World Modeling into Multimodal LLMs

Can visual transition reasoning serve as a shared training primitive for multimodal LLMs?

Zheyu Fan1,Yue Zhang3,Mingkai Deng2,Kangrui Wang1,Qineng Wang1,Canyu Chen1,

Jie Hao4,Xing Fan4,Chenlei Guo4,Eric P. Xing2,Mohit Bansal3,Manling Li1

1Northwestern University2Carnegie Mellon University3UNC Chapel Hill4Amazon

Key Takeaway

Training subsets of only about 2,000 WOVEN items each collectively improve 22 of 26 external benchmarks, by up to 27.3 percentage points. The gains come from one shared capability, and controlled comparisons yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks.

Overview figure. Panel 1: MLLMs fail on spatial, embodied, physical and temporal tasks. Panel 2: these failures share reasoning about how visual states, actions and outcomes relate. Panel 3: WOVEN generates (s, a, s′) transitions with a video generation model and turns them into error-typed multiple-choice supervision. Panel 4: training on WOVEN improves 22 of 26 external benchmarks and yields a training recipe.
Overview. (1) MLLMs fail across spatial, embodied, physical, and temporal tasks. (2) These failures share a weakness: reasoning about how visual states, actions, and outcomes relate. (3) WOVEN generates (s, a, s′) transitions with a video generation model, organizes them along controlled axes of reasoning, action, and scene, and turns them into error-typed multiple-choice supervision. (4) Training on WOVEN improves 22 of 26 external benchmarks; the gains come from a shared primitive that different sources build and different tasks use; and controlled comparisons yield a training recipe.

Visual transition reasoning: inferring the missing part of (s, a, s′)

Given a transition (s, a, s′), where s and s′ are the visual states of a scene before and after an action a, which may be agent-driven or passive, visual transition reasoning is the process of inferring unobserved parts of the transition from its known or observed parts. Predicting outcomes, inferring actions from observed changes, and reasoning about alternative actions are different inferences over the same triplet. Each choice of what is given and what is asked defines a reasoning operation; WOVEN covers eight, in four families.

Causal dynamics
Forward dynamics
P(s′ | s, a)
What will the scene look like after the action?
Inverse dynamics
P(a | s, s′)
Which action caused the change?
Counterfactual reasoning
Counterfactual removal
P(s | a, s′)
What would the scene look like if the action had not happened?
Counterfactual substitution
P(s′aalt | a, s′, aalt)
What would the outcome be under a different action?
Physical modeling
Outcome prediction
P(s′ | s, aexo)
How will a passive physical event end?
Cued prediction
P(s′ | s, aq)
How will the event end, given a cue naming the physical principle?
Temporal coherence
Temporal ordering
P(s | s̃, a)
In what order did these states occur?
Temporal adjacency
P(s−, s+ | sref, s̃, a)
Which states come right before and after a reference state?

Notation follows the paper: aalt is an alternative action from the same initial state, aexo a passive event, aq a passive event with a cue naming the physical principle, s states in time order and s̃ the same states shuffled.

Controlled transitions from a video generation model

Studying this capability requires controlled comparisons of supervision: varying one axis such as scene, action, or reasoning operation while holding the others fixed. Real video offers neither controlled actions nor alternative outcomes from the same initial state, and simulators offer control but limited realism and scene diversity. WOVEN therefore generates rollouts with a video generation model (Wan2.2-I2V-A14B) from specified initial frames and action descriptions, following two principles: the resulting and intermediate states come from the rollout rather than from the conditioning prompt, and several rollouts under different actions start from the same initial state, so that alternative outcomes are available.

Every item is a four-option multiple-choice question labeled along three axes: 20 scene types, 5 action types (passive physical events, camera motion, object inspection, navigation, and object manipulation), and 8 reasoning types. The three wrong options are usually outcomes of other rollouts from the same initial state, and each is labeled with its error type, such as “reversed direction” or “violating gravity”.

The WOVEN taxonomy. Left: the passive physical event type with its eight principles and the four agent-driven action types. Middle: the eight reasoning types in four families, each marking which element of (s, a, s′) the question asks for. Right: example scenes; restaurant and beach are held out.
The WOVEN taxonomy. Every item is labeled along three axes. Left: passive physical events with their eight principles, and the four agent-driven action types, where manipulation combines subject (human, humanoid) and view (egocentric, allocentric). Middle: the eight reasoning types in four families; the dashed box marks the element of (s, a, s′) that the question asks for. Right: example scenes; restaurant and beach are held out for the held-out-scene test.
36,076multiple-choice items
22,728Training
2,508Validation
6,228In-distribution test
3,496Held-out-scene test
1,116State-perturbation test

Current MLLMs show a substantial and systematic deficit

To establish the need for training, we evaluate 38 frontier MLLMs on the in-distribution and state-perturbation test sets. The best model, GPT-5.4, reaches 65.8% in-distribution accuracy against 92.3% for humans (the mean of three annotators), and the next three fall between 57.8% and 63.2%. The gap persists across model scales, and it has a consistent structure across model families: for every model, temporal adjacency is harder than temporal ordering, counterfactual removal is harder than counterfactual substitution, and accuracy drops more under geometric than under appearance changes.

Action typeReasoning typePerturbation
Accuracy (%) OverallExoPercInspNaviManiFwdInvRmvSubOtmCueOrdAdjGeoApp
Human 92.392.892.892.690.692.892.892.891.192.293.391.192.891.797.398.7
GPT-5.4 65.862.461.258.368.881.283.881.764.576.081.985.857.127.062.795.5
Intern-S1-Pro 63.256.252.853.770.187.177.677.566.773.361.569.163.426.533.379.2
Qwen2.5-VL-72B 58.953.747.747.866.483.474.774.754.872.059.476.049.428.038.583.0
Gemma-4-31B 57.850.849.453.161.274.376.376.663.664.564.970.530.825.359.095.5

In-distribution accuracy overall, by action type (Exo: passive physical events; Perc: camera motion; Insp: object inspection; Navi: navigation; Mani: object manipulation) and by reasoning type (Fwd/Inv: forward/inverse dynamics; Rmv/Sub: counterfactual removal/substitution; Otm/Cue: outcome/cued prediction; Ord/Adj: temporal ordering/adjacency), and state-perturbation accuracy under geometric and appearance changes. Results for all 38 models are in the paper.

Visual transition reasoning is learnable and transfers broadly

Supervised fine-tuning of Qwen2.5-VL-3B-Instruct on the full WOVEN training set raises WOVEN accuracy as follows; the gains also hold at larger model scales and in a second model family.

In-distribution26.4%89.3%
Held-out scenes27.8%88.2%
State perturbation21.1%38.3%

Each controlled subset of the training set (one action type × one reasoning family, about 2,000 items) also raises in-distribution accuracy on its own, by 6.1 to 25.3 percentage points, including on action types it does not train on. The learning is data-efficient: with about 7 hours of generated video, WOVEN post-training yields a larger gain on a spatial-reasoning benchmark than mid-training the same backbone on 125,000 hours of video (Orca).

Transfer to 22 of 26 external benchmarks

Evaluated zero-shot on 26 external benchmarks, the full training set improves 12 of them, by up to 24.0 percentage points on SAT. The controlled subsets transfer as well: with the best subset chosen for each benchmark, the coverage rises to 22 of 26, that is, every benchmark that involves visual transitions, with gains of up to 27.3 percentage points. The 4 benchmarks of static perception and general video QA do not improve under any training, so the gains do not come from an overall stronger model or from better multiple-choice answering.

Causal dynamics subsetCounterfactual reasoning subsetPhysical modeling subsetTemporal coherence subset Static perception / general video QA
SAT +27.3 Camera motion · causal dynamics
ActionEQA +12.5 Camera motion · causal dynamics
TemporalBench-long +11.7 Object manipulation · counterfactual
InPhyRe +10.6 Object inspection · counterfactual
Robo2VLM +8.9 Object inspection · counterfactual
WorldPrediction +8.1 Object manipulation · causal dynamics
DSI-Bench +6.0 Camera motion · causal dynamics
MindCube +5.9 Camera motion · temporal
CLEVRER +5.1 Object inspection · counterfactual
WM-ABench +4.6 Object manipulation · counterfactual
MVP +4.5 Camera motion · causal dynamics
ERQA +4.5 Object manipulation · causal dynamics
SpatialViz +4.5 Navigation · temporal
ViewSpatial +4.3 Object manipulation · counterfactual
PaiBench +3.5 Object inspection · causal dynamics
CV-Bench +3.5 Object inspection · causal dynamics
Cosmos +3.3 Object inspection · causal dynamics
BLINK +3.3 Passive physical events · physical modeling
TOMATO +3.1 Object inspection · temporal
TVBench +2.8 Camera motion · causal dynamics
BlackSwan +2.5 Object manipulation · causal dynamics
TempCompass +2.4 Object manipulation · causal dynamics
3DSRBench +1.4 Navigation · causal dynamics
EgoTaskQA +1.1 Passive physical events · physical modeling
PerceptionTest +0.8 Camera motion · causal dynamics
CoreCognition −1.0 Passive physical events · physical modeling
Gain over the base model on each external benchmark (percentage points; Qwen2.5-VL-3B-Instruct), using the best-performing controlled subset for that benchmark, named on the right; the gain is the larger of the SFT and GRPO gains. Bars are colored by the reasoning family of that subset; gray bars are the four benchmarks of static perception and general video QA. The dashed line marks the +2-point threshold for counting a benchmark as improved.

Visual transition reasoning serves as a shared training primitive

The gains so far could still be a sum of separate effects, each subset helping the tasks closest to it. For visual transition reasoning to serve as a shared training primitive, two requirements must hold.

(a) Different sources improve the same capability

Each subset raises WOVEN accuracy on action types it does not train on, in 43 of the 44 such cells below. The capability also spans reasoning operations: training on forward prediction alone raises counterfactual substitution and removal by 48 and 25 percentage points.

Training subset Passive physical eventsCamera motionObject inspectionNavigationObject manipulation Temporal adjacency
Camera motion · causal dynamics +5.1+33.0+13.4+26.9+39.8 +0.2
Object inspection · causal dynamics +6.8+8.9+39.9+27.2+37.7 +0.9
Navigation · causal dynamics +7.6+10.1+12.7+30.0+50.9 +4.9
Object manipulation · causal dynamics +8.2+3.9+5.7+26.2+61.4 +2.3
Object inspection · counterfactual +6.6+7.3+38.5+24.8+42.7 +1.3
Object manipulation · counterfactual +10.6+5.2+5.5+27.2+64.2 +2.9
Passive physical events · physical modeling +27.4−1.2+5.7+6.7+9.6 −1.5
Camera motion · temporal +8.4+10.9+4.9+8.0+1.3 +29.9
Object inspection · temporal +9.2+6.4+3.4+11.2+1.2 +33.1
Navigation · temporal +8.1+12.2+4.0+8.6+2.9 +32.0
Passive physical events · temporal +8.4+7.1+5.5+10.1+2.3 +30.6

Gain in WOVEN in-distribution accuracy (percentage points over the base model) after training Qwen2.5-VL-3B-Instruct on each controlled subset, by action type of the test items and on temporal adjacency. Outlined cells mark the action type the subset trains on; every other column is an action type the subset does not train on.

(b) The downstream gains result from that capability

If so, all subsets should improve largely the same benchmarks, and subsets that improve the capability more should transfer more. The transfer profiles of the 11 subsets are all positively correlated, and the more a subset improves WOVEN accuracy, the larger its average gain on the external benchmarks (r = 0.86 across the 11 subsets). The same WOVEN format with permuted or mirrored transitions does not reproduce the downstream gains, and the gains are specific to benchmarks that involve visual transitions.

We further test this by intervening on the training data. With the training set fixed at 2,000 items, we replace a growing share of task-specific items with a WOVEN subset on three benchmarks from three domains (SAT, ActionEQA, and CLEVRER), averaging over three seeds. Replacing 30–50% of the task-specific items leaves accuracy comparable to training on those items alone, and above a control that simply deletes the replaced items.

Three panels (SAT, ActionEQA, CLEVRER) plotting accuracy against the share p of task-specific training data replaced by WOVEN items under a fixed 2,000-item budget. Accuracy stays near the p = 0 level up to 30 to 50 percent replacement and stays above the deletion control.
WOVEN supervision substitutes for in-domain training data. Each panel retrains one external benchmark on a fixed 2,000-item budget in which a share p of the benchmark’s own training data is replaced by items from a selected WOVEN subset (blue dots: 3 seeds; blue line: seed mean; band: seed range; gray dashed: p = 0 mean; orange dashed: deletion control). On SAT, training exclusively on WOVEN yields the highest accuracy in the sweep.

A training recipe for visual world modeling

From the transfer profiles, error-typed analyses, and perturbation robustness of the controlled subsets, we distill four recipe items, then test the recipe prospectively on held-out benchmarks.

(1) The reasoning operation decides where supervision transfers

Similarity in scenes, actions, or application domains is an intuitive basis for selecting transferable supervision. To test it, we compare subsets that share only an action type with subsets that share only a reasoning family. All transfer profiles contain the common component from Question 3, so we remove it by standardizing the gains within each benchmark and correlate the residual profiles pairwise. Subsets teaching the same reasoning family have similar residual profiles (mean pairwise correlation +0.32), whereas subsets sharing only an action type are no more similar than subsets sharing neither (−0.22 versus −0.23). The same ordering holds under three other normalizations of the gains.

Three correlation matrices of the 11 subsets' transfer profiles: raw profiles with all pairs positive; residual profiles ordered by action type with no block structure; residual profiles ordered by reasoning family with diagonal blocks.
Correlation of transfer profiles under the two candidate groupings. Pairwise Pearson correlation of the 11 subsets’ 26-benchmark transfer profiles. Left: raw profiles, all pairs positive. Middle and right: the same residual matrix after the per-benchmark z-score, ordered by action type and by reasoning family; only the latter produces diagonal blocks.

(2) Larger state changes give more robustness

Among the causal-dynamics subsets, gains on the state-perturbation test set increase with how much the action changes the scene, and passive physical events give the largest gain of all subsets.

Camera motion +1.7
Object inspection +3.0
Navigation +13.6
Object manipulation +23.2
Passive physical events +29.9

Gain on the state-perturbation test set (percentage points over the base model) after training on each subset; the first four are causal-dynamics subsets, the last is the passive-physical-event subset.

(3, 4) Where transfer breaks

Supervision on one reasoning operation often teaches others, but two kinds of transition reasoning are not learned this way, nor are capabilities beyond state transitions. Temporal localization is learned only from temporal supervision: the four temporal subsets raise temporal-adjacency accuracy from 25.7% to 55.6–58.7%, whereas the seven other subsets leave it near the baseline. Agent-driven transitions are not learned from passive physical events: that subset contains no agent action to learn from, and it lowers accuracy on the two benchmarks built on an agent’s action and its outcome, WorldPrediction (−10.2) and ActionEQA (−1.7). The 4 benchmarks of static perception and general video QA do not improve under any training.

The recipe predicts transfer on held-out benchmarks

We use the recipe to select training supervision before any training, and evaluate on external benchmarks that played no part in its derivation (questions from MMSI-Bench, VLM4D, PhysBench, SPAR-Bench, MVBench, IntPhys 2, and CameraBench). For questions about agent actions and their outcomes, a mixture of WOVEN items composed by the recipe is compared with mixtures of the same size composed by matching the action (camera-heavy), by a different reasoning operation (temporal-heavy), by passive physical events (exogenous-heavy), or with no selection (uniform). Each mixture trains Qwen2.5-VL-3B-Instruct with 8 seeds. All five comparisons come out as the recipe predicts.

Temporal-heavy item (1) Agent-action questions +3.5 [+2.03, +4.92] 0.0002
Uniform item (1) Agent-action questions +1.5 [+0.18, +2.81] 0.03
Camera-heavy item (1) Camera-motion questions +0.6 [+0.02, +1.14] 0.04
Exogenous-heavy item (3) Agent-action questions +4.7 [+3.19, +6.17] 0.0002
Exogenous-heavy item (3) Passive-physics questions −0.5 [−1.05, −0.02] 0.04
Prospective test of the recipe. Δ = accuracy of the recipe mixture minus the comparison mixture (percentage points, mean over 8 seeds) with 95% confidence intervals. Item (1): selecting by reasoning operation beats the other mixtures on agent-action questions and beats the camera-matched mixture on camera-motion questions. Item (3): on passive-physics questions the exogenous-heavy mixture is ahead instead, as the recipe predicts.

Citation

If you find WOVEN useful in your research, please consider citing:

@article{fan2026woven,
  title={WOVEN: Weaving Visual World Modeling into Multimodal LLMs},
  author={Fan, Zheyu and Zhang, Yue and Deng, Mingkai and Wang, Kangrui and Wang, Qineng and Chen, Canyu and Hao, Jie and Fan, Xing and Guo, Chenlei and Xing, Eric P. and Bansal, Mohit and Li, Manling},
  journal={arXiv preprint arXiv:2610.12417},
  year={2026}
}