Natural correlations support predictive shortcuts
Context, motion, and future appearance are correlated in the training distribution, permitting appearance-based prediction without exclusive reliance on physical history.
Causal Writability in Video Models
When appearance cues conflict with physical history, which training-supported continuation controls generation, and does the rejected continuation remain causally accessible within the model?
Physical structure can remain causally available without controlling natural generation.
Natural generation identifies the selected continuation; intervention identifies alternative continuations that remain causally accessible.
decoded futures in the 64 × 11 cue sweep
endpoint shortcut failures corrected at the purple cue
compact route retaining nearly the full causal effect
successful directed transfers across run pairs
MECHANISTIC OVERVIEW
The analysis separates behavioral selection, causal accessibility, commitment with depth, and downstream realization.
Context, motion, and future appearance are correlated in the training distribution, permitting appearance-based prediction without exclusive reliance on physical history.
Physical history is held fixed while the appearance cue is varied continuously. The decoded continuation can switch between shortcut-consistent and history-consistent modes.
A state-conditioned internal edit changes the final decoded motion, establishing a causal effect on generation rather than a correlational probe readout.
A future is writable while condition-side intervention can redirect it. Commitment occurs when condition-to-target writes make the same intervention ineffective at subsequent depths.
NATURAL SOLUTION SELECTION
Fitted frequency and future RGB are measured directly from decoded rollouts. The trajectory, rather than the individual cue-conditioned rollout, is the statistical unit.
As the input cue changes from red to blue, decoded futures move from the slow mode through an intermediate region toward the fast mode.
The corresponding sweep under fast history separates the effect of cue strength from the direction of the physical evidence.
switch to the history-consistent frequency at the purple cue: 8/10 fast histories and 15/16 slow histories.
The denominator contains trajectories for which the strongly conflicting endpoint already produces the shortcut-associated frequency. This quantity is not an unconditional accuracy estimate.
With physical history and future window fixed, changing only the cue alters both fitted frequency and generated appearance. The selected object is therefore a joint appearance–dynamics continuation.
Long histories retain physics-consistent behavior under stronger cue conflict than short histories, providing a same-seed control for the strength of physical evidence.
These are natural generations. Within each pair, physical history, generation seed, and future window are fixed; only the appearance cue changes.
Same fast physical history; only the appearance cue changes.
Same slow physical history; only the appearance cue changes.
STATE-STRUCTURED CAUSAL CONTROL
Matched replacement identifies the writable route; low-rank analysis and a fit-only controller then test whether the route can be synthesized prospectively.
Replacing the receiver state with the matched target difference recovers the alternative continuation and identifies a functional route site. Donor activations are used only for route discovery.
The frozen coordinates trace a smooth phase-dependent family.
The opposite target direction remains structured by the same low-order boundary variables.
The four-dimensional edit approaches the decoded effect of the complete matched difference.
For each requested direction, the controller uses [1, cos θ*, sin θ*] to predict the top-four route coordinates from input-boundary phase.
After fitting and freezing the controller, a requested target direction and held-out boundary phase are sufficient to synthesize the condition-state edit. Held-out coordinate R² is .83–.88, compared with .51–.60 for direction alone.
The intervention changes the final pixel trajectory under fixed receiver input and noise.
The same held-out receiver, noise, and future window are used. The fit-only controller changes only the condition prefix at the frozen rewrite site.
The same fit-only PCA basis projects held-out raw activations and matched differences. B6, seed 3408, 50K; 128 fit pairs and 128 held-out pairs. Color denotes decoded frequency, with aligned frequency used for differences.
CAUSAL WRITEABILITY
Layerwise intervention profiles quantify where a selected future remains causally revisable.
The depth range over which intervention redirects shortcut failures becomes progressively narrower.
Across 15 seeds, mean full-grid physics-follow rate rises from .52 to .60 while remaining-error writability falls from 10.88 to 9.55 sites between 5K and 100K.
Failures later corrected by training (n = 157) are writable at 3.79 [1.57, 6.53] more sites than persistent failures (n = 953).
The effect appears in all three matched-seed 50K comparisons. Short and Long use checkpoint-local strict banks, not trajectory-paired banks.
DOWNSTREAM CAUSAL AUTHORITY
Target-route persistence, cross-run transfer, and path restoration localize how the shared route acquires downstream authority.
Route-aligned target responses remain measurable after condition-side frequency writeability has closed.
All six directed run-pair transfers achieve decoded recovery R = .94–.99 using fit-only scale and rotation.
Within a successful matched edit, restoring conflict K or K+V reduces final recovery to approximately zero, whereas restoring Q leaves recovery near one. The causal bottleneck is therefore localized to condition-to-target key/value writes.
Changing one of nine condition-token value heads yields a continuous dose response. Clean physics peaks at gain 8; continuous frequency recovery can overshoot at larger gains and increasingly leave the supported modes.
The same selection-clean receiver and future window are used; only condition-token V in head 8 is edited.
Network depth tells us where a write acts within one forward pass. Video generation also unfolds over 20 flow-matching calls, each passing through the network again. We therefore test when the same physical write can influence the final video.
Here, early and late refer to the first and second halves of the denoising sequence, not training checkpoints, network layers, or the first and second halves of the generated video.
In a separate 16-receiver comparison, neither calls 10–14 nor calls 15–19 rescues motion alone, whereas their joint 10–19 window does. The result is not simply “edit the last few steps”: it identifies an effective late window for this intervention. These tests do not establish a universal timing rule for every model, task or write.
SECOND-SYSTEM REPLICATION
Appearance and physical history compete in decoded rollouts, while low-order boundary state predicts a donor-free edit that recovers held-out dynamics.
Physical history, trajectory, renderer, generation seed, and future window are fixed. Only cue appearance changes.
64 fit pairs, 64 held-out pairs at B12. The raw view shows 128 endpoints; the difference view shows 64 edits. Point color is decoded frequency, using the aligned endpoint for differences.
Short/Long frequency_color_circle, using the paper's final checkpoints.
Balanced 32/32 by target direction; top-four coordinate basis and scale frozen on fit only.
Target-low and target-high edit coordinates follow direction-specific phase-aligned planes.
Both top-four oracle and fit-only top-four; median recovery .926 and .917, respectively.
ΔL₅₀ = +9.46 sites for target-low and +7.95 for target-high on n=38 receivers per direction.
All displayed rollouts pass the current appearance-tolerant geometry and frequency-fit gates.
NON-OSCILLATORY REPLICATION
A fit-only gravity controller changes the decoded trajectory in a non-oscillatory system, using the same 64-fit/64-held-out protocol as the paper.
B1, hist32, 100K. A basis fitted on 64 pairs projects 128 endpoints from the other 64 pairs; each input-color/gravity group contains 32 points. Color is decoded gravity, not frequency. Displaying PC3 does not change the rank-2 controller.
The same frozen pairs, 32 observed frames, and all 20 flow-matching calls are used. Endpoint projections reproduce the archived difference coordinates. The paper's gravity-error success rate is 61/64 for the controller and 0/64 for natural conflicts; this is distinct from normalized recovery.
SUPPLEMENTARY RESULTS
Figures 7–33 from the final-paper appendix cover behavior, pretrained adaptation, controller controls, Pendulum and Free Fall, training-time writability, and downstream mechanisms.
Changing appearance redirects both motion and generated color under fixed slow history.
Caption and methods in paper ↗The cue-dependent allocation of generated motion recurs across 15 independently trained models.
Caption and methods in paper ↗The evidence supports a compact, state-dependent causal route whose availability can be dissociated from its natural causal authority. The route jointly controls appearance and dynamics.
Alternative continuations remain internally callable even when they do not control natural generation.
The results do not imply classical disentanglement, a universal mechanism across solutions, or calibrated model confidence.
Xingyun Wang*, Haomin Zheng*, Man Yuan, Leqian Yang, Ziming Liu
Tsinghua University · Peking University · University of Science and Technology of China · MetaCircle · Shanghai Qi Zhi Institute
* Equal contribution. xingyun-24@mails.tsinghua.edu.cn · zmliu@tsinghua.edu.cn
@misc{wang2026causalwritability,
title={A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models},
author={Xingyun Wang and Haomin Zheng and Man Yuan and Leqian Yang and Ziming Liu},
year={2026},
note={Preprint}
}