A memory for what comes next

FrameMorrow

Future-Guided Frame Selection with Prospective Tokens
for Long-Horizon Video Generation

1 National University of Singapore2 Harbin Institute of Technology, Shenzhen

See the difference

Same backbone. Better memory.

Native generation alongside
generation with FrameMorrow.

Native backboneCausal Forcing
+ FrameMorrowCausal Forcing
0:00 / 0:59

Long video generation

Follow the subject and scene as the video progresses. Both videos use Causal Forcing, with historical frame selection added on the right.

Synchronized playback

The intuition

Relevant now ≠ useful next.

Select history for future needs.

A corridor may match the current view. But returning to a room requires remembering its earlier appearance. FrameMorrow learns to retrieve historical evidence for the upcoming generation.

Conceptual illustration: context matching retrieves a corridor, while future-guided selection retrieves the earlier room and armchair needed for the next scene.
Conceptual illustration of current similarity versus future usefulness.

A compact future proxy

Look ahead. Select history.

No full future rollout required.

01

Anticipate information needs

Prospective tokens summarize what may matter next, using available history, recent context, and the generation condition.

02

Retrieve explicit frames

The tokens score historical candidates. A small set of selected frames carries relevant visual evidence forward.

03

Reuse across generators

Selected frames enter each backbone through its own conditioning interface. The selector stays outside the generator.

FrameMorrow overview: inference uses prospective tokens to select historical frames, while training distills future-derived teacher rankings into the selector.
Realized future observations supervise the selector during training only. Inference uses available history and the current generation condition.

Quantitative results

Measured across generators.

Orange rows: + FrameMorrow
↑ Higher is better · ↓ Lower is better

01Long video generationMovieGenBench · 60 seconds
VBench-Long on 128 MovieGenBench prompts
Model / methodSubject ↑Background ↑Motion ↑Dynamic ↑Aesthetic ↑Imaging ↑Avg. rank ↓
Self-Forcing
Native95.8495.2798.2051.7256.0562.224.33
+∞-RoPE97.2496.2498.5846.6456.0963.283.17
+Deep Forcing96.0895.3898.2441.4456.6860.814.17
+LongLive-RAG97.6096.5198.7044.6957.1964.972.00
+FrameMorrow97.7196.4598.6064.5657.2768.471.33
LongLive 1.0
Native97.1395.8998.6144.5658.1767.563.83
+∞-RoPE97.0095.8598.5353.3657.4866.944.25
+Deep Forcing97.1796.0498.7345.1357.4867.273.42
+LongLive-RAG97.3296.0898.6249.9058.3067.792.25
+FrameMorrow97.3296.1798.7550.4258.7567.821.25
Causal Forcing
Native93.5294.1295.7472.3251.2462.304.83
+∞-RoPE93.8193.7896.0992.4754.4267.503.33
+Deep Forcing94.2794.1896.6278.5952.1264.253.17
+LongLive-RAG94.2994.2496.4888.2054.9568.162.33
+FrameMorrow94.5494.2896.5989.7456.2768.761.33

Bold marks the best score within each backbone. Average rank is computed across the six metrics; lower is better.

02Interactive video generationSingle-shot & multi-shot · 60 seconds
Interactive video generation
Model / methodQuality ↑Consistency ↑Aesthetic ↑Avg. CLIP ↑
Single-shot
LongLive 1.083.4992.6264.2128.75
+FrameMorrow84.0693.2164.5629.08
Self-Forcing78.7184.9558.1626.39
+FrameMorrow81.7489.3060.6228.05
Causal Forcing76.0080.4655.4524.53
+FrameMorrow78.6683.4858.7525.24
CausVid81.9489.7763.5328.63
+FrameMorrow82.2690.5263.6429.03
Multi-shot
LongLive 2.082.7793.1461.3228.35
+FrameMorrow83.1893.4961.6828.44
ShotStream83.1292.7562.3029.09
+FrameMorrow84.8596.7661.7429.16

Bold marks the better score within each backbone pair. Average CLIP score summarizes six consecutive 10-second segments.

03Interactive world modelsAction-conditioned generation
Action-conditioned interactive world generation
Model / methodSubject ↑Background ↑Imaging ↑Anti-flicker ↑Motion ↑Action alignment ↑
Matrix-Game 3.00.8010.89468.200.9360.9500.842
+FrameMorrow0.8160.91068.350.9390.9490.861
WorldMem0.7820.92470.230.9100.9160.866
+FrameMorrow0.7930.93270.470.9130.9180.889
YuMe 1.50.7650.87250.980.9440.9650.883
+FrameMorrow0.7980.88651.800.9470.9670.912

Bold marks the better score within each backbone pair. Imaging uses a 0–100 scale; the other metrics lie in [0, 1].

One idea, multiple settings

Memory beyond a single generator.

Evaluated across long-video generation, interactive generation, and action-conditioned world models.

5benchmarks
11generative models