WAN2.2 TI2V-5B · DANCEGRPO

See where training moved the needle.

Task-level change, sample-level evidence, and every scored video from the strict In-Domain sweep—always grounded against the same SFT baseline.

20
evaluated runs
10,500 videos · 100 tasks

01 / RUN SELECTION

Choose a checkpoint and sampler

BASELINE SFT epoch 1
Sampling mode

02 / SCORE TRAJECTORY

Overall score across training

Step 0 is the shared SFT ODE baseline; all runs use the same 500 samples.

03 / TASK LEDGER

Every task, ranked by change

Task Domain / category Current Baseline Δ vs baseline Inspect

Evaluation prompt

SAMPLE EVIDENCE

Selected run vs SFT baseline

Five matched samples. Scores are exact values from the scorer JSON.

V

Indexing the evaluation archive