Narrative and visual consistency
VBench evaluates overall consistency, subject consistency, aesthetic quality, and motion smoothness. Each scene is 5.375 seconds. Baselines are evaluated with both repeated-seed generation and scene descriptions from CineAGI’s Script Writer; scores are averaged across the two strategies.
Scroll sideways to see all metrics →
Main paper, Table 1. Values as reported in the ICME manuscript.| Method | Overall consistency ↑ | Subject consistency ↑ | Aesthetic quality ↑ | Motion smoothness ↑ |
|---|
| CogVideoX | 0.096 | 0.823 | 0.379 | 0.960 |
|---|
| VideoCrafter2 | 0.076 | 0.885 | 0.364 | 0.920 |
|---|
| Hunyuan | 0.185 | 0.909 | 0.569 | 0.976 |
|---|
| CineAGI | 0.259 | 0.949 | 0.600 | 0.987 |
|---|
All four metrics are better when higher. The paper reports paired tests over 100 prompts against the best baseline: p < 0.01 for overall and subject consistency, and p < 0.05 for aesthetic quality and motion smoothness.
Overall consistency improves from 0.185 to 0.259 over Hunyuan (+40.0% relative). Subject consistency increases from 0.909 to 0.949 (+4.4% relative).
How people experience the stories
Twenty participants, including ten with professional multimedia experience, rated five aspects of the videos on a 1–5 Likert scale.
Scroll sideways to see all metrics →
Main paper, Table 2. Values as reported in the ICME manuscript.| Method | Visual quality ↑ | Narrative coherence ↑ | Character consistency ↑ | Audio coherence ↑ | Overall quality ↑ |
|---|
| CogVideoX | 3.16 | 2.52 | 2.21 | — | 2.63 |
|---|
| VideoCrafter2 | 2.75 | 2.26 | 1.98 | — | 2.45 |
|---|
| Hunyuan | 3.52 | 2.91 | 2.44 | — | 2.88 |
|---|
| CineAGI | 3.83 | 3.57 | 3.14 | 3.26 | 3.37 |
|---|
Higher is better. A dash under audio coherence means that the baseline does not support the evaluated audio feature; it is not assigned a zero score.
Character consistency increases from 2.44 to 3.14 / 5 over Hunyuan (+28.7% relative). Narrative coherence increases from 2.91 to 3.57.
Planning and integration both contribute
Each ablation removes one part of the framework: narrative synthesis, the Quality Inspector, or Decoupled Character Integration.
Scroll sideways to see all metrics →
Main paper, Table 3. Values as reported in the ICME manuscript.| Method | Overall consistency ↑ | Subject consistency ↑ | Aesthetic quality ↑ | Motion smoothness ↑ |
|---|
| Without narrative synthesis | 0.232 | 0.924 | 0.575 | 0.974 |
|---|
| Without quality inspector | 0.245 | 0.938 | 0.570 | 0.982 |
|---|
| Without character integration | 0.206 | 0.911 | 0.583 | 0.971 |
|---|
| Full CineAGI | 0.259 | 0.949 | 0.600 | 0.987 |
|---|
These are automatic metrics from the main paper’s ablation table; higher is better throughout.
Removing character integration reduces overall consistency from 0.259 to 0.206 and subject consistency from 0.949 to 0.911—the largest drops among these three variants.