Designing 36 LTX-2 scenes and clip continuity
A video-generation AI workflow design separating scenarios, shots, and continuity. It covers fixed schemas, audio, camera instructions, concatenation conditions, and limitations in generated samples.
LTX-2 scenarios and shot design
I separated scenario structure, individual shots, and clip continuity for a 36-scene LTX-2 workflow. These are prompt and downstream-processing conditions for integrating video generation into AI systems.
- 36-scene horror scenario generation template — system prompt for LLM-driven scenario writing
- Cinematic prompt design principles — structure for maximizing individual shot quality
- Multi-scene continuity control — pipeline design for preventing inter-clip visual breakage
Prerequisites
- Video generation engine: LTX-2 (5-second clip generation)
- Resolution: 3840x2160 (4K), 24fps
- Scenario generation: Local LLM (Mixtral, etc.) generating STORY→SCENES→LTX2_PROMPTS in one pass
- Pipeline: Per-scene generation → ffmpeg concatenation → audio composited downstream
1. 36-Scene Horror Scenario Generation Template
Design Intent
Monstral-123B NVFP4 scenario generation had these problems:
- Scene counts varied between 20–50; achieving a fixed 36-scene structure was impossible
- Narrative structure collapsed; act structure (setup/escalation/resolution) became unclear
- Dialogue disappeared midway; mute scenes proliferated
- Endings became ambiguous or devolved into abrupt “it was all a dream” tropes
I fixed the output schema and rules to obtain 36 scenes from mid-size models. The review addressed missing format constraints as well as model capability.
Output Schema (Fixed)
Allow only these five output sections, without commentary.
- STORY_PROMPT — narrative skeleton (one paragraph)
- LABELS — genre, tone, motifs
- OUTLINE — 5-stage plot summary
- SCENES_36 — Scene 01 through Scene 36 (1-3 sentences each)
- LTX2_PROMPTS_36 — Shot 01 through Shot 36 (generation prompt per shot)
Require Scene 01 through Scene 36 without gaps, extras, or surrounding commentary for downstream parsing.
Scene Allocation
Scenes 01–08 : Daily life and subtle unease (first whispers, minor anomalies)
Scenes 09–20 : Escalation and repetition (voices repeat, reflections speak back)
Scenes 21–28 : Investigation and confrontation (identifying the cause, revealing truth)
Scenes 29–36 : Resolution and closure (cause identified → concrete action → safety restored)
Prohibit dream-only, imagination-only, and unresolved endings.
Required Elements Per Scene
Scene 04:
Visual: A woman stands before a mirror. Camera slowly zooms into her reflection.
Whisper: "You see it too, don't you?"
Sound: Fluorescent light flickers with a low humming sound.
- Visual: what is on screen
- Camera: framing or camera movement
- Dialogue: at least one spoken line (Dialogue / Whisper / Voice(V.O.) / Heard Voice / Inner Voice)
Require speech in every scene. Use whispers, distorted voices, inner monologue, reflected voices, or remembered phrases where dialogue is unnatural. This puts audio rhythm into scene planning.
Setting Rules
- Japan (urban or suburban), deliberately vague location names
- Cultural elements: quiet residential streets, small houses, mirrors, shrines, corridors, sliding doors, evening TV noise, cicadas, wind, fluorescent lights
- No graphic gore or explicit violence. Fear is built through psychological means: whispers, shadows, reflections, repetition, isolation
Avoid real place names and use ordinary Japanese settings and sounds to create psychological unease.
Horror Constraints
Allowed elements:
- Whispers, shadows, reflections, repetition
- Voices that cannot be confirmed as real
- Emotional isolation
Not permitted: graphic gore, explicit violence, dream-only endings, “it was all imagination” endings, unresolved ambiguity.
Use repetition, sound, and reflection changes rather than jump scares or graphic violence. The aim is small anomalies that remain manageable across generated clips.
Prohibition Block
State the prohibitions explicitly.
The ending must be resolved.
No dream ending.
No insanity-only explanation.
No ambiguity.
Story Flavors (Presets)
Presets for vague user input:
| Flavor | Premise |
|---|---|
trapped_room_mummy | Locked in before summer break, only traces remain |
water_reflection | Water surface reflections gradually rewrite reality |
station_beep_phrase | Station chimes become words that guide people |
fox_shrine_wrath | Removal of a small shrine triggers quiet anomalies |
Each flavor has a logline, key beats, and a resolution pattern, defining the ending before generation.
story_core Template
A narrative skeleton with placeholders:
{SETTING}
{PROTAGONIST}
{HELPER}
{FLAVOR}
{RESOLUTION_PATTERN}
{MOTIFS}
Use psychology, sound, shadow, water, and repetition, with restrained violence. This layer defines narrative progression separately from shots.
LTX2_PROMPTS_36 Shot Definition
Shot 01:
DURATION=5s FPS=24 RES=3840x2160
PROMPT: <concrete visual instruction, short>
NEGATIVE: low quality, blurry, distorted hands, deformed face, gore, blood
CAMERA: handheld POV / slow push-in / over-the-shoulder
LIGHTING: low-key, streetlight, fluorescent hum
AUDIO_CUE: <ambient, dialogue, silence>
CONTINUITY: prev=none next=shared_object:mirror
SEED_HINT: episode_seed+01
Avoid face close-ups and prefer hands, backs, or silhouettes. Add subtitles/telop later in ffmpeg.
Separate fields allow shared NEGATIVE, replaceable SEED_HINT, generated CONTINUITY, validation, and retries.
LTX-2 Style Defaults
DURATION : 5 seconds
FPS : 24
RESOLUTION: 3840x2160 (4K)
STYLE : cinematic, realistic, subtle horror, Japanese suburban atmosphere
CAMERA : handheld POV / slow push-in / over-the-shoulder (avoid face close-ups)
LIGHTING : low-key, streetlight, fluorescent hum, rainy reflections
Shared negative prompt and seed strategy:
2. Writing each shot
Start with camera position
Specify camera position first using cinematography terms rather than “the camera shows”.
static camera/slow pan/close framing/wide interior shot/shallow depth- Specify camera position, field of view, and spatial compression
Position and lens cues come before narration to reduce variation.
Keep lighting, color, and texture consistent
Use shared lighting, palette, and texture parameters.
| Element | Examples |
|---|---|
| Lighting | warm, cold, flickering, fluorescent, natural, overcast |
| Color palette | muted pastels, warm golds, sickly greens, neutral grays |
| Surface textures | fogged glass, worn metal, dust in light beams |
| Mood | cozy, oppressive, sterile, playful |
These are intended to reduce scene-to-scene drift.
Action as Continuous Physical Sequence
Describe the physical progression in order:
Leaning into frame → Hesitating near the handle → Exhaling slowly to fog the glass
Use arrows (→) for intermediate movement. Abbreviated actions can produce teleportation, so describe how the body moves through space.
Character Definition Through Behavior
Specify posture, expressions, timing, habits, age, clothing, and emotion through movement. Pixar-style acting is the reference; avoid static attribute lists.
Camera Movement as Narrative Tool
- Slow pan: observation and tension building
- Static shots: tension or comedic timing
- Avoid sudden or unmotivated movement unless it supports a punchline
- Maintain consistent pan direction and speed across adjacent clips
- When continuing from previous scene, explicitly state
Continue the same pan
Continue the same pan carries camera motion into an adjacent clip.
Audio Is Mandatory
Audio also defines timing:
- Ambient Sound: oven humming, steam, distant noise
- Dialogue Beats: quoted dialogue,
[Beat]and[Silence]markers for pauses - Speaking mode:
whisperingvsspeakingvsmuttering - Music must be explicitly included or excluded
Include audio from the start so later sound matches the action.
Prompt Structure Template
Use this seven-step order:
- Shot Establishment
- Environment & Lighting
- Character Position & Emotion
- Core Action Sequence
- Camera Movement (with timing)
- Audio & Dialogue
- Ending Visual Beat (pose, pause, reaction)
Specify Ending visual beat separately to define how the shot ends.
Example: Cinematic Prompt (approx. 600 words)
The INT. OVEN – DAY. example uses all seven steps.
INT. OVEN – DAY.
Static camera positioned from inside the oven, looking outward through the slightly fogged glass door. The frame is tight and claustrophobic, bordered by dark metal edges. Warm golden light fills the interior, reflecting softly off the glass and illuminating a tray of freshly baked cookies resting just below frame center. Fine steam curls upward, leaving faint streaks across the glass.
The color palette is rich and inviting: honeyed browns, soft ambers, and glowing highlights. The air feels warm and dense, as if the oven itself is breathing.
A baker leans into frame from outside the oven door. His face slowly fills the glass, eyes narrowed in intense concentration. His breath fogs the glass further as he exhales, creating overlapping clouds that briefly obscure his features before clearing again. He doesn't blink. His posture is rigid, shoulders slightly hunched, as though this moment demands absolute precision.
The camera remains static as subtle reflections ripple across the glass from the rising heat.
Baker (whispering, reverent):
"Today… I achieve perfection."
He tilts his head slightly, adjusting his angle to catch the light just right. His eyes track the edges of the cookies, watching for the slightest change in color. One hand lifts slowly into frame, hovering near the oven door handle but not touching it yet.
The steam grows thicker for a moment, then thins.
Baker (still whispering, building intensity):
"Golden edges. Soft center."
He leans even closer, nose nearly pressing against the glass. The camera holds, letting the silence stretch just long enough to feel deliberate.
Baker (a hushed gasp):
"The gods themselves will smell these cookies… and weep."
A beat.
The baker's eyes widen slightly as a new thought intrudes. His hand freezes mid-air.
Baker:
"Wait—"
Silence. The hum of the oven is suddenly very noticeable.
Baker (uncertain, quieter):
"Did I… forget the chocolate chips?"
Another beat. Just long enough to be uncomfortable.
Cut to a side angle outside the oven. The camera now slowly pans from left to right, revealing a coworker casually stepping into frame. The coworker chews loudly, crumbs at the corner of his mouth, utterly unconcerned. The lighting here is flatter and cooler, contrasting the oven's warmth.
Coworker (mouth full, casual):
"Nope. You forgot the sugar."
The pan continues for a moment as the coworker shrugs and wanders off.
Cut back inside the oven.
A quick push-in as the baker's face slams back into frame, pressed against the glass. His expression is pure horror. His breath fogs the glass completely now, obscuring the cookies behind it.
Behind the fog, the cookies visibly sag and deflate in slow motion, their once-perfect shape collapsing inward. Steam rises dramatically, drifting upward like a defeated sigh.
No music. Only the oven's hum and the faint crackle of heat.
The baker stares, unmoving.
Baker (barely audible):
"...No."
The camera holds on the fogged glass as the steam slowly dissipates, revealing the ruined cookies one last time before the shot ends.
Its components are:
- Shot establishment:
static camerainside the oven, looking outward — frame and spatial constraints defined before any character appears - Environment:
warm golden light,honeyed browns,fogged glass— palette and texture locked early - Character: defined through posture and behavior —
He doesn't blinkdoes more work than a paragraph of description - Action: leaning, exhaling, hovering, freezing — continuous sequence with no teleports
- Camera: the slow pan appears only when the scene has a reason to reveal new information; the static holds are load-bearing
- Audio: whisper, silence, oven hum, and heat crackle are all embedded in the dramatic timing
No music. Only the oven's hum and the faint crackle of heat. removes music and leaves oven sound and silence.
3. Multi-Scene Continuity Control
The Problem
Multiple scenes introduce:
- Visual drift: lighting and texture shift between shots
- Camera discontinuity: pan speed and direction mismatch across clips
- Timing loss: inability to control acting “beats” and pauses
Last-Frame Continuity
- Continuation directive: explicitly state
Continue the same panat the start of the next scene - Seed incrementing: use a fixed episode seed plus shot index (
episode_seed + shot_idx). Allows micro-variation while maintaining macro-consistency - Texture inheritance: visual states from previous shots (e.g., fogged glass) become preconditions for the next shot
CONTINUITY Field in Practice
Use CONTINUITY with prev= and next= to connect objects, sounds, actions, or locations.
CONTINUITY: prev=shared_sound:fluorescent_hum next=location_transition:hallway_to_kitchen
These directives keep adjacent shots from being planned independently.
Avoiding Teleportation
Describe intermediate physical motion rather than skipping directly to the end position.
Evaluating the Sample Outputs
OUTPUT 1 (Fox Curse)
STORY_PROMPT, LABELS, OUTLINE, and SCENES_36 follow the schema and form a story. Continuity and camera instructions remain too sparse for direct LTX-2 generation.
OUTPUT 2 (Gon’s Mutation)
Gon loses his sister to a bear, seeks revenge, is cursed, mutates, and reaches acceptance and reunion. The ending is clearer than OUTPUT 1, but camera instructions remain sparse.
Variable Parameterization
MAIN_CHARCTER="Gon", SUB_CHARCTERS="Mosuke, Yayoi", and FLAVOR=... are experimental notation, not a finished schema. The aim is reusing a scene skeleton with different characters and flavors.
Proposed workflow changes:
- Consolidate the repeated system-prompt blocks into one canonical version
- Move
OUTPUT 1andOUTPUT 2into a dedicated examples section - Replace ad hoc variables with YAML or JSON inputs to stabilize parsing
- Add automatic validation between
SCENES_36andLTX2_PROMPTS_36
Automate checks for prev/next consistency, concrete resolution actions, and accidental silent scenes.
Reproduction Steps
Minimum Setup
- LTX-2 video generation environment (ComfyUI or API)
- Scenario generation LLM (local or API)
- ffmpeg (clip concatenation and telop compositing)
Horror Scenario Generation Workflow
- Select a Story Flavor (or provide free-form input)
- Feed the system prompt template from this article to the LLM
- Retrieve SCENES_36 and LTX2_PROMPTS_36 from output
- Feed each shot’s PROMPT to LTX-2 sequentially
- Concatenate generated clips with ffmpeg
- Composite audio and telop downstream
Cinematic Quality Checklist
- Camera position specified at the start of each prompt
- Lighting and color palette consistent across scenes
- Actions described step-by-step with arrows (→)
- CONTINUITY defined between adjacent shots
- Audio directives included
- Face close-ups avoided; hands/backs/silhouettes preferred
