Lore

AI video generation

https://www.youtube.com/watch?v=3rDs6FhFoUQ

A full AI-generated product commercial can be made on a laptop through a repeatable three-step workflow — build tested, locked visual assets; turn the script into a single connected shot list via a Claude skill; then generate and iterate scenes in a video model — and the real skill is not writing a perfect prompt once but iterating hundreds of times and cutting together the best seconds.

unknown · English

Key ideas

  1. The workflow has three steps: assets (Soul Cinema + GPT Image 2.0), a Claude skill that converts the script into a structured shot list, and scene generation in SeaArt/Citas 2.0, all inside Kixal AI.

  2. A single reference image of a product or character is insufficient; the video model needs multi-angle sheets or it hallucinates mid-scene.

  3. Generate multiple candidate assets, test each in a simple scene, and only lock the ones that hold up in motion — a still that looks strong can collapse once it moves.

  4. Every edit to an asset degrades its quality toward a flat 'AI slop' look, so state changes (sweat, new outfit) should be pre-generated as new reference sheets rather than text-edited mid-pipeline.

  5. The shot list should function as one connected document with a shared style-prefix block, so a single edit propagates to every prompt instead of managing loose separate prompts.

  6. Generic motion language produces generic results; choreography has to be spelled out move-by-move by name.

  7. Text alone cannot reliably pin down spatial layout; feeding the model a schematic/map of the scene anchors character and prop position and replaces brute-force 20-generation guessing.

  8. The finished ad is assembled from the best few seconds out of roughly 100 generation attempts — iteration, not first-try perfection, is presented as the actual skill.

  9. Product sheet generation — Using GPT Image 2.0 to turn a single product photo into a reference sheet with front and 3/4-angle views. Apply: Feed one product image into GPT Image 2.0 and prompt for a front + 3/4-view product sheet before using the product in any scene.

  10. Character sheet (multi-pose, gray background) — A Soul Cinema-generated two-panel reference (headshot + full body, front and back) shot on a neutral gray background for consistency. Apply: Generate character sheets on gray backgrounds with both close-up and full-body, front and back poses, to raise win rate and give the video model a single clean reference.

  11. Multi-candidate generation + lock-in testing — Generating several candidate versions of an asset and testing them in motion before committing to one. Apply: Generate multiple face/asset candidates, run each through a simple test scene, and only lock the ones that hold up once they move.

  12. Canvas comparison workspace — A side-by-side visual layout for comparing candidate assets. Apply: Arrange generated candidates next to each other on the canvas and drag the locked winners to the top for reference during later steps.

  13. AI Cast — A tool that instantly produces character sheets with multiple styling variants from a single text prompt. Apply: Use AI Cast to quickly generate secondary/background characters instead of building full manual character sheets for minor roles.

  14. Three-quarter angle location shots — Composing location reference images at a three-quarter angle rather than head-on. Apply: Shoot/generate location references at a 3/4 angle to give the video model more depth to track during camera movement.

  15. Controlled variable (single-change) testing — Isolating which asset affects generation quality by changing only one element per test. Apply: When troubleshooting a scene, change only one asset (e.g., the face, the location) per test run while holding everything else constant.

  16. Pinpoint image-to-image edits — Using GPT Image 2.0's image-to-image mode to add, remove, or modify specific elements within an existing scene image. Apply: Target a specific object in a location image (e.g., remove clutter, add a stove) via image-to-image editing instead of regenerating the whole scene.

  17. Face consolidation — Removing duplicate faces from a character sheet so only one face remains for the model to lock onto. Apply: Erase extra/duplicate faces from a character sheet before feeding it to the video model to prevent facial drift.

  18. Outfit compositing via layer masking — Combining a GPT Image 2.0-edited outfit with the original Soul Cinema shot using photo-editor masking to preserve facial/skin detail. Apply: Layer the original Soul Cinema image over an outfit-edited GPT Image 2.0 version and mask out the old outfit so the new one shows through while keeping the original face and skin.

  19. Claude shot-list skill — A preloaded Claude skill/instruction file that teaches Claude to write shot lists optimized for the video generation tool. Apply: Load the skill into Claude and have it convert a written script into a structured, tool-ready shot list.

  20. Shot list as one connected document — Structuring all scene prompts as a single linked document rather than separate isolated prompts. Apply: Keep every scene prompt inside one shot-list document with a shared style-prefix block so a single edit updates every prompt at once.

  21. Named project elements — Saving locked assets as named elements within the project so they auto-attach when referenced. Apply: Save each locked character/prop/location as a named element and reference it by name in prompts so it attaches automatically.

  22. Frame salvage across takes — Pulling the best individual frames from multiple separately generated video clips of the same prompt. Apply: Run the same prompt several times, then extract and composite the best keeper frames from different resulting clips.

  23. Move-by-move choreography scripting — Replacing vague motion instructions with an explicit named sequence of actions. Apply: Instead of prompting 'he dances,' list the exact moves in order (e.g., two head nods, a shoulder roll, a knee dip, a finger snap) for the video model to execute.

  24. Match cutting on action — Editing technique of cutting between scenes on matching motion for visual continuity. Apply: Align the ending action of one generated scene with the opening action of the next so the cut reads as continuous motion.

  25. Multi-version character sheets for state changes — Pre-generating a second full character sheet for a needed state change (e.g., sweated, different outfit) instead of text-editing the state mid-pipeline. Apply: When a character needs to change state or outfit later in the story, build a dedicated new reference sheet for that state up front rather than prompting the change via text edits, since edits degrade quality and videos are costly to regenerate.

  26. Outfit-locking via reference image — Generating a consistent locked outfit asset by feeding a written prompt plus the character sheet and prop images into GPT Image 2. Apply: Before generating multiple video takes, lock the outfit as a reference image using GPT Image 2 fed with the character sheet and relevant props, to prevent outfit drift when stitching clips from different generations.

  27. Scene-specific prefix override — Overriding the shot list's global style-prefix block for an individual scene whose lighting/setting differs from the default. Apply: When a scene (e.g., a midday stadium) needs different lighting than the global style prefix (e.g., a kitchen setup), write a scene-specific prefix override instead of editing the shared block.

  28. Spatial schematic/map for layout anchoring — A generated map/schematic image fed into the prompt to fix character and prop positions, since text instructions can't reliably pin down a location. Apply: Generate a schematic of the scene layout in GPT Image 2, attach it to the prompt to lock subject/prop positions, and reuse it across related scenes instead of brute-forcing many generations for positional luck.

  29. Music track as model input — Embedding the actual audio track into the prompt as an input so the model can synchronize character movement to it. Apply: Drop the music track into the prompt as an input and instruct the model to have the character dance in time with its beat.

  30. Iterative single-prompt refinement — Improving a shot by repeatedly refining and rerunning the same prompt rather than writing separate new prompts. Apply: When a generated scene isn't right, iterate on the same prompt across multiple passes and keep the best resulting clips, rather than starting over with parallel new prompts.

Insights

A gray background on character sheets is claimed to raise win rate significantly over colored backgrounds because there's no clutter competing visually with the subject.

Putting more than one face on a single character sheet makes the video model uncertain which face to lock onto, causing facial drift during generation.

Outfit continuity breaks when clips from different generation runs are stitched together (shoes, shorts, etc. vary run to run) unless the outfit is explicitly locked via a generated reference image first.

A body-rig/snorricam-style locked tracking shot, which traditionally requires a full day shoot with a dedicated camera operator and costs thousands of dollars, is reproduced here through the generative pipeline instead.

The spatial schematic/map technique is explicitly framed as the single highest-leverage trick in the whole video, eliminating the need to brute-force 20 generations to fish for a few usable seconds.

Embedding the actual music track as a prompt input lets the video model synchronize the character's dance moves to the beat, rather than relying on text descriptions of rhythm.

«what you just watched is a full commercial I made completely with AI on this laptop»

— 00:53

«A real cinematic ad. And honestly, it's way simpler than it looks.»

— 01:09

«The model needs to see them from every angle, or it'll hallucinate halfway through our scene.»

— 02:08

«A face that looks great as a still might fall apart the moment it moves»

— 04:00

«Every time you edit an asset, the quality takes a hit.»

— 11:09

«This isn't a pile of separate prompts. It's a shot list. One connected document.»

— 14:10

«You have to spell it out move by move.»

— 19:37

«The fix here isn't words. Text just can't pin a location down. So, I give the model a map instead.»

— 22:42

«Honestly, this is the number one hack to steal from this video. Normally, you'd use brute force, running 20 generations just to fish out a few good seconds. That's slow and burns through credits. The map skips the lottery.»

— 22:46

«the finished ad is just the best few seconds out of 100 tries cut together. That's the whole game. Iteration is the skill.»

— 25:00

A methodical, technique-dense walkthrough that frames AI ad production as an engineering-and-iteration problem rather than a single clever prompt, with the shot-list-as-connected-document and spatial-schematic tricks standing out as its most transferable contributions.

↳ unknown · YouTube

Watch original