AI video generation
The video, hosted by Adil of Higgsfield AI, argues that a professional, cinematic football/robot commercial can be built from scratch with no studio and a tiny budget using a three-step AI workflow — asset creation, a purpose-built 'cloth skill' prompting framework, and iterative scene generation in Cines 2.0/Seedance — where Claude does most of the prompt-writing, and the real skill is directed iteration across many generation batches rather than nailing a single prompt.
Three-step workflow: (1) build assets, (2) load the cloth skill prompting framework, (3) generate scenes in Cines 2.0/Seedance.
Location is called the most important image in the whole project; a sloppy or plasticky location ruins the video regardless of prompt quality.
Soul Cinema is presented as the best model for likeness consistency, built from many reference images into a character sheet.
Props (the can, football, goal) must stay visually consistent across every scene.
Most prompting failures come from overloading prompts with unnecessary information, wasting time and credits.
A shared style header (lighting, lens, acting style, physics) is kept identical across every shot's prompt.
You don't hand-edit prompts; you watch the output, give plain-language director's notes, and have Claude rewrite the prompt.
First scene is treated as decisive — if it feels off, viewers stop watching.
Shot-by-shot camera directing (describing what the camera sees, not just what happens) enables specific angles like low-angle handheld shake or over-the-shoulder.
Complex action scenes need many generation batches, then the best phases from different generations are cut together.
A borrowed sports-film storytelling rule: the hero must lose or get knocked down before the comeback win, or the win won't land.
Objects 'teleporting' between generations happen when the model has no spatial anchor; text-only position prompts repeatedly fail.
A drawn location scheme (with compass, distances, dimensions) fixes spatial-consistency failures that text prompts couldn't.
Single-use props don't need a separate asset build; the model can generate them inline within a scene prompt.
Overly packed scenes should be split into multiple prompts so each beat has room to breathe.
Physical effects (falls, impacts, armor breaking) need explicit physics language (gravity, metal snapping) to look real.
Macro close-up shots on small details (chest plate, wrists, helmet) raise perceived production value.
Unused batch phases and shots from earlier scenes are reused later instead of regenerating from scratch, to save credits.
For a missing shot, only that shot is generated and stitched into already-approved footage, rather than regenerating a whole scene.
The final pack shot (last scene) is treated as the one that has to sell, since it's what viewers remember.
The final film is described as the best few seconds selected out of roughly 100 generation attempts; iteration, not first-try perfection, is the actual skill.
Three-Step Workflow (Assets → Cloth Skill → Scene Generation) — The video's overarching production framework: build assets first, load a prompting-framework system prompt, then generate scenes in the video model. Apply: Lock all characters, locations, and props before writing a single scene prompt, then run every scene through the cloth-skill-primed Claude before pasting into the video generator.
GPT Image 2.0 — An image-editing model the video calls the best current option for editing, used to build product sheets and location schemes from reference images. Apply: Feed it a reference image and an instruction like 'make a product sheet with front, back, and top shot' or turn a location photo into a formal scheme with compass and distances.
Soul Cinema — A character/asset generation model described as best for preserving a person's likeness and producing movie-still-quality images. Apply: Upload roughly 20 reference photos of a face, generate a two-panel character sheet (full body + close-up) at 16:9/2K across many batches, and pick the best pairing to reuse across all scenes.
Higgs Field Colab project workspace — A project mode inside the Higgs Field platform that keeps all image and video models and generated assets for one production in a single shareable space. Apply: Create a dedicated Colab project per shoot so every generated asset stays in one place and is easy to hand off to a team.
Cines 2.0 / Seedance — The video generation platform where finished, Claude-written prompts are pasted to produce the actual scene clips. Apply: Paste each scene prompt at 16:9, 1080p, 15 seconds, running multiple batches (e.g., four) per scene to get several options to choose from.
Cloth skill — A downloadable system prompt the creators say took months to perfect, meant to skip the most painful part of manual prompt iteration. Apply: Load the cloth skill into Claude so it automatically enforces one consistent style header, lighting, lens, acting style, and physics across every shot it writes.
Named element auto-attach system — A tagging feature in Colab (referenced as 'Hexel/Hexol') that lets locked assets be named so they auto-attach whenever a prompt referencing that name is pasted. Apply: Name each locked asset once (e.g., 'city location', 'Robo player two') and reference it by name in later prompts instead of re-uploading its image each time.
Product sheet prompting — A prompt pattern that asks GPT Image 2.0 for front, back, and top shots of a single product from one reference photo. Apply: Drop in a product image and prompt 'make a product sheet with front, back, and top shot for the product from image one' to create a locked, multi-angle product asset.
Character sheet prompting — A two-panel image format (full body left, close-up face right) used to lock a character's appearance before any video generation. Apply: Generate the sheet at 16:9/2K with prompt enhancer on across dozens of batches, then select the best result as the reference for every later scene.
Cinematic-look prompt elements (anamorphic lens, shallow depth of field, film grain) — A set of prompt terms the video says are needed to get a wide cinematic field and remove the 'overly clean digital look.'. Apply: Include phrases like 'cinematic anamorphic lens, shallow depth of field, film grain' in location and scene prompts by default.
Style header consistency — Keeping an identical style header — lighting, lens, acting style, physics — across every shot's prompt so disparate generations read as one continuous film. Apply: Template every scene prompt off the same shared header block and vary only the action and camera description per shot.
Director's-notes iteration loop — The core prompting method of never hand-editing prompt text, but instead watching a generation and giving plain-language directorial notes for Claude to turn into a revised prompt. Apply: After each batch, describe the problem like a director (e.g., 'the shot's too static, do the dolly in') and send that note to Claude for a rewritten prompt.
Shot-by-shot camera directing — A refinement of the notes loop that describes what the camera sees rather than what happens in the scene, enabling explicit per-cut camera control. Apply: When a scene fails, rewrite the note as a camera description (low angle, handheld shake, over-the-shoulder) instead of an action description.
Batch-and-mix editing — Running many generation batches for a complex action, then cutting together the best individual phases from different generations rather than expecting one perfect take. Apply: Run several batches of a complex scene and edit the strongest segment from each generation into one coherent sequence.
Sports narrative arc (hero loses first) — A storytelling rule invoked in the video that a sports-film hero should be knocked down before the comeback win, or the win won't land emotionally. Apply: Structure the commercial's action beats so the protagonist suffers a setback before the payoff win.
Map/scheme-based spatial prompting — A technique for fixing objects 'teleporting' between generations by replacing text-only position descriptions with a visual location scheme (compass, distances, dimensions) as a reference image. Apply: When text prompts repeatedly fail to keep an object in place, mark object positions on a location image, formalize it into a scheme via GPT Image 2.0, and attach that scheme plus a keyframe to the prompt.
Scene splitting — The rule that an overly packed, multi-action scene should be divided into two or more separate prompts/generations instead of forced into one. Apply: When a scene has too many simultaneous actions, split it into sequential prompts so each beat gets its own generation and room to breathe.
Physics-explicit direction — A note-writing technique that specifies physical behavior explicitly (gravity, weight, metal snapping instead of sliding) to fix unrealistic falls, impacts, or material effects. Apply: When a fall or impact looks wrong, add explicit physics language describing how the object or material should behave under force.
Macro close-up technique — Inserting super-close-up shots on small details (a chest plate, wrists, a helmet) to raise the perceived production value of a scene. Apply: Add at least one macro/detail close-up per scene where a prop or costume detail can carry visual impact.
Minimal regeneration / asset reuse strategy — A credit-saving approach of generating only the specific missing shot and stitching it into already-approved footage, while reusing unused batch phases or earlier scenes' shots later. Apply: Before generating a new batch, check whether an earlier unused take already covers the need, and generate only the true gap (e.g., just a slide tackle) to cut into existing footage.
Prompt mutation over rebuild — Taking an already-working prompt from an earlier scene and appending new actions or details rather than writing a new prompt from scratch. Apply: Copy a proven scene prompt, add the new action (e.g., 'step overs, nutmegs'), and regenerate instead of starting over.
Comedic-option sourcing from Claude — Asking Claude to generate multiple comedic pitch options for a beat rather than writing the jokes yourself. Apply: Prompt Claude for several comedic options for a given moment, then pick the strongest one to prompt into the scene.
Pack shot prioritization — Treating the final scene (the 'pack shot') as the one that has to sell the whole ad, since it's the last thing viewers see and remember. Apply: Iterate harder on the closing branded shot than on earlier scenes — matching lighting to scene one, adding camera movement, and building the logo/name reveal with extra effects.
Soul Cinema's credit pricing (8 images per credit) means running large batches for a character sheet is cheap relative to how much locking in the right asset matters downstream, incentivizing over-generating early rather than settling.
The named-element auto-attach system in Colab (tagging assets like 'Adil's top up' or 'Robo player two') removes the need to re-upload reference images for every new prompt, turning asset management into a naming problem rather than a file-handling one.
The 'object teleporting' failure mode reveals a structural limitation of the video model — it has no persistent spatial memory across generations — and the fix (a drawn map/scheme) is a workaround for a model limitation, not a prompting-wording issue.
Requiring armor pieces to already exist as physical objects on the ground at the start of a shot (rather than materializing) shows that physical plausibility has to be established at frame one, since the model won't infer object history.
Reworking an already-successful prompt by appending new actions (mutation) is used as a faster, cheaper alternative to writing a new prompt from scratch for a similar shot.
Framing the finished ad as '3 seconds of 100 tries' reframes AI filmmaking as fundamentally a curation/editing exercise built on volume, not a single-generation craft.
«I figured you needed a huge budget and a real studio to make something like that. Turns out, you just needed one website.»
— 00:08
«By the end of this video, you will have everything you need to start creating professional cinematic commercials from scratch.»
— 00:23
«It's a simple three-step workflow. Step one, the assets. Step two, the cloth skill and the prompting framework. Step three, generating the scenes in Cines 2.0.»
— 00:38
«With a tiny budget, whatever crazy idea you've been carrying around can actually become real.»
— 03:00
«Location is the most important image you will create. If this looks sloppy or plasticky, your videos are ruined no matter how good your prompt is.»
— 05:29
«This is where most people fail. They overload their prompts with unnecessary information, spending their time and credits.»
— 11:20
«You don't fix prompts by hand. You watch, you give notes, and Claude rewrites.»
— 19:10
«Before, I told Claude what happens. Now, I'm telling it what the camera sees.»
— 21:46
«Think about every sports film you've ever seen. The hero never wins right away. He loses first. He gets knocked down and then comes back. That's what makes the win mean something.»
— 25:42
«When words keep failing, stop writing and show the model a picture. A map beats a paragraph.»
— 29:49
«When a scene gets too packed, don't try to force everything into one generation. Break it into smaller parts and give each moment room to breathe. It's very important.»
— 32:32
«The truth nobody tells you, the final film is just the best 3 seconds of 100 tries cut together.»
— 45:36
«Iteration is the skill.»
— 45:43
The video is a dense, tool-specific production walkthrough — part tutorial, part case study, part product showcase for Higgsfield AI's own Colab/Soul Cinema/Cines pipeline alongside Claude — that surfaces genuinely transferable prompting and iteration techniques (director's notes, spatial scheme prompting, scene splitting, physics-explicit direction) inside an otherwise promotional frame.

46:05