Lore

AI filmmaking

How I Built a Car Commercial With AI (Every Prompt I Used)

Building a full AI-generated car commercial is presented as a repeatable three-step workflow (assets, setup, generations) that works for any product category, and the video argues that quality comes less from any single AI model than from disciplined asset organization, continuity control, and iterative, actor-directing-style prompting across multiple models.

Higgsfield AI · 2026-08-03 · English

Key ideas

  1. A full commercial can be built with AI using just three steps: assets, setup, and generations, and the same workflow is claimed to transfer to any product (sneakers, perfume, cars) by swapping the assets.

  2. Rigorous nested folder organization (scene → asset type → iteration) is necessary because a project can reach a few hundred generations and otherwise becomes unmanageable.

  3. No single model should be asked to do everything; different models are used for different strengths within the same shot (e.g., one model acts the scene, another handles expression, then a face-swap merges them).

  4. Continuity — spatial position, lighting, geography — is entirely the creator's responsibility; the model does not track it across separate generations.

  5. Generating a video sequence instead of separate stills preserves lighting and detail consistency across frames, since consecutive frames of one video inherently match while separate image generations do not.

  6. 'Clean your plate': reference images should be edited to remove clutter, extra text, and stray objects before being used to generate video, since video models amplify these flaws.

  7. A brief beat of character hesitation or doubt is claimed to be what makes a generated performance read as human rather than as an NPC.

  8. Rule-governed elements like traffic lights, multi-lane traffic, and parking behavior are presented as current failure points for video models, and wide/aerial shots expose more of these failures at once than tight shots.

  9. Complex multi-beat scenes are broken into smaller sub-scenes (e.g., arrival, dialogue, exit) generated and directed separately, then stitched together.

  10. Blocking is locked using tools like a position-log screenshot, a hand-drawn diagram, or colored position markers uploaded to Claude, to stop characters, cars, or objects from teleporting between shots.

  11. Directing the AI is framed like directing real actors: multiple takes are shot for the same beat, with incremental notes (breath, emotion, physicality) added pass by pass.

  12. Final shots can be composited from multiple separate generations, taking the best-performing lead actor from one take and the best-performing supporting actor from another.

  13. All prompts and assets used in the video, plus a free custom 'Cinesoul 2.0' skill, are shared in the video description.

  14. Three-Step Workflow (Assets → Setup → Generations) — The overall production framework taught in the video, breaking any AI commercial into three phases — assets, setup, and generations — claimed to apply to any product type. Apply: Build and organize all reference assets first, configure the project setup, then run generations, swapping in different product assets to reuse the same pipeline for other categories like sneakers or perfume.

  15. Hierarchical Project Folder Structure — An organizational system with a main folder per scene, subfolders per asset type, and a fresh subfolder for every generation iteration. Apply: Set up nested folders (scene → asset type → iteration) before generating anything, since it becomes the only way to locate specific outputs once a project reaches a few hundred generations.

  16. @-Prefixed Asset Elements — A Higgsfield naming convention where saved reference images are stored as 'elements' with names prefixed by the @ symbol (e.g., @car sheet). Apply: Save each finalized asset image as a named @element in Higgsfield so it can be referenced consistently across future prompts instead of re-describing it each time.

  17. Claude-Assisted Element Auto-Insertion — A workflow where a reference image and its @ element name are given to Claude so Claude automatically writes that element name into future prompts it drafts. Apply: Attach the same reference image to Claude and tell it the element's name; when Claude's generated prompt is pasted back into Higgsfield, the element attaches itself automatically.

  18. Three-Panel Character Sheet — A reference format for a lead character consisting of a close-up, a full-body front view, and a full-body rear view. Apply: Generate all three panels for a new lead character before using them across scenes so downstream generations have consistent close-up/front/rear reference.

  19. Multi-Model Testing/Selection — The practice of running the same prompt through multiple AI models to identify which produces the best result for a given asset or shot. Apply: Test candidate models (e.g., SeaArt Dream Pro 5.0, Niji Journey Pro, Soul Cinema) on the same prompt and pick the winner for that specific task instead of assuming one model handles everything.

  20. Multi-Model Pipeline + Face-Swap — A technique of splitting a single character shot across models — one model handles the scene/acting, a different one the detailed facial expression — then merging them via face-swap editing. Apply: Let one model act out the scene, generate detailed expression separately if needed, then use face-swap editing to combine the desired face into the acted scene.

  21. Outfit/State Variation Generation — Generating alternate outfits or physical states of an established character, e.g. for disguise or narrative state changes. Apply: Reuse the locked character asset and prompt for outfit or costume variations needed later in the story rather than creating a new character from scratch.

  22. Color Transfer / Palette Locking — A technique for locking a scene or location's visual palette after the first successful generation and applying that palette to every subsequent image in that setting. Apply: Once a color grade/look is approved, use color transfer (e.g., in Soul Cinema) to apply it consistently across all later shots and locations in the same scene.

  23. "Clean Your Plate" — A pre-generation editing step, named after the VFX term 'plate' for a background image, where clutter, extra text, random objects, and extra cars are removed from a still before it is used to generate video. Apply: Edit out anything a video model is likely to distort — signage text, background clutter, stray vehicles — from the source image before submitting it for video generation.

  24. Position Log — Using a screenshot of a parked car (or other element) taken from an already-generated video as a fixed spatial reference for all later shots in that location. Apply: Pull a frame from an approved video showing correct object placement and reuse it as the positional reference image for every subsequent interior/exterior generation in that scene.

  25. Video-Sequence Generation for Frame Consistency — Generating a short video clip instead of a series of separate still images, since consecutive frames of one video inherently match each other in lighting and detail, unlike independently generated stills. Apply: When several look-alike shots are needed, generate one continuous video and extract frames from it instead of generating each still separately.

  26. Frame Extraction from Video — Pausing a generated video and screenshotting individual frames to use as standalone image assets. Apply: After generating a consistent video sequence, pause on the desired moment and capture the frame as a usable still asset.

  27. Prompt Splitting — Asking Claude to divide an overly complex single prompt into two or more separate, more focused prompts. Apply: When a shot description tries to do too much at once and results degrade, have Claude split it into simpler, sequential prompts instead.

  28. Transition Scene Generation — Creating intermediate shots between two major scenes specifically to preserve geographic and spatial continuity. Apply: Insert an extra generated shot bridging two locations whenever a cut between major scenes would otherwise break geographic logic.

  29. Iterative Casting — Repeatedly regenerating a character until a version's face and energy convincingly reads as the intended lead. Apply: Reject and regenerate character takes (about 3 iterations in the video) until the performance and appearance match the intended character, rather than accepting the first result.

  30. 3/4 Angle Establishing Framing — A camera angle convention treated as critical for establishing a new location clearly. Apply: Use a 3/4 angle shot when generating or selecting the establishing image for a new location.

  31. Multi-Shot Burst + Selective Editing — Generating five to seven unique shot variations per scene beat and then hand-picking and patching together the best individual takes. Apply: Generate a burst of shot options per beat instead of a single attempt, then assemble the final sequence by cherry-picking the strongest take from each burst.

  32. Tighter-Frame Rule — The claim that tighter camera framing produces less visible AI artifacting/error ('slop') than wide or aerial shots. Apply: Favor closer, tighter framing over wide or aerial shots, especially for scenes with rule-governed elements (traffic, multiple actors) the model tends to get wrong.

  33. Asset Library Reuse — Extracting elements that worked well (e.g., a steering wheel, furniture) from earlier generations and re-embedding them as named elements in later prompts. Apply: Screenshot successful sub-elements from prior outputs and save/attach them as elements for reuse in follow-up prompts rather than regenerating them from scratch.

  34. Hard-Rule Prompting — Prompting with explicit numeric/structural constraints — exact shot count, cut count, and forbidden insert types — to control output structure. Apply: Specify hard limits in the prompt (e.g., 'exactly five shots and four cuts, no extra inserts') to prevent the model from adding unwanted extra shots or elements.

  35. Location Cleanup via Image Editing — Using an image editor to erase unwanted structural elements (duplicate buildings, extra booth rows, extraneous details) from a location image before generation. Apply: Manually erase problematic duplicated or extraneous elements from a location still before it's used as a generation reference.

  36. Pre-Rendered Asset Insertion — Generating a needed prop or detail asset (e.g., a newspaper headline, pre-blurred) separately and dropping the finished asset into the video prompt. Apply: Create fine-detail props as standalone pre-processed images rather than relying on the video model to render them correctly in-scene.

  37. Diagram-Based Positional Prompting — Drawing a simple diagram of a desired path or parking layout by hand and uploading it as an image reference to Claude to lock vehicle trajectory. Apply: When textual position descriptions fail to control blocking, hand-draw the intended layout or path and upload it as a reference image alongside the prompt.

  38. Position Marker Annotation — Scribbling colored marks (e.g., red, blue) directly onto a location image to indicate exact stop points or trajectories. Apply: Annotate a location still with colored markers denoting precise stop or trajectory points before using it to constrain a generation.

  39. Scene Deconstruction (Sub-Scene Splitting) — Breaking one complex multi-beat scene (arrival, dialogue, exit) into labeled sub-scenes, e.g. 5A/5B/5C, generated and directed separately. Apply: Split a scene with complex geography into ordered sub-parts, generate each separately, and let the dialogue-critical sub-scene constrain the blocking of the others.

  40. Dialogue-First Shooting — The practice of locking dialogue and blocking in the core conversation shot of a scene before shooting the arrival or exit shots that must match it geographically. Apply: Generate and finalize the dialogue-heavy middle shot first, then constrain the arrival and exit shots to match the car's position and surroundings established there.

  41. Cross-Generational Compositing — Taking the best-performing actor from one full generation and combining them with the best-performing supporting actor from a different generation of the same shot. Apply: When no single generation nails every actor's performance, pull the strongest character performance from each of several separate takes and composite them into one final shot.

  42. NPC Character Sheet — A reference sheet for background/non-lead characters specifying age range, appearance, uniform, and lighting directives. Apply: Before generating scenes with recurring background characters (e.g., toll controllers), define their age range, look, uniform, and lighting in a dedicated sheet for consistency.

  43. Iterative Constraint Refinement Directing — A directing method where each successive generation pass adds a more specific directive (emotion, breath, body physics) on top of a broad initial prompt, likened to giving notes across multiple takes. Apply: Start with a broad prompt for a shot, then regenerate multiple times, adding one specific refinement (e.g., a breath, an emotional beat) per pass until the performance reads correctly.

  44. Cinesoul 2.0 Skill — A custom, free-to-download 'skill' (workflow/prompt toolkit) the creator built for this pipeline and shares for others to use. Apply: Download the free Cinesoul 2.0 skill from the video description to replicate the creator's workflow directly.

Insights

The @-prefixed element system combined with Claude auto-inserting element names into prompts effectively turns visual assets into reusable variables, removing the need to redescribe them in every new prompt.

Shot scale is treated as a way to hide model weaknesses, not just an aesthetic choice: wide/aerial shots expose more simultaneous rule-governed elements (multiple traffic lights, lanes) for the model to get wrong, so tighter framing is used deliberately to reduce visible errors.

A generated video frame is treated as more reliable continuity ground truth than the original static asset — the workflow screenshots an approved video frame ('position log') and uses it as the spatial reference for all later shots in that location.

Prompt-splitting and scene deconstruction (e.g., 5A/5B/5C) are presented not just as workarounds for prompt length but as the direct fix for compounding physics and composition errors that appear once a single prompt is asked to do too much.

Uploading a hand-drawn diagram of a desired parking or driving path is used to bypass the failure of purely textual spatial-position language, suggesting that visual/spatial instructions outperform verbal descriptions for blocking control.

Iterative prompting is explicitly reframed as directing a performance rather than debugging a technical output — each regeneration pass adds one specific human-performance note (a breath, a beat of doubt, an emotional shift) rather than a purely visual correction.

«The whole workflow is just three steps: assets, setup, and generations. Today, it's a car, but the same workflow works for pretty much any commercial. Sneakers, perfume, whatever you're selling. You just swap the assets and that's it.»

— 01:43

«A few hundred generations in, it's the only way to find anything.»

— 02:27

«don't try to get everything out of one model. I let Soul act the scene, then swap my face in after.»

— 05:46

«In VFX, the background image is called a plate. So, clean your plate. Anything a video model can break, multiple text, background clutter, random cars. Edit them out of the image before you generate a single video.»

— 12:18

«In AI filmmaking, continuity is on you. The model won't keep track of it.»

— 13:51

«That one second of doubt is what makes you feel like a character and not just an NPC.»

— 16:18

«He takes the glasses off like he's the face of a luxury eyewear brand. The flares streaks across the frame and the eyes come up half closed.»

— 18:13

«If you remember one rule from this video, make it this one. The tighter the frame, the less slop you get.»

— 21:01

«If you have multiple people or objects in the video and they keep teleporting from shot to shot, start with one that can clearly define who should be where.»

— 21:54

«This is the part of AI filmmaking you rarely get to see, directing actors that don't exist.»

— 24:22

«Let's try something instead. I drew a simple diagram of where I want the car parked, and I'll upload that straight into Claude.»

— 29:28

«The dialogue is the heart of the scene. It decides where the car sits and what's outside every window.»

— 30:36

«That's the full workflow. But remember, that works for any product. Can be perfume, sneakers, cars in our case.»

— 35:20

Reception

Audience is overwhelmingly impressed by the technical quality and realism of the AI-generated content, with strong engagement and requests for tutorials, though minor skepticism and copyright concerns appear.

A dense, well-organized production diary that functions as a genuinely reusable named-technique framework for AI-assisted commercial filmmaking — heavy on concrete workflow detail (folder structure, model-switching, continuity fixes, iterative directing) rather than a simple highlight reel of finished shots.

35:47

↳ Higgsfield AI · YouTube

Watch original