Lore

AI filmmaking

How to Make AI Films That Don't Look AI (Full Tutorial)

The video is a day-4, unedited field diary of Higgsfield's 15-person team racing to finish the first fully AI-generated feature film (80 minutes, roughly 10 million credits, 14-day deadline) in time for a Cannes premiere, and it uses that day's work on scenes 21 and 23 to teach the concrete prompt-engineering and production frameworks — custom Claude skills, style prefixes, spatial layout blocks, batching, emotional prompting — the team uses to turn failed AI generations into usable cinematic shots.

Higgsfield AI · 2026-05-09 · English

Key ideas

  1. Project scope: first fully AI-generated feature film, 80 minutes, premiering at Cannes in under a month, built with ~10 million credits, a 15-person team, and a 14-day deadline; the video shows day 4 unedited rather than the presenter's usual polished first-try footage.

  2. Today's assigned work: the presenter generates scenes 21 and 23 while a coworker handles scene 22; whatever gets made must clear director approval to make the final film.

  3. Scenes 21/23 are emotional flashback/reminiscence scenes after the character Lulu is taken away, and must convey that Roco's bond with her is family, not just friendship.

  4. A custom skill ('shotlist-builder'/'short list builder') is uploaded into Claude and, given the script, splits target scenes into a shot list broken into roughly 15-second shot chunks with prompts tailored to the video model.

  5. A shared Canvas document holds all assets (locations, characters, props) plus a 'style prefix' controlling lighting, color, composition, and audio, prepended to every prompt to keep 15 people's output consistent.

  6. Style prefix specifies natural light only from sky/windows, no music, no subtitles, environmental SFX only — because the video model (Cannes 2.0/Cedense 2.0) bakes a single audio track and tends to add unwanted music that can't be separated in post.

  7. The image model (Nano Banana Pro) struggles with spatial awareness and tends to produce plasticky textures; adding a 'light atmospheric haze' keyword measurably improves lighting and texture.

  8. Location reference image quality is treated as critical, since most of a generated video's lighting and style derive from it.

  9. Claude Collab is used to organize generation work by scene, manage sharing/permissions, and keep every asset, prompt, and tool-provenance record visible across the team.

  10. Repeatedly re-editing the same image degrades the model's grasp of spatial layout; the fix is to download a preferred version and restart edits from that new base rather than iterating on one file indefinitely.

  11. Locking the apartment's layout took 44 iterations; an additional ~20 were estimated for a hallway reverse-angle view, with fridge and mirror placement key to fixing spatial orientation.

  12. Stacking two different-angle reference images (front-facing wide shot + hallway view) into a single composite image helps give the model spatial context.

  13. Uploading character sheets without tagging them by name caused the model to invent new characters instead of using the provided faces; tagging fixed this.

  14. Generating all Polaroids on one photo-wall prompt simultaneously caused quality loss and facial drift; the team switched to generating one image at a time and compositing them in Photoshop.

  15. Cedense 2.0 originally imposed a hard 3,000-character prompt limit; since Chinese uses roughly 1-2 characters per concept versus 5-10 for English, the team writes prompts in Chinese to fit more meaning within the cap.

  16. Guest VFX artist Patrick Kalin (credited on Avatar, Dune, Blade Runner 2049, Deadpool 2) reviews a prior Higgsfield AI film ('Hellgrind'), calls the character work 'the illusion of life,' and estimates the equivalent traditional VFX/production budget at $15-20 million.

  17. The model struggles with spatial awareness specifically when a character enters a room from outside.

  18. A mid-production script change to a flashback scene's location/dialogue forced scrapping roughly 2 minutes of already-generated footage out of a 2.5-minute scene that had already cost nearly 100,000 credits.

  19. Only 1 of 64 generations for one shot, and only 1 of 100 total attempts, made the final cut.

  20. Cost comparison cited: the director of 'The Mask' reportedly estimated a minimum traditional-film budget of $5 million for an equivalent project, versus this AI production's roughly $70,000.

  21. Typical AI film productions reportedly use teams of 3-4 people; this project deliberately uses 15, which the presenter frames as part of why it needs the full 14 days ('creativity does need time').

  22. A 'spatial layout block' — a new prompt section explicitly describing room layout and exact camera position — is introduced and is said to measurably improve camera positioning accuracy.

  23. Splitting a shot's motion into second-by-second sub-phases forces unwanted extra cuts in the output; describing continuous motion/framing instead works better.

  24. Emotional-state prompting (naming the character's feeling, e.g. heartbroken, nostalgic) is said to yield better character performance than literally describing the physical action.

  25. Failed generations are not charged credits.

  26. The video model can under-deliver on requested frame rate (24 fps requested, ~12 fps observed in one case); removing duplicate frames fixes the stutter but destroys the shot's audio and costs a few seconds of runtime.

  27. The model struggles to place props absent from its reference images, illustrated by repeated failures placing a Polaroid on a fridge.

  28. Overlong, heavily-edited prompts are said to 'overwhelm' the model; the fix demonstrated is scrapping the prompt and rewriting it from scratch rather than continuing to patch it.

  29. GPT Image (2.0) is said to outperform the team's usual image tool at non-destructive edits to an existing image.

  30. An unscripted camera pullback generated by the AI is kept and retroactively interpreted by the presenter as a deliberate cinematographic choice — 'pulling back from memories to let them settle' — and tied to the standard technique of retreating the camera to give a character emotional privacy.

  31. End-of-day-4 tally: the team spent 4,441,352 credits (~$260,000) over 4 days, generated 48,336 images/videos and roughly 800 total assets, of which only 8 made the final cut, prompting the presenter to question whether the 10-million-credit budget will be sufficient.

  32. Episode 3 is promised to show the first full draft of the finished film.

  33. Shotlist-builder skill — A custom skill uploaded into Claude (via Customize → Create New Skill → drag-and-drop zip) that, given a script, breaks target scenes into a shot list split into roughly 15-second chunks with prompts tailored to the video model, then auto-requests needed character/location/prop assets. Apply: Feed the script into the skill, ask for a shot list for specific target scenes, and let it request the asset references it needs before generating per-shot prompts.

  34. Style prefix — A standing block of text, kept in a shared Canvas asset document, that fixes lighting, color, composition, and audio rules (e.g., natural light only, no music, no subtitles) and is prepended to every shot prompt. Apply: Write the visual/audio constraints once in a shared document and prefix every team member's prompts with it so a large crew's output stays consistent.

  35. Claude Collab — A collaboration workspace for organizing generation work by scene, sharing projects with configurable roles/permissions, and keeping every asset, prompt, and tool-provenance record visible to the whole team. Apply: Create a project per scene, invite collaborators with the appropriate role, and use it as the single source of truth for which prompts produced which generations.

  36. Spatial layout block — A dedicated new prompt section explicitly describing a room's layout and the camera's exact position, added after repeated camera-placement failures. Apply: Append a short paragraph naming the room layout and stating precisely where the camera sits whenever a shot's camera position keeps coming out wrong.

  37. Reference anchoring — Describing camera position relative to a named, already-established set piece (e.g., 'next to the fridge') to stop the model from moving furniture toward the camera instead of moving the camera itself. Apply: Anchor camera-position language to a specific object already visible in the reference images rather than giving an abstract distance or direction.

  38. Emotional prompting — Prompting a character's emotional state (e.g., heartbroken, nostalgic) rather than literally scripting the physical action, on the claim that it produces stronger performance. Apply: Replace action-blocking language with an emotional cue and let the model infer the physical performance from it.

  39. Batch generation — Generating several variants of the same shot at once (batches of 4, 8, or more) and culling for the best result instead of trying to perfect a single generation through repeated prompt tweaks alone. Apply: Run a batch, skim the first few frames of each result, and if multiple outputs share the same flaw, stop and fix the prompt in Claude rather than continuing to batch blindly.

  40. Multi-view reference composition — Stacking two different camera-angle images of the same location (e.g., a wide front-facing shot and a hallway view) into one composite reference image to give the model more spatial context. Apply: Combine multiple angle references of a set into a single composite image before uploading it for a room the model needs to reconstruct from more than one viewpoint.

  41. Character sheet tagging — Marking up character reference sheets with the character's name so the model uses the provided face rather than inventing a new character, after untagged uploads caused the model to fabricate faces. Apply: Label each character reference image with its character name before uploading it for a shot needing that character's face to persist.

  42. Reference image substitution for continuity — Swapping in an updated version of a reference image mid-sequence (e.g., a fridge shown without a Polaroid once the character has already taken it) so later shots stay continuous with earlier story events. Apply: After a prop's state changes within the story, generate and upload a new reference image reflecting that state for every subsequent shot.

  43. Frame-by-frame duplicate-frame removal — A manual fix for video output rendering at a lower actual frame rate than requested (24 fps asked, ~12 fps delivered): scrubbing the timeline to find and cut doubled frames. Apply: When output looks stuttery relative to the requested fps, scrub frame-by-frame to remove duplicated frames, accepting that the fix strips the shot's audio and shortens its runtime.

  44. Director's call — A named decision point where the human director consciously keeps a generation (including an AI-originated creative choice) that departs from the written script because it works better. Apply: When a generation deviates from the script in a way that improves the scene, explicitly approve it as a director's call rather than forcing a re-generation that matches the script exactly.

  45. Edit-as-you-go — Placing accepted shots into the timeline and rough-trimming them while the rest of the scene is still being generated, instead of waiting until all shots exist to start editing. Apply: Drop each accepted generation into the rough cut immediately so you can see how shots intercut and spot what's still missing.

  46. "Light atmospheric haze" keyword — A specific prompt keyword found to meaningfully improve lighting quality and counter the plasticky textures the image model tends to produce. Apply: Add 'light atmospheric haze' (or similar atmospheric-lighting language) to image prompts when outputs look flat or plasticky.

  47. Long-prompt reset — Scrapping an over-elaborated, repeatedly-edited prompt and rewriting it from scratch once its length starts 'overwhelming' the model, rather than continuing to patch it. Apply: When a heavily-edited prompt stops improving results, delete it and write a fresh, simpler version instead of adding another patch.

Insights

The video reframes near-total generation waste as pedagogically productive: with roughly 1% of generated assets (8 of ~800) making the final cut, the presenter explicitly argues every rejected batch 'taught' the prompt for the next attempt rather than being wasted spend.

Across the whole session the recurring bottleneck is geometric/spatial consistency (room layout, hallway orientation, camera position, prop placement) rather than motion or emotional performance, which the presenter repeatedly describes as comparatively easy for the model to get right.

Because the video model bakes one single audio track per generation, any AI-added music becomes essentially unfixable without discarding the whole audio track — forcing the team to suppress music entirely at the prompt level rather than mix it out later.

Emotional-state prompting (naming a feeling rather than blocking an action) is presented as a specific technique on the premise that these models model performance/affect more reliably when given a motivation than an explicit physical instruction, similar to actor direction.

Prompt density is language-dependent in a way that has real economic consequences under a hard character cap: the team's use of Chinese prompts over English is a direct workaround for Cedense 2.0's 3,000-character prompt limit, since Chinese needs far fewer characters per concept.

An AI-originated creative deviation (the unscripted camera pullback) is accepted and narrativized after the fact by the human director, illustrating how authorship in this workflow can shift from prompted intent to curated AI improvisation.

The team scaled team size (15 vs. the stated AI-production norm of 3-4) as a deliberate strategy to compress a feature-length pipeline into 14 days, treating headcount rather than tool speed as the main lever against the deadline.

«We're presenting a full feature film in can in less than a month. 80 minutes long, fully AI generated.»

— 00:00

«I'm going to teach you the exact frameworks I use to turn those failures into cinematic shots.»

— 01:03

«We have to convey to the viewers that they aren't just friends, they're family.»

— 02:49

«Natural light only, uh key light coming from sky and windows only.»

— 04:03

«This little line in the bottom over here makes a big difference in terms of lighting and textures.»

— 11:56

«It's easier to just batch a couple images. And if there are any parts I don't like, I'm just going to ask Claude to fix those.»

— 15:04

«This is an incredible piece of work. Um, I was blown away… there were times when I was watching it and I completely forgot that I was watching something that was like generated with AI… that's got to be the highest compliment.»

— 32:19

«you'd be looking at like 15 to 20 million US for something like [this]»

— 33:09

«I prefer doing batches of eight and then seeing if it's doing what I wanted to do. Uh you'll notice right away if a few of your videos are making the same mistake, uh that's when you stop them and head over to Claude and fix your prompts.»

— 48:56

«that's how you get greatness. That's how you get cinema.»

— 58:30

«Creativity does need time.»

— 76:10

«The spatial layout block definitely worked.»

— 76:49

«if you're new here, this is like 90% of AI film making. You just have one little issue and spend 15 batches trying to fix it.»

— 96:07

«I'm not describing the exact actions, but rather the emotions that they should be feeling uh when they perform those actions.»

— 99:09

«these aren't waste. Like you've noticed, every single one of these taught me the prompts for the next one. And keep iterating and keep iterating and eventually we're going to get the shot.»

— 107:55

«if your images look sloppy or the lighting is off, Cance has a tendency to grab those imperfections and add them to the video. So, we really want to make sure that the images are perfect.»

— 119:43

«As you can see, each frame feels cut or like delayed. Uh, this is happening because in the prompt we asked for 24 fps. So, 24 frames per second. But in this case, it's lower. So, like maybe 12 fps.»

— 123:42

«I think it's getting really long and it's just overwhelming the model.»

— 143:14

«I'm praying to the prompt gods at this point.»

— 145:05

«This is my workflow for pretty much any shots. Uh the prompts are really long and you can't really tell. Well, especially cuz it's Chinese. Uh but even with English prompts, you can't really read through all of it and understand exactly what the output is going to look like. Uh so I recommend running a couple batches, uh testing it, and then you can go from there.»

— 169:29

«Claude added this on its own, but I really like it. It's like we're pulling back away from these memories and letting them settle and land, if that makes sense.»

— 174:56

«this is a pretty common technique in cinematography where it's like you're leaving the character alone, giving them some space to really feel their emotions.»

— 179:51

«During the shooting, I generated nearly 800 assets and only eight of them made the final cut. Now, it took a lot of effort.»

— 185:24

«Now, by episode 3, we'll have already finished the first draft of the full film, and we'll have a lot to share with you guys. But until then, stay tuned.»

— 185:38

Reception

Audience appreciated the extensive effort and educational value of the walkthrough, with constructive feedback on technical consistency issues but strong interest in future content.

The video is a genuinely technique-dense production log — its named frameworks (style prefixes, spatial layout blocks, emotional prompting, batching, reference anchoring) are concrete and reproducible — but its central 'failures become frameworks' pitch is demonstrated through the presenter's own selected wins rather than across the roughly 99% of generations that didn't make the final cut.

185:51

↳ Higgsfield AI · YouTube

Watch original