Lore

AI filmmaking

How to Make Ultra Realistic AI Videos (28 Best Tips)

Higgsfield AI's production team documents building an 80–90 minute fully AI-generated feature film in a 14-day, 15-person, ~10-million-credit sprint for a Cannes debut, and distills that process into roughly 28 concrete prompting, pre-production, teamwork, and editing techniques for making AI-generated video look consistent and realistic.

Higgsfield AI · 2026-05-16 · English

Key ideas

  1. A 15-person team is producing an 80-minute fully AI-generated feature film in under a month for Cannes, working with a ~10 million credit budget over a 14-day sprint.

  2. By day 8 of 14 the team has a full draft (110 minutes) that must be cut to 90 minutes, with 40 scenes still needing fixes; three are used as on-camera case studies.

  3. The episode walks through fixing an early scene, identifying sloppy/plasticky imagery, a cheap-looking portal effect, and inconsistent character positioning across cuts as the recurring failure modes.

  4. The team says it discovered 28 distinct techniques across 8 different generation types, spanning prompting, image generation, pre-production asset prep, teamwork structure, and editing.

  5. A team-shared system prompt is presented as encoding seed/safe-language rules, on-screen text handling, prompt density (one idea + one action + one camera move per shot), reference-image discipline, generation-mode selection, and an X/Y frame-coordinate system for placing characters — claimed to 'save months of work.'

  6. Concept artist Jama Jurabaev (credited with work on The Mandalorian, Ready Player One, Avengers: Age of Ultron, Jurassic World, and 10+ years at Lucasfilm/ILM/Marvel/Framestore) appears and frames the work as requiring talented people, the right tools, the right time, and motivation.

  7. The finished film consumed 108,859 generations over 14 days (about 7,700/day, ~320/hour), 9,540,147 credits, $400,000 in generation costs, and $500,000 total production cost for a 90-minute feature.

  8. The team claims this is roughly 1% of the ~$50 million a comparable live-action production with the same VFX would cost.

  9. The video frames AI filmmaking's core structural advantage as unlimited prompt-based 'retakes' replacing the single-take finality of traditional filmmaking.

  10. Emotional scene anatomy prompting — Providing situational context plus precise anatomical/body-position detail (exact arm position, jaw, eyes, tears) instead of vague emotion words to eliminate robotic AI reactions. Apply: When prompting an emotional beat, specify exact body position and facial anatomy (e.g., 'right arm fully extended, fingers spread, mouth open, jaw dropped, teeth visible, no tears') rather than just naming the emotion.

  11. Photographer-style image prompting — Prompting image generation using cinematography vocabulary — lens type, aperture, depth of field, grain, film stock, halation, bokeh — instead of generic descriptors like 'photorealistic, cinematic, high quality.'. Apply: Replace vague quality adjectives with concrete camera specs (e.g., '35mm lens, stopped down two stops from wide open, deep focus, cool blue light spills through curtains'), optionally asking Claude to identify camera/lens choices from a reference movie still.

  12. Team system prompt for prompt adherence — A shared, team-approved master system prompt covering seed/safe-language rules, on-screen text handling, prompt density (one shot = one idea + one action + one camera strategy), reference-image discipline, generation-mode selection, and a frame coordinate (X/Y axis) system for character placement. Apply: Build and reuse a single system prompt across the whole team so every generation follows the same conventions, including specifying character placement as percentages on an X/Y axis of the frame.

  13. Iterative regeneration — The observation that first-batch generations almost never work, so getting a usable result requires inspecting what landed and rerunning with targeted adjustments. Apply: After a generation, identify which portion worked (e.g., the first 5 seconds of a 15-second clip) and instruct the model to only change the remaining portion rather than regenerating from scratch.

  14. Single shots — A generation approach defining what happens at each second within one continuous camera/character motion, used for most non-action, non-dialogue content. Apply: For simple continuous action, script the shot second-by-second within a single continuous camera move to maximize control.

  15. Multi shots — A single generation that holds multiple cuts while keeping lighting and character positioning consistent, used for action and dialogue scenes. Apply: Generate several cuts in one pass for a dialogue or action sequence, then cut out only the successful phases instead of generating each cut as a separate clip.

  16. Phonetic dialogue spelling — Writing out how complex or invented words sound phonetically instead of using their standard spelling. Apply: Replace invented or unusual proper nouns/words in dialogue prompts with a phonetic spelling so the model pronounces them correctly.

  17. Cut/angle prompt precision — The claim that giving more detail on exactly when to cut, which angle to use, and who to cut to reduces hallucination in multi-shot generations. Apply: Specify precise cut timing, camera angle, and subject for every cut in a multi-shot prompt rather than leaving it implicit.

  18. Reference frames — Screenshotting a liked moment from a generation and uploading it to Claude as a reference for positioning and lighting in follow-up prompts. Apply: Capture a successful frame, feed it to Claude, and use it as a reference image to keep later generations consistent with that positioning and lighting.

  19. Shot list optimization — Instructing Claude to update only the changed prompt within a long shot list rather than rewriting the entire list, claimed to cut token usage by roughly 80%. Apply: Ask Claude to 'optimize your work with a shot list and update only what changed' whenever revising one shot within a large shot list.

  20. Prompt saturation cleanup — When a prompt grows too long from repeated editing, having Claude study the context and sanitize it by removing conflicting or unnecessary elements. Apply: Periodically have Claude 'optimize, study context, sanitize' an over-grown prompt to remove clutter before continuing to iterate on it.

  21. 60-30-10 color rule — A cinematography color-balance rule of 60% dominant color, 30% secondary color, and 10% accent color, described as the most visually appealing split. Apply: Compose a scene's palette around a 60/30/10 split (e.g., a forest scene at 60% green/blue, 30% red, 10% white) when prompting for image or video color.

  22. Practical lighting only — Restricting light sources in a prompt to practical (diegetic) sources visible or implied in-frame — sun, windows, lamps — to avoid plasticky, 'plopped-in' looking characters. Apply: Specify 'practical light only' and name the actual in-scene light sources instead of letting the model invent artificial extra lighting.

  23. Establishing shots for spatial consistency — Starting a scene or sequence with a wider or medium establishing shot showing all characters and the location so the model maintains correct positioning across subsequent cuts. Apply: Always generate an establishing shot first for dialogue and multi-cut sequences; the model uses it to keep character placement consistent as you move into closer cuts.

  24. Anchor points for character placement — Assigning a distinct visual reference object in a location (a tree, books, a fixture) that generations can use as an anchor for precisely placing characters relative to the environment. Apply: Designate a visual anchor for every location and reference it in prompts when specifying where a character should stand.

  25. Non-frontal location generation angle — The model's reportedly weak depth perception when a location is generated head-on, addressed by generating locations at a three-quarter angle or from a ceiling/CCTV-style angle instead. Apply: Generate location reference images from a three-quarter or overhead CCTV-style angle rather than straight-on to get usable depth.

  26. Multi-view location asset splitting — Because the model reportedly confuses foreground/background views of a location with multiple distinct sides (e.g., front vs. back), generating and referencing separate image assets for each view instead of one composite image. Apply: Create separate front-view and back-view (or other distinct-view) images of a multi-perspective location and reference the correct one per prompt.

  27. Top-down schematic map via Claude — Generating a top-down schematic map of a scene through Claude before writing generative prompts, iterating character/object positions verbally, then copying the resulting prompt prefix into the video model. Apply: Use Claude to draft and refine a top-down blocking map of a scene as text first, then paste that positional description as a prefix into the video-generation prompt.

  28. Beat-by-beat action scripting — Writing action scenes as a beat-by-beat script specifying participants, positions, weapons, and choreography, delegating the writing to Claude if needed. Apply: Before generating an action scene, script out each beat (who, where, what weapon, what movement) explicitly rather than describing the action loosely.

  29. Face extraction from character sheet — Extracting a character's face from a wide shot on the character sheet and substituting it in place of a directly generated close-up, to avoid the plasticky look of close-up generations. Apply: When a generated close-up looks plasticky, take the face from an existing wide character-sheet shot instead of generating a fresh close-up.

  30. Multi-view prop sheets — Creating prop reference sheets with multiple views (e.g., a monster claw's interior and exterior) for ambiguous or complex props. Apply: For any prop whose appearance is ambiguous from a single angle, generate a multi-view reference sheet before using it in shot prompts.

  31. Script supervision / continuity tracking — Tracking each character's state across a long-form production — injuries, clothing, props, missing items — and locking that state before generation, akin to a script supervisor's job. Apply: Maintain a per-character continuity log of current injuries, clothing, and props, and check it before generating any new scene involving that character.

  32. Minimum-two-person teams — The claim that a single long-term solo contributor becomes blind to AI artifacts ('slop'), so generation work should be organized in teams of at least two to distribute load and cross-check. Apply: Assign at least two people to any generation workstream rather than one, to prevent bottlenecks and artifact-blindness.

  33. Slop director — Appointing one dedicated person whose sole job is to review and veto AI visual defects ('slop') before any generation is locked into the timeline, because a dedicated reviewer reportedly catches what the full team misses. Apply: Designate one team member as the 'slop director' to review every generation before it ships to the timeline.

  34. Shared canvas asset library — A centralized shared canvas (aided by a 'Hexi' assistant for optimization) holding the asset library — characters in all their states, all locations — so any team member can grab correct, up-to-date assets. Apply: Maintain one shared canvas of characters (in every state) and locations that the whole team pulls from when generating scenes, rather than scattered individual asset copies.

  35. Frame-by-frame review — Reviewing generated video frame by frame because, unlike filmed footage, a single bad frame can ruin an entire generated shot. Apply: Scrub generated clips frame by frame before accepting them, since generation artifacts can appear in individual frames.

  36. Duplicate-frame deletion for FPS correction — When a generation stutters and drops below the target frame rate (e.g., 18 FPS instead of a specified 24 FPS), manually deleting the duplicated frames to restore the correct rate. Apply: If a clip stutters below your target FPS, go frame by frame and delete duplicate frames to bring it back to spec.

  37. Salvaging failed generations — The principle that a failed generation is not wasted — 1–2 seconds of it may be usable — so those segments should be cut out and patched into the timeline instead of discarding the whole clip. Apply: Review failed generations for short usable segments (1–2 seconds) and splice those into the edit rather than re-rendering the whole shot.

  38. Regeneration pipeline (move on and revisit) — A workflow where a stuck scene is temporarily placed on the timeline as-is so work can continue on other scenes, then revisited later with fresh eyes to improve the prompt — contrasted with traditional filmmaking's single-take finality. Apply: If stuck on a scene, drop the current best version into the timeline, move on to other scenes, and return to iterate on the prompt later rather than blocking progress.

Insights

Vague quality descriptors like 'photorealistic, cinematic, high quality' are explicitly said to 'do nothing for your image' — the video argues specificity (lens, aperture, light source) is what actually changes output, not generic quality adjectives.

The generation model is described as having weak depth perception on head-on location shots, so the team generates locations at a three-quarter angle or overhead/CCTV-style angle instead of front-facing.

Multi-perspective locations (e.g., the front vs. back of the same room) reportedly get confused by the model when combined into one image, so the team generates and stores separate image assets per view.

The video claims a single long-term solo contributor goes blind to AI-generation artifacts ('slop'), which is why it recommends institutionalizing a dedicated 'slop director' role rather than relying on the whole team to catch defects.

Failed generations aren't treated as waste: the stated workflow explicitly harvests 1–2 second usable fragments from failed takes and splices them into the final edit rather than discarding the clip.

Token usage is treated as a production cost lever, not just a technical detail — having Claude rewrite only the changed prompt within a long shot list is claimed to cut token usage by 80%.

The team explicitly reframes AI-production economics against live-action VFX budgeting, positioning heavy prompt iteration ('free retakes') as the reason a $500K, 90-minute AI film can substitute for what they estimate as a $50M live-action equivalent.

Practical/diegetic-only lighting (sourcing light exclusively from sun, windows, or lamps actually present in frame) is presented as the specific technical fix for the common 'plopped-in,' artificially lit look on characters.

«We're presenting a full feature film in Cannes in less than a month. 80 minutes long, fully AI generated. Now, the issue is that we don't have the film yet. But, what we do have is 10 million credits, a team of 15 people, 14 days, and no [ __ ] clue if we'll make it.»

— 00:00

«These actually do nothing for your image.»

— 08:51

«The idea is that your input for the location, it's going to split it into X and Y axis, and then tell the model exactly where to place your characters or your objects.»

— 11:27

«To make something cool, you need to have talented people, right tools, right time, and the motivation.»

— 13:00

«Tell Claude to optimize your work with a shot list and update only what changed. This will cut your tokens by 80% and save you tons of time, I guarantee you.»

— 18:00

«But the model understands an establishing shot and then moving forward places everyone correctly.»

— 20:32

«appoint a slop director. So, one person whose job is to catch AI slop before they ship.»

— 25:15

«In traditional filmmaking, if you have a bad take, well, you're pretty much done. There's not much you can do about it. In AI filmmaking, you try again, you change your prompt, you try again, and you get what you're looking for.»

— 27:19

«All right, 28 hacks, roughly half a million dollars worth of work compressed into 14 days.»

— 27:28

«For comparison, a 90-minute live-action with this kind of VFX would cost about 50 million dollars. We did it for 1% of that.»

— 28:14

Reception

Viewers praised the technical tutorial and workflow details, but criticism of the film's quality, AI limitations, and perceived marketing overreach created a mixed reception.

A technique-dense, self-documented case study that pairs a real production's numbers and failure examples with 28 specific, reusable prompting and pipeline tips, making it a strong reference episode despite being a self-reported, single-production account.

28:37

↳ Higgsfield AI · YouTube

Watch original