Lore

AI filmmaking

The Complete AI Short Film Workflow Everyone's Missing

Adele (Higgsfield AI) argues that a complete, emotionally resonant short film can be made entirely with AI tools — no camera, crew, or large budget — if the creator starts from a feeling rather than a prompt, compresses the story into a three-sentence/three-beat foundation, and treats generation as an iterative loop.

Higgsfield AI · 2026-03-13 · English

Key ideas

  1. Start ideation from an emotional core/feeling, not a visual prompt.

  2. Compress the story into a 3-sentence foundation before any production.

  3. Structure the short film around only three beats: setup, midpoint, climax.

  4. Train a personal 'Soul ID' from ~20 varied reference photos so the creator's own face/likeness can be inserted into any scene.

  5. Generate a character sheet from source images to lock visual consistency across shots.

  6. Use Soul Cinema for cinematic-default image generation, claimed to beat Nano Banana Pro at photorealism.

  7. Feed plain-language scene descriptions into Claude to generate stronger prompts for the image/video models.

  8. Carry visual continuity between shots by screenshotting a prior shot's frame and using it as the next shot's start frame.

  9. Vary color grade, camera style, and lens per scene to give each scene its own visual identity.

  10. Use camera choices like handheld shake to signal narrative/emotional tone, not just style.

  11. Break up dialogue scenes with short atmospheric B-roll shots layered under one continuous audio track.

  12. Treat generated shots as revisable indefinitely — missing shots can be generated post-production and matched to the existing aesthetic ('nothing is locked until you lock it').

  13. Clone the creator's own voice via Pixel Audio's 'change voice,' claimed to preserve original emotion, pacing, and delivery while only changing timbre.

  14. Split dialogue into separate tracks per character to voice-clone or assign each character individually; use preset voices for secondary characters.

  15. Frame the whole workflow as: no camera, no crew, no big budget — just the feeling plus an all-in-one AI tool suite.

  16. Iteration (generate, review, tweak the prompt, regenerate) is presented as the core skill, with many discarded 'bad takes' behind the final film.

  17. Closes with a contest call-to-action: submit a 3-sentence concept (character, location, moment) in the comments for a chance at a free 'ultimate plan.'

  18. Feeling-first ideation — Adele's method of starting short-film ideation from an emotional core question rather than a visual prompt. Apply: Before writing any AI prompt, ask 'What's a story with real emotional stakes?' and let the answer drive the concept instead of visual style.

  19. Three-sentence story foundation — Distilling the entire story premise into exactly three sentences before any production begins. Apply: Write the story's foundation in three sentences (e.g., character, situation, stakes) and use it as the anchor for all scene mapping.

  20. Setup/midpoint/climax structure — A three-beat scene structure the video claims is sufficient for a full short film. Apply: Map the short film into exactly three scenes: a setup where something's off, a midpoint that builds tension, and a climax where everything shifts.

  21. Soul ID / face training (HeyGen) — Uploading around 20 varied reference photos into HeyGen's character selection tool to train a personal AI likeness that can be inserted into any scene. Apply: Gather ~20 photos of yourself at varying angles and lighting and upload them to HeyGen to train a Soul ID before generating scenes featuring yourself.

  22. Character sheet generation (Nano Banana Pro) — Generating a reference character sheet from source images to lock a character's visual identity across shots. Apply: Run trained character images through Nano Banana Pro to produce a character sheet, then reference that sheet in every subsequent shot prompt for consistency.

  23. Soul Cinema image generation — Higgsfield's Soul Cinema tool, claimed to generate cinematic images by default and to outperform models like Nano Banana Pro at photorealism. Apply: Use Soul Cinema for generating key frames when photorealism and a cinematic look are the priority.

  24. Multi-shot single-prompt generation — Generating multiple camera angles or shots for a scene from a single prompt pass. Apply: Write one detailed scene prompt and request several shot variations from it rather than a separate prompt per angle.

  25. Start-frame continuity via screenshot — Taking a screenshot of an already-generated shot's final frame and feeding it back in as the start-frame reference for the next shot. Apply: After generating a shot, screenshot its last frame and use it as the anchor/start-frame input for the following shot in the sequence.

  26. Prompt expansion via Claude — Feeding a plain-language scene description into Claude to produce an improved, more detailed prompt for the image/video model. Apply: Draft the scene in plain language first, then run it through Claude to expand it before submitting it to Soul Cinema or Cinema Studio.

  27. Per-scene visual identity — Assigning each scene its own distinct color grade, camera style, and lens choice rather than keeping the whole film visually uniform. Apply: Before generating a scene, define its specific color grade, camera style, and lens, and keep that treatment consistent only within that scene.

  28. Dash-cam/surveillance aesthetic — An intentionally lower-quality, DVR-style visual treatment used for a specific POV shot such as a patrol-car dispatch view. Apply: Specify a dash-cam or DVR-quality look in the prompt when a scene calls for a surveillance or in-vehicle POV.

  29. Handheld camera as emotional signal — Using handheld camera shake narratively to signal a scene's emotional tone rather than purely for style. Apply: Choose handheld shake for emotionally charged or unstable moments and static/smooth camera for calmer beats.

  30. B-roll interstitials with dialogue overlay — Short atmospheric B-roll shots generated to break up a scene, with one dialogue audio track spanning multiple B-roll visuals. Apply: Generate one continuous dialogue track, then cut in several short atmospheric B-roll shots underneath it instead of holding on the speaking character for the whole line.

  31. Post-production shot recovery — The claim that any missing or needed shot identified during editing can be generated afterward using the same character sheet and prompt workflow to match the existing aesthetic. Apply: During editing, flag missing shots and regenerate them later with the same character sheet and prompt style rather than reshooting the whole scene.

  32. Voice cloning with preserved performance (Pixel Audio 'change voice') — A Pixel Audio feature that swaps vocal timbre while claimed to preserve the original emotion, pacing, and delivery of the source recording. Apply: Record your own natural line reading first, then run it through Pixel Audio's 'change voice' to swap in a different voice while keeping the original performance.

  33. Voice track splitting per character — Separating audio into distinct tracks per character so each can be individually voice-cloned or assigned a preset voice. Apply: Record or generate each character's lines on a separate audio track, then apply a distinct voice clone or preset voice to each track independently.

  34. Preset voice selection — Using a pre-built voice from Pixel Audio's library for secondary characters instead of cloning a real voice. Apply: For characters who don't need a personal voice clone, pick a fitting preset voice from the tool's library.

  35. Iteration as the core skill — The video's framing that generative filmmaking is fundamentally a loop of running, watching, tweaking the prompt, and rerunning. Apply: Budget for multiple failed generations per shot and treat each failure as a signal to adjust the prompt rather than a stopping point.

  36. Character + location + moment concept framework — A three-part framework (character, location, moment) the video uses for its audience-submission call, mirroring its own three-sentence story method. Apply: Pitch or scope any short film idea by naming its character, its location, and its key moment in three sentences.

Insights

The video's own three-sentence story (a depressed cop, his birthday, friends trying to cheer him up) draws on the creator's autobiographical detail — her own birthday ritual and favorite Lagman noodle spot — fictionalized into the cop character, suggesting the 'feeling-first' method in practice means mining real personal experience rather than inventing from scratch.

The claimed voice-cloning behavior (only timbre changes; emotion, pacing, and delivery are preserved) implies the human's original recorded performance is what actually carries the acting — the AI functions as a timbre filter, not a performance generator.

'Nothing is locked until you lock it' reframes AI production as non-linear and recoverable: unlike traditional film, where a missing shot means an expensive reshoot, any shot can be regenerated after editing and dropped in later to match the established look via the character sheet.

The screenshot-to-start-frame technique chains generations together — each new shot is conditioned on a screenshot of the previous one — building continuity incrementally instead of generating every shot independently from scratch.

The closing 'submit your 3-sentence concept' ask functions simultaneously as audience engagement and as top-of-funnel marketing for Higgsfield's paid tiers, embedding product promotion directly inside the tutorial's call-to-action.

«No camera, no crew, no $50,000 production budget.»

— 00:18

«Most people start with a prompt. That's the wrong move. Start with a feeling.»

— 02:24

«What's a story with real emotional stakes? Something an audience actually feels. Not what looks cool, what will get views. What's a story worth telling?»

— 02:32

«A cop who suffers from depression. It's his birthday, and his close ones try to cheer him up. That's it. Three sentences. That's your entire foundation.»

— 02:43

«Set up, midpoint, climax. You don't need more than that for short film.»

— 03:11

«That whole exchange happened in one shot. No cuts, just two guys in a locker room, one of them trying way too hard to pretend he's fine, and the other one not letting him get away with it.»

— 11:06

«Nothing is locked until you lock it.»

— 22:46

«The emotions, the pacing, the delivery, all of that stays exactly the same.»

— 23:19

«No camera, no budget, just a feeling that I wanted to put on screen. And all the tools I needed to do so, all in one place.»

— 25:30

«That's just part of the process. You run it, you look at it, and then you tweak a thing or two in the prompt, and run it again. That's the whole game. Iteration is the skill.»

— 25:46

Reception

Highly engaged creative community excited about the tool, with abundant participation in a film concept contest and enthusiasm for AI video quality, though tempered by concerns about cost and artistic limitations.

The video functions as much as a promotional walkthrough for Higgsfield's own tool suite (Soul Cinema, Cinema Studio, HeyGen, Pixel Audio) as a tutorial, but it does lay out genuinely reusable story-structuring and shot-continuity techniques, while product superiority claims (e.g., Soul Cinema's photorealism over competitors) are asserted rather than demonstrated head-to-head.

26:28

↳ Higgsfield AI · YouTube

Watch original