Lore

AI video generation workflow

5-Step Workflow To Make Ultra-Realistic AI Short Films (Seedance 2.0 4K)

Adil (Higgsfield AI) argues that Runway ML 2.0 / Seedance 2.0 in 4K (called "C Dance"/"C-DANCE" throughout the video) crosses a realism threshold where AI-generated shots become indistinguishable from real footage, and he demonstrates a repeatable asset-based workflow — building characters, locations, and props separately via Claude-written prompts and Hexel's Cinema Studio/Soul Cinema/GPT Image 2.0 — for assembling a hyperrealistic 1-minute short film.

Higgsfield AI · 2026-07-02 · English

Key ideas

  1. The film's premise (a teen glued to hyperrealistic C-DANCE 4K videos on TV while his mom calls him to dinner) is used as the vehicle to showcase the tool, with the TV screen embedding the actual demo clips.

  2. Every scene is claimed to need three ingredients — characters, locations, and props — built and locked in as individual reference assets before being combined into a scene prompt.

  3. Claude is used exclusively to draft/refine prompts from uploaded reference images or videos; Hexel (via Cinema Studio) is used to generate and organize the actual images/video.

  4. 4K resolution is presented as the key qualitative leap: skin pores, individual eyelashes, forehead wrinkles, distant background texture, and large dynamic crowds hold up in 4K where 1080p reportedly blurs, glitches, or dissolves into 'slop.'

  5. No single model is claimed to be best at everything; the workflow explicitly switches between GPT Image 2.0, Nana Banana Pro, and Soul Cinema depending on which produces better depth/realism for a given asset.

  6. Locked-in character/location/prop assets are saved as named 'elements' inside per-project Cinema Studio folders for reuse across scenes.

  7. Scene-to-scene transitions (including a room-to-battlefield 'scale transition') are achieved without hard cuts by carrying continuity from one generated clip into the next.

  8. The finished 1-minute film spans multiple genres inside its diegetic TV-within-video framing: a snow-leopard wildlife documentary, a fantasy jungle scene, a post-apocalyptic chase, and a full epic battle.

  9. Adil states he is sharing the exact prompts, a Claude 'skill,' and the full workflow in the video description so viewers can reproduce the process from scratch.

  10. Character Sheet Method (Gray Background + Three-View) — A reference-image technique where character sheets are generated against gray backgrounds with front/side/back views plus close-up face and full-body shots to maximize model consistency. Apply: When generating a character sheet, request a gray-background image showing close-up face, full-body, and side/back angles before feeding it into the video model.

  11. Three-Quarter Camera Angle for Locations — The claim that location reference images shot from a three-quarter angle give the model better depth perception than a flat frontal shot. Apply: Generate location assets from a three-quarter camera angle rather than a straight-on view to improve depth in the resulting video.

  12. Cinema Studio Element Library — A per-project asset-management feature inside Hexel where saved characters, locations, and props are stored as named 'elements' that can be referenced automatically in later prompts. Apply: Create a dedicated project folder in Cinema Studio and save each locked-in character, location, and prop as a named element so it can be reused without re-uploading.

  13. Model Switching / Comparative Generation — The claim that no single model (e.g., GPT Image 2.0 vs. Nana Banana Pro) performs best on every task, so results should be compared across models. Apply: If an image lacks depth or realism, regenerate the same prompt in a different model rather than only re-tweaking the prompt within one model.

  14. Red Arrow Annotation — Drawing a red arrow directly on a prop reference sheet to unambiguously point to which button or element a character should interact with. Apply: When a prompt repeatedly fails because the model interacts with the wrong object, annotate the reference image with a red arrow pointing at the correct target.

  15. Final-Frame Continuity Stitching — A scene-transition technique where a segment of one generated clip (e.g., its first six seconds) is reused as the reference for the next scene to avoid a hard cut. Apply: Extract a portion of a completed clip and feed it as the starting reference for the next scene's generation to carry visual continuity across the cut.

  16. Iterative Prompt Refinement via Claude — A feedback loop where a generated video's flaws are reviewed, then Claude is asked to add specific details (camera movement, slow motion, emotion, environment) to the next prompt. Apply: After each generation, identify what's missing and describe those additions to Claude so it writes an updated, more detailed prompt for the next pass.

  17. Video-to-Voiceover Generation — A workflow where a finished video clip is uploaded to a narration tool (called 'the supercomputer' in the video) that analyzes the footage, writes a documentary-style script, and overlays a selected AI voice. Apply: Upload the rendered clip to the narration tool, request a documentary-style script based on the visuals, choose a voice, and overlay the narration onto the video.

  18. Exact Reference-Duration Matching — The claim that C-DANCE/Seedance will fabricate content instead of following a reference video if the requested output duration doesn't precisely match the reference clip's length. Apply: When using a clip as a continuation reference, set the generation duration to match the reference clip's exact length to prevent the model from inventing unrelated content.

  19. Multiple Character Sheets for Cast Diversity — The claim that reusing one character sheet across a crowd scene produces an identical 'clone army,' while supplying several distinct character sheets produces a varied cast. Apply: For scenes needing multiple distinct-looking characters, generate and supply several different character sheets rather than reusing a single one.

  20. Seedance 2.0 4K (referred to in-video as "C Dance"/"C-DANCE") — The 4K video-generation model central to the tutorial, claimed to preserve fine detail (skin pores, eyelashes, distant textures, crowd geometry) that 1080p output loses. Apply: Generate final video shots in the 4K model rather than 1080p when fine texture, distant detail, or large dynamic crowd scenes need to stay sharp.

  21. Claude for Prompt Writing — Using Claude as the prompt-generation layer: uploading a reference image/video and describing the desired change so Claude drafts the detailed prompt fed into the image/video models. Apply: Upload the current asset to Claude, describe the desired change in plain language, and use Claude's returned prompt as the input for the image/video generator.

  22. Soul Cinema — An image-generation tool used for cinematic location and character/prop stills, priced at one credit per eight generated images. Apply: Use Soul Cinema for generating cinematic location reference images and character/prop stills when the one-credit-per-eight-images pricing matters.

  23. GPT Image 2.0 as High-Fidelity Input Reference — Using GPT Image 2.0 at 4K quality to produce the still reference image fed into the video model, on the claim that a sharper input image yields a sharper final video. Apply: Generate character/location/prop reference stills in GPT Image 2.0 at 4K quality before passing them to the video generator, since input image quality is claimed to cap output video quality.

  24. Prop Sheet Volume & Lighting Detail — The claim that prop reference sheets need clear physical volume, highlights, and shadows for the object to animate realistically rather than looking flat or plasticky. Apply: When generating a prop reference sheet, ensure it shows clear highlights and shadow falloff rather than a flat, evenly lit rendering.

  25. Specificity Over Prompt Permissiveness — The claim that giving the model too much creative freedom causes visual artifacts and degraded on-screen text/imagery ('slop'), while specifying exact on-screen content prevents it. Apply: When a scene includes on-screen text, screens, or UI elements, specify exactly what should appear rather than leaving it to the model's discretion.

Insights

Gray backgrounds are claimed to outperform white or black backgrounds for character-sheet consistency — a specific, non-obvious production detail rather than a general best practice.

Higher resolution is framed as changing the failure mode of generation itself, not just cosmetic sharpness: at 1080p, dynamic crowd/battle scenes are claimed to 'glitch out' or dissolve into blur, while 4K is claimed to hold geometry together even under a shifting camera perspective.

Output quality is described as bottlenecked more by input reference quality (image resolution, camera framing like the three-quarter angle for locations) than by the video model itself — 'your video quality depends mostly on this one image.'

Reusing a single character sheet across a crowd scene is claimed to produce an unwanted 'clone army'; achieving a diverse cast requires generating and feeding in multiple distinct character sheets.

Prompt permissiveness is treated as a liability rather than a stylistic choice: leaving on-screen details (like TV/UI content) unspecified is claimed to increase visual-artifact risk, so more explicit prompting is framed as safer, not just more literal.

The model is claimed to fabricate content instead of following a reference clip if the requested generation duration isn't set to match the reference's exact length, implying reference-based continuation is duration-sensitive in a way that isn't obvious from the interface alone.

A visual annotation trick (drawing a red arrow on a prop reference image) is presented as a fix for a specific class of failure — the model repeatedly interacting with the wrong element — suggesting prompt text alone wasn't sufficient to disambiguate intent.

«Runway ML's 4K is unreal.»

— 00:39

«Wait, is that 4K?»

— 01:02

«I'm going to show you how to make cinematic shots so convincing that people won't even question whether it's AI. They'll just assume it's real.»

— 01:15

«After tons of testing, we've discovered that character sheets on a gray background just perform way better than those on white or black backgrounds.»

— 03:17

«It's really important to pick the correct location because your video quality depends mostly on this one image, so make sure to examine each image, and if something looks off, regenerate.»

— 05:36

«There's no one best model that does everything perfectly. You try GPT Image, if the result isn't what you wanted, you try a different model or tweak your prompt.»

— 07:14

«If you look closely here, you can see the pores in detail, every single eyelash, how natural the hair turned out, and even the wrinkles on the forehead.»

— 14:31

«C Dance 4K keeps the details all the way into the distance, where 1080p usually falls apart.»

— 17:04

«Normally in standard 1080p, this is where the whole generation would just glitch out. But here, the dynamics are insane.»

— 21:47

«Otherwise, C dance will just start making things up»

— 22:49

«What blows my mind the most is how realistic I turned out»

— 26:53

«It looks exactly like a scene out of a hundred million dollar blockbuster.»

— 30:41

«C-DANCE 4K is unreal.»

— 32:19

«After making this, I really don't think I can go back to 1080p.»

— 32:54

«High above the world, where the air grows thin and the wind never sleeps.»

— 00:17

Reception

Viewers appreciate the video quality and tool capabilities, but widespread frustration over high credit costs and lack of transparency about pricing significantly dominates the reception.

This is a detailed, reproducible production tutorial — asset-by-asset prompts, specific tool combinations, and named troubleshooting fixes — rather than a general theory piece, and its practical value is tightly coupled to the current state and pricing of Higgsfield's Hexel/Cinema Studio/Soul Cinema/Seedance ('C-DANCE') 4K toolchain, which the video itself frames as a just-released capability.

33:29

↳ Higgsfield AI · YouTube

Watch original