AI video generation
AI videos look cheap not because the underlying AI is bad but because of how people use it; Adil (Higgsfield AI) argues there are five specific, fixable mistakes behind that cheap look — skipping a start frame, using a low-quality start frame, writing vague prompts, failing to keep characters consistent across shots, and overloading a scene with too many simultaneous actions — and pairs each mistake with a concrete fix.
Mistake 1: typing a long prompt straight into text-to-video is 'basically a slot machine' with zero control over composition; the fix is to use a start frame — an image the video begins from or animates.
Mistake 2: using a blurry or 'plasticky' AI-looking start frame carries that quality through the whole video; the fix is generating a production-ready image first, using Soul 2 and the prompt enhancer, since 'your start frame is responsible for 80% of your result.'
Mistake 3: vague prompts like 'a cinematic scene' or 'a cool shot' produce vague results; the fix is deciding who, where, what's happening, what camera move, and what mood before prompting.
Mistake 4: stitched-together clips often show the same actor with different hair/eye color or a birthmark that jumps position; the fix is building a character sheet (via Nana Banana Pro), uploading and tagging that character in Cinema Studio, and using multi-shot generation for consistency across angles and cuts.
Mistake 5: AI models can only convincingly render so much at once, so packing a scene with many simultaneous events (e.g., 15 explosions, an earthquake, and an emotional conversation) overwhelms both the tool and the viewer; the fix is one primary action plus one or two secondary actions per scene, splitting complex scenes into separate shots (wide/medium/close-up) and stitching them, or using multi-shot with per-scene prompts.
Pro tip: feed the start-frame image to an LLM like Gemini or ChatGPT, describe the animation simply, and ask it to write a detailed prompt — but read and tweak it rather than pasting it blindly, and add 'Do not generate just give me the prompt in all caps at the end' because Gemini tends to start animating the image instead of returning text.
Pro tip: use motion control (record yourself performing the action, transfer it onto your character) with scene control turned on for pinpoint camera/action precision.
Pro tip: brief image generation with full production detail — clothing, makeup, lighting, composition — the way a real photo shoot would be briefed, i.e. 'channel your inner movie director' rather than thinking like a prompt engineer.
Pro tip: for fast-motion scenes that keep morphing, generate a slow-motion version of the scene and speed it up afterward in any editing tool.
Start frame — An image used as the starting point of a video generation, as opposed to relying on a pure text-to-video prompt. Apply: Provide a start-frame image before generating, so the video has a defined beginning and the text prompt only needs to describe what happens next.
Soul 2 — An image-generation model the video calls 'hands down the best model right now for generating production ready images.'. Apply: Select Soul 2 in Images when creating the start-frame image you plan to animate.
Prompt enhancer — A toggle in the image workflow that expands a simple description (e.g., 'a man skating') into a fuller prompt. Apply: Write a simple description of the desired image and turn on the prompt enhancer so it automatically enriches the prompt.
Soul ID — A personal identity model trained on a set of one's own images (the video mentions over 20) used to keep a specific likeness consistent in generated images. Apply: Train a Soul ID on 20+ images of the person you want represented, then select it when generating the start-frame image so the face stays consistent.
Cinema Studio camera-movement and genre selection — A feature in Higgsfield's Cinema Studio letting a user pick a camera movement (e.g., handheld) and a genre/mood from a list instead of describing them in text. Apply: Instead of prompting for 'a cinematic scene,' select a specific camera movement such as handheld for natural shake, and choose a genre from Cinema Studio to set the mood.
Who/Where/What/Camera/Mood prompt structure — A five-element prompt framework recommended in place of vague prompts: subject, location, action, camera move, and mood. Apply: Before writing a prompt, decide the subject, the location, what's happening, the desired camera movement, and the mood, then write those explicitly into the prompt.
Character sheet via Nana Banana Pro — A reference set produced by generating the same character from multiple angles using Nana Banana Pro, used to lock in a consistent look before animating. Apply: Generate the same character from different angles with Nana Banana Pro to build a character sheet before uploading the character into a video tool.
Character tagging in Cinema Studio — A feature where an uploaded character (given a name and description) can be tagged in a prompt so it appears consistently across shots instead of being re-described each time. Apply: Upload the character sheet into Cinema Studio, name and describe the character, and tag that character in each shot's prompt rather than writing 'he is doing something' from scratch.
Multi-shot generation — A generation mode producing a sequence with scene cuts built in, keeping a tagged character consistent across the cuts and angles. Apply: Use multi-shot generation with a pre-built, tagged character to produce a sequence of scene-cut shots without face drift or morphing between cuts.
One-primary-action-per-scene rule — A guideline that a scene should center on one primary action plus at most one or two secondary actions, since AI models can only handle so much happening at once. Apply: Limit each generated scene to a single primary action rather than combining many simultaneous events, prompting each scene separately if multiple things need to happen.
Shot-splitting and stitching by shot scale — A technique for handling a scene with multiple simultaneous events by generating each event as its own shot at an appropriate scale (e.g., explosion as wide shot, earthquake as medium shot, conversation as close-up) and combining them. Apply: Generate each major event of a complex scene as its own shot at a fitting scale in a video editing tool, or use multi-shot with a separate prompt per scene, then stitch the shots together.
Motion control — A Higgsfield feature that transfers a user's recorded real-life action onto an AI-generated character for precise action or camera control. Apply: Record a quick video of yourself performing the desired action, bring it into motion control, and transfer it onto your character.
Scene control — A setting the video says should be enabled alongside motion control for 'the best results.'. Apply: Turn on scene control whenever using motion control to transfer a recorded action onto a character.
LLM prompt-drafting hack (Gemini/ChatGPT) — A workflow of uploading a start-frame image to an LLM with a simple animation description and asking it to return a detailed prompt. Apply: Upload the image to Gemini, ChatGPT, or another LLM, describe the animation in simple terms, ask for a detailed prompt, then read and tweak the result rather than pasting it unedited; since Gemini tends to start animating the image instead of returning text, add 'Do not generate just give me the prompt in all caps at the end.'
Slow-motion-then-speed-up workaround — A fix for fast-motion generations that keep morphing: generate the scene in slow motion, then speed the clip up afterward in an editing tool. Apply: When a super-fast-motion generation keeps morphing, generate a slow-mo version of the same scene instead, then speed it up in any editing tool.
Detailed image-prompting (clothing/makeup/lighting/composition) — A principle that image-generation prompts should include as much production detail as possible, briefed the way a real photo shoot would be, rather than a minimal prompt-engineering request. Apply: When generating images, specify clothing, makeup, lighting, and composition in detail, approaching the prompt as a movie director briefing a shoot rather than as a terse prompt-engineer request.
The 80% weighting given to the start frame reframes where effort should go: most of the work should happen at the image-generation stage, not in tweaking the video prompt afterward — summarized in the video as 'garbage in, garbage out.'
Character consistency is framed as a workflow problem solved by pre-building an asset (a multi-angle character sheet plus a tagged, named character) rather than a prompting problem solved case-by-case for each shot.
The five-part prompt structure (who, where, what's happening, camera move, mood) is explicitly modeled on how a film director briefs a shot, not on 'prompt engineering' vocabulary — the video repeats this framing ('channel your inner movie director') as its core mental model.
The fix for fast-motion morphing artifacts is a post-production workaround (generate slow motion, then speed up in editing) rather than a generation-side prompting fix, treating a model limitation as solvable outside the tool itself.
A specific behavioral quirk of Gemini is named and worked around explicitly (it tries to animate the uploaded image instead of returning the requested text prompt), which is a granular, tool-specific friction point rather than a general AI-video principle.
The video's own call-to-action — commenting personal pro tips to enter a giveaway — visibly generated additional, non-scripted technique content in the comments (e.g. lighting direction as a control lever, cross-checking prompts across multiple LLMs, negative prompts), functioning as crowdsourced validation of the video's central claim that specificity and control beat vague prompting.
«Most AI videos look bad. Not because of the AI, but because of how people use it.»
— 00:00
«That's text to video. And it's basically a slot machine. You have zero control over the composition.»
— 00:34
«Your start frame is responsible for 80% of your result. If you have a bad start frame, no amount of prompt engineering is going to save it.»
— 02:14
«And say it with me, garbage in, garbage out.»
— 02:22
«You don't necessarily need more words. You just need clear intent.»
— 02:41
«And once you see it, you just can't unsee it.»
— 04:42
«Focus on one primary action per scene and maybe one or two secondary actions.»
— 06:10
«That's how masterpieces are born.»
— 06:30
«Do not generate just give me the prompt in all caps at the end.»
— 07:20
«If you want to get ahead of 90% of the people, you can't keep thinking like a prompt engineer. You need to channel your inner movie director.»
— 07:56
Reception
Predominantly positive reception with strong constructive engagement from practitioners, though a vocal minority expresses skepticism about AI tools while most substantive comments offer helpful tips and testimonials of practical value.
This is vendor content — a Higgsfield creator using Higgsfield's own models and Cinema Studio to illustrate the five fixes — but the diagnostic frame (start frame quality, prompt specificity as director's brief, character asset-building, one-action-per-scene, and post-production workarounds like slow-mo-then-speed-up) is generalizable beyond the specific tools used to demonstrate it.

09:14