A start frame is the image a video generation begins from or animates, used instead of relying on a pure text-to-video prompt. Typing a long prompt straight into text-to-video is "basically a slot machine" — there's zero control over composition, since the model has to invent the entire opening frame from words alone.
The start frame is claimed to be responsible for roughly 80% of the final result: "your start frame is responsible for 80% of your result. If you have a bad start frame, no amount of prompt engineering is going to save it." Whatever quality (or flaws) exist in the start frame — blur, "plasticky" AI-looking rendering, wrong likeness — propagates through the entire generated video. Garbage in, garbage out.
Generate a dedicated, production-ready start frame image first — using a capable image model (e.g. Soul 2: Production-Ready Start-Frame Image Model) — before animating it, rather than skipping straight to text-to-video or accepting whatever low-effort image is on hand. Brief that image generation with full production detail; see Brief Image Generation Like a Movie Director, Not a Prompt Engineer.
Related: Who/Where/What/Camera/Mood Prompt Structure for structuring the animation prompt itself once the start frame exists.
Из тем: Unsorted, Prompting Foundations: Briefing the Model Like a Director