WAN 2.5
The video, sponsored by Higgsfield, argues that WAN 2.5 is a powerful, HD-capable, audio-enabled video model that is especially good at following instructions in image-to-video generation, and it demonstrates this by walking through a full product-commercial workflow (image generation with Seedream/Nano Banana, storyboarding, then animating each shot with WAN 2.5) rather than delivering a direct technical comparison against VEO 3 despite the title.
WAN 2.5 is presented as the newest, 'craziest uncensored' video model, capable of producing full HD video up to 10 seconds long with audio, and supporting image-to-video generation.
The video is sponsored by Higgsfield, which gave Thomas 10 Pro plan promo codes to give away to viewers for hitting 10,000 subscribers, plus a separate 10% discount code for everyone else.
Higgsfield's main appeal, per Thomas, is that it aggregates access to many image and video AI tools (WAN 2.5, Nano Banana, Seedream) in one place, so an image model and WAN 2.5 can be combined in one workflow.
Community examples shown include a Dumbledore-as-iPhone-spokesperson clip and a Leonardo DiCaprio 'Wolf of Wall Street'-style Nutella ad, both using celebrity likenesses, which Thomas says he'd personally be careful about using.
Thomas demonstrates a full commercial-making workflow: generate a product image (his fictional 'Banana Pringles') with Seedream, brainstorm and storyboard a narrative (tree in the jungle, sad gorilla, gorilla discovers the Pringles, gorillas fight over it, hero shot + tagline), then animate each shot with WAN 2.5.
He used the WAN 2.5 'fast' model with 5-second clips at 1080p, though the platform supports up to 10-second clips at 1080p.
His assessment: WAN 2.5 is strong at following instructions, especially for image-to-video; not every clip was a first-try success, but only minor prompt tweaks were usually needed to get the intended result.
A stated downside is that generations are on the slower side, but Higgsfield lets users queue multiple requests to run in parallel while doing other tasks.
Thomas's closing take: WAN 2.5 is 'a very powerful video model, but with great power comes great responsibility,' so he advises using it responsibly.
WAN 2.5 — A newly released video generation model, described in the video as 'uncensored,' capable of producing full HD video up to 10 seconds long with audio and supporting image-to-video generation. Apply: Use it inside Higgsfield to animate a still image (e.g., a product or scene keyframe) into a short video clip, choosing the 'fast' variant for quicker 5-second/1080p generations or the standard mode for up to 10-second 1080p clips.
Higgsfield — An AI platform (this video's sponsor) that aggregates access to multiple third-party image and video generation models in one place, including WAN 2.5, Nano Banana, and Seedream. Apply: Sign up for a free account, generate images with one of its image models, then hand those images to WAN 2.5 within the same platform to animate them, and queue multiple generation requests at once to work around WAN 2.5's slower render times.
Nano Banana — One of the image generation models made available inside Higgsfield, cited by the video as an option for creating starting images. Apply: Use it to generate a product or scene image, which can then be animated into video using WAN 2.5.
Seedream — An image model available in Higgsfield (rendered in the transcript as 'Cream 4'/'Crereem 4') that Thomas used to brainstorm and generate the concept images for his Banana Pringles commercial. Apply: Use it iteratively to generate each new keyframe image (establishing shot, character shot, product shot, etc.) as the ad's concept develops, before animating each frame.
Image-to-video generation — A generation mode where a still image is used as the starting point/input for producing a video clip, rather than generating video from text alone. Apply: Feed a generated or real product image into WAN 2.5 to animate it into a short clip; the video reports WAN 2.5 handles this well, closely following the accompanying instructions.
"Next frame" storyboarding technique — Thomas's method of generating a still image representing the next section/beat of the commercial before animating it, effectively storyboarding the ad shot by shot. Apply: For each new scene in the narrative, generate a keyframe image with an image model, then animate that specific frame with WAN 2.5, chaining the resulting clips together into a longer, coherent commercial.
Prompt-embedded voice direction — Per Thomas's reply in the audience comments, WAN 2.5 can generate spoken dialogue/audio directly from an instruction describing the voice and line written into the prompt (e.g., 'Deep commercial male voice says: Banana Pringles etcetc'). Apply: Write the desired voiceover line and voice style directly into the WAN 2.5 prompt text to have the model generate matching audio alongside the video.
The 'next frame' storyboarding technique — generating a still keyframe image for each upcoming scene beat and then animating that specific frame with WAN 2.5 — is the concrete method Thomas uses to keep multi-shot narrative continuity across separately generated clips.
Higgsfield's request-queuing turns WAN 2.5's stated weakness (slow generation) into a workaround: batch multiple prompts and multitask while they render, rather than waiting on each clip serially.
Per Thomas's reply in the audience comments, WAN 2.5's audio/dialogue is prompt-driven rather than a separate voice step: he generated the ad's voiceover by writing a direct instruction into the prompt ('Deep commercial male voice says: Banana Pringles etcetc').
Audience comments confirm WAN 2.5 has no native feature to seamlessly stitch multiple clips into one longer video — any 'longer ad' result comes purely from Thomas's manual consecutive-clip/next-frame workflow, not a built-in platform capability.
A viewer's comment adds a practical prompting tip not stated by Thomas himself: keeping very similar wording across prompts from scene to scene helps maintain consistency.
«Today we're checking out WAN 2.5, the newest, craziest uncensored video model capable of producing full HD video up to 10 seconds with audio.»
— 00:00
«But I would personally be a bit careful using celebrities likenesses in this type of way.»
— 01:18
«The best thing about Hicksfield in my opinion is that you have access to all the image and video AI tools available.»
— 01:30
«Banana Pringles. A taste worth going apeshit for.»
— 02:15
«Basically, I always tried to create kind of a next frame for a next section of the video, which I would then animate using one 2.5.»
— 03:18
«From my limited testing, it seems to be really good at following instructions, at least when it comes to image to video generation.»
— 03:56
«So, as a summary, when 2.5 is a very powerful video model, but with great power comes great responsibility. So, use it responsibly.»
— 04:47
Reception
Enthusiastic engagement with the concept and results, but promo code distribution problems and tool feature limitations dampened some satisfaction.
Despite its VEO 3-comparison title, this is a sponsored, workflow-focused walkthrough of WAN 2.5 inside Higgsfield rather than a rigorous head-to-head benchmark; its concrete value is the reusable 'storyboard-frame-then-animate' production technique and Thomas's brief, self-described 'limited testing' impressions of instruction-following and speed.

05:24