Lore

Higgsfield AI

How to Make Viral AI UGC for TikTok Ads (Higgsfield Tutorial)

The video's core claim is that Higgsfield AI's Canvas — a node-based workflow builder that chains an LLM assistant, image generation, and video generation nodes together in one place — is the best way to produce believable, non-polished 'UGC'-style AI video ads (e.g., for TikTok), because it lets the creator connect and iterate on a full pipeline (product image → target-audience analysis → actor image → script → video → audio/B-roll fixes) inside one organized canvas instead of switching between separate single-purpose tools.

Artturi Jalli · 2026-05-13 · English

Key ideas

  1. Use Higgsfield's Canvas rather than the standalone UGC Ad Studio or the regular video generator, because Canvas lets nodes be chained into one continuous, organized, iterative workflow in a single place with one subscription.

  2. Start the workflow by uploading a product image (a screenshot taken from the product's landing page) as the first node.

  3. Connect the product image to an LLM Assistant node (Claude) and prompt it to analyze the product via its landing page link and describe its ideal target user, to ground later prompts in that audience.

  4. Feed the LLM's audience description plus the reference product image into an image generation node to create a 'UGC actor' holding the product, matching the described target audience.

  5. First-pass results using NanoBanana2 looked too polished/advertisement-like with 'no UGC vibe'; adding explicit 'non-polished and slightly amateurish, like a friend sending a recommendation' language, setting a 9:16 aspect ratio, and switching to the TPT image model produced much better UGC-style images.

  6. Manually select the best of the generated images (out of four) to use as the seed frame for video generation.

  7. Prompt the LLM again (Claude Sonnet 4.6) to write a ~15-second UGC-style script to be read aloud, explicitly instructing it that the result 'should not sound like an ad but rather just a candid and genuine recommendation made by a friend,' referencing the chosen actor image for scene context.

  8. Feed the chosen image and generated script into a video generation node using the Seedance 2.0 model, with settings: 1 video, 480p resolution (for cheap/fast testing), 9:16 aspect ratio, 'generate audio' turned on (off by default, and skipping it is called 'a catastrophe'), and 15-second duration.

  9. Iterate and regenerate based on quality-control review: the first video had unwanted background music and an unnatural-looking prop (remote controller) and was discarded; the second had no background music and looked much better but had a continuity error (the drone appears to be filming itself from the air with no explanation); a third regeneration fixed the continuity issue.

  10. Fix AI audio that sounds 'manufactured and artificial' by sourcing free ambient sound effects from pixabay.com and layering them under the video in a free video editor, adjusting volume to match the scene.

  11. Optionally create AI-generated B-roll: regenerate the seed image with the person removed, reframe it as if shot from 50 ft above the same spot, then animate it with a prompt describing slow drone-like movement across the scene, to intercut with the main clip for added realism.

  12. Recap of the full loop: product image → AI-written target-audience description → AI-generated actor image → AI-written script → AI-generated video → manually added ambient audio → optional AI B-roll, all built and connected inside one Canvas.

  13. Higgsfield AI Canvas — A node-based AI workflow builder inside Higgsfield AI that lets you chain image, LLM, and video-generation nodes together in one continuous canvas instead of using separate single-purpose tools. Apply: Sign up to Higgsfield, click 'Canvas' in the top bar, click 'Create new canvas,' then right-click on the empty canvas to add and connect nodes (product image, LLM assistant, image generation, video generation) into a single pipeline.

  14. LLM Assistant node (Claude AI / Claude Sonnet 4.6) — A canvas node that runs a large language model — Claude, later specifically Claude Sonnet 4.6 — to analyze inputs and produce text such as an audience description or a video script. Apply: Connect a product image and a prompt (e.g., 'analyze this product [link] and describe its ideal user accurately') into the node and generate an audience description; reuse the same node type later with the chosen actor image plus the product link to generate a 15-second UGC script.

  15. NanoBanana2 — An image generation model available in Higgsfield's image node, described as fast and good for producing an initial set of images. Apply: Select it for a first-pass batch of UGC actor images when speed matters, then evaluate whether the 'UGC vibe' is right before deciding whether to switch models.

  16. TPT image model — A second image generation model in Higgsfield, described as the best current model for producing the intended non-polished, amateurish UGC look once the prompt is adjusted. Apply: Switch to this model (with a 9:16 aspect ratio set) after adding 'non-polished and slightly amateurish' language to the prompt, then rerun the image pipeline to get more UGC-realistic actor photos.

  17. Reference image conditioning — The practice of connecting a real product photo (and later a chosen actor photo) into a generation node as a 'reference picture' so the AI output matches the product's or person's exact appearance. Apply: Connect the source image into the image or LLM node alongside the text prompt so later outputs stay visually consistent with the real product or the previously generated actor.

  18. UGC-style prompt engineering ('non-polished/amateurish' framing) — A prompting technique that explicitly requests a non-polished, slightly amateurish look 'just like a friend sending a recommendation to their friends,' and, for scripts, a tone that 'should not sound like an ad but rather just a candid and genuine recommendation made by a friend,' to stop the AI from defaulting to slick, ad-style output. Apply: Add this exact framing language into both the image-generation prompt and the script-generation prompt whenever default output looks too polished or advertisement-like.

  19. Seedance 2.0 (video generation model) — The video generation model selected in the video generation node, described as currently the best option, used to turn a starting image and script into a 15-second 9:16 clip. Apply: Select it in the video generation node along with 480p resolution for cheap testing, 9:16 aspect ratio, 'generate audio' enabled, and 15-second duration, then scale up to higher quality once a shot is validated (an audience comment estimates roughly $5–$10 per video at the highest-quality Seedance 2.0 settings).

  20. Low-resolution test-run practice — Deliberately generating a single video at 480p first, since output quality is unpredictable, before spending time and money on full-quality renders. Apply: Set resolution to 480p and quantity to 1 for early iterations; raise resolution/quantity only once the shot's framing, continuity, and audio are confirmed usable.

  21. Iterative regeneration / QC review loop — A repeat-and-inspect workflow where each generated video is checked for problems (unwanted background music, unnatural prop rendering, continuity errors such as an unexplained aerial drone shot) and regenerated with adjusted prompts until usable. Apply: After each generation, review the clip for audio issues, object rendering, and narrative plausibility, add specific corrective notes to the prompt, and rerun before accepting the shot.

  22. Ambient SFX layering via Pixabay — A manual post-production fix for AI-generated audio sounding 'manufactured and artificial' due to missing ambient background noise, using free sound effects from pixabay.com added in a separate video editor. Apply: Download a matching ambient sound effect from pixabay.com, drag both the generated video and the sound effect into a free video editor's timeline, and adjust the SFX volume so it matches the depicted scene.

  23. AI-generated B-roll technique — A method for producing supplementary drone-style B-roll by regenerating the seed image with the person removed and reframed as if shot 50 ft above the same spot, then animating it with a slow camera-movement prompt. Apply: Ask the AI to regenerate the actor's starting frame without the person, reframed 50 ft above the ground at the same location, connect that frame to a video generator with a prompt describing slow movement across the scene 'as if it was a drone shot,' and intercut the result with the main clip.

Insights

Model choice affects stylistic 'vibe,' not just technical quality: NanoBanana2 was fast and 'good' but defaulted to polished, ad-like images with no UGC feel, while the TPT image model combined with explicit 'amateurish' wording was needed to get an authentic-looking result — implying default/fast image models trend toward polish that has to be deliberately fought with both prompt language and model choice.

Testing at low resolution (480p, one clip) before scaling up is framed as a cost/risk-mitigation step specifically because 'we have no idea what kinds of results we're about to get,' i.e., output unpredictability is treated as a given, not an edge case.

AI-generated video can introduce logic/continuity errors that are invisible until playback — the drone footage appearing to film itself with no in-scene explanation — meaning the review pass has to check narrative plausibility of a shot, not just its visual polish or audio.

The realism gap in AI-generated UGC audio (lack of ambient background noise) is treated as something the generation step itself does not solve; the fix is explicitly manual post-production (external stock SFX plus a separate video editor), so the presented workflow is AI-generation-plus-manual-audio-finishing rather than fully automated end to end.

The same reference images are deliberately reused across pipeline stages (the real product photo feeds both the initial actor-image prompt and, later, is echoed in the person-removed B-roll frame) specifically to keep visual identity consistent across multiple separately generated shots.

«Okay, so this thing is actually insane. It's the DJI Avata 360 and it fits in your hand. I brought it to the pitch today and the footage it gets, like flying around while you're playing, is something I've never seen from a drone this small. Genuinely didn't expect it to be this good.»

— 00:05

«we're going to use the Higgsfield AI canvas because it allows us to build this kind of a continuous and organized workflow to get much more consistent and better results all in one place inside one canvas.»

— 00:37

«make it non-polished and slightly amateurish just like a friend sending a recommendation to their friends.»

— 03:37

«this is a good way to test things around before spending time and money doing the highest quality.»

— 06:32

«if you do that, you will not get any audio to your shot, so that is a catastrophe.»

— 06:48

«how on earth is the drone filming itself? Or is there another drone? It doesn't really make sense in this context.»

— 08:05

«what I actually recommend doing is if you find yourself in a situation where you don't have the ambient background noise, just go to a website like pixabay.com»

— 08:47

«the best way to do this is by just using the AI workflow builder or the AI canvas because it allows you to put everything inside one canvas with one subscription.»

— 10:11

Reception

Generally positive with genuine enthusiasm, but tempered by technical concerns about sound consistency and some criticism of execution.

This is a hands-on, iterative product tutorial for a specific paid tool (Higgsfield's Canvas) rather than a general theory piece — it maps a concrete node-by-node pipeline and documents several real regeneration cycles (wrong vibe, bad audio, a continuity error) instead of presenting AI UGC video generation as a one-shot solved problem.

11:28

↳ Artturi Jalli · YouTube

Watch original