AI agents
Axel Dunn (head of product), presenting on behalf of Higgsfield AI, argues that "Supercomputer" is the first AI agent to unify research, visual-content generation, and end-to-end marketing production (including website builds) inside one chat, and that its "skills" system plus automatic model routing let users get production-grade output without writing prompts or knowing which model to use.
Supercomputer is pitched as the first agent that does research, generates visual content, and runs marketing end-to-end, plugging into top reasoning models (Claude, Gemini, ChatGPT) and top content generators (Seedance/"Cida" 0.2, ChatGPT image, "Hexel Soul"), auto-choosing the right one per step.
MCP-based import lets users pull their full memory and skills from other agents (Claude, ChatGPT, "Hermes", "Open Claude") into Supercomputer in two clicks, replacing the manual exports and fragmented configs of migrating between AI platforms.
"Skills" are production-tested workflows built by Higgsfield's own creative and engineering team, encoding the same playbook the in-house production team uses; every new model release ships a new skill that all users get.
Supercomputer also writes its own personal skills by watching a user's repeated workflows and saving the pattern, so the agent becomes increasingly personalized over time (described as inspired by "Hermes").
Native multimodal understanding: the agent analyzes actual video/image frames directly rather than a text description of them, reading hooks, tracking pacing, and identifying what made a clip perform, then flags patterns that translate across formats.
"Model Orchestrator" automatically handles model selection, parameter tuning, asset routing, format choices, and retry logic per task; the video claims this produces a 15x reduction in cost versus running every step on an overpowered model.
Demo 1 (Hicks, an AI-generated clothing brand): from one URL and one instruction, Supercomputer researched the brand and viral TikTok/Instagram fashion trends in parallel, then produced 100 matched video briefs, scripts and hooks, and generated 100 finished UGC videos with no manual prompting.
Demo 2 (cinematic long-form video): "Picsel Elements" is described as a reference layer that pins characters, environments, props, and brand assets so they can be reused across every shot in a session; "Sol ID" is a trained identity model that holds facial likeness across generations.
The long-form demo ran on a named "Cinematic Flow" skill: characters and locations were generated in "Sol Cinema" and saved as elements, each clip was generated in "Cidase 2.0" using Cinema Studio tools while maintaining scene continuity, and Supercomputer stitched the clips into one sequence from a single command.
Demo 3 (beauty brand website): from one prompt and one product image, sent via a Telegram connector from a phone and synced live to the desktop app, Supercomputer researched top-performing sites in the category, generated product imagery and a hero video, assembled the site around that content, and deployed it.
Supercomputer ships with 27+ connectors (Telegram, Gmail, Slack, Notion, Superbase, GitHub, Docs), described as an open, weekly-expanding layer that gives the same agent, memory, and skills across different surfaces.
Closing claim: no other agent on the market combines content generation and site/product building at this level of quality in one place; Supercomputer is framed as "an entire marketing department running on one chat" and "the operating system for a new era of AI content."
MCP (Model Context Protocol) import — An open standard connecting agents to each other, used by Supercomputer to pull a user's entire memory and old skills from other agent platforms in two clicks. Apply: Use Supercomputer's MCP-based import to migrate memory and skills from Claude, ChatGPT, Hermes, or Open Claude instead of manually re-exporting configs and content.
Skills system (production-built) — Production-tested workflows built by Higgsfield's creative team and encoded by engineers, shipped with the product so the agent runs the same playbook the in-house production team uses. Apply: Invoke a skill for a task instead of hand-writing a prompt; the skill is picked automatically before any model call and every new model release adds a new skill for all users.
Personal skills (self-generated) — A mechanism where Supercomputer watches how a user works and, when it sees a repeated workflow, saves the pattern as a personal skill unique to that user. Apply: Repeat a workflow inside Supercomputer over time and let the agent turn it into a reusable personal skill, building a personalized skill library.
Native multimodal / direct visual analysis — An architectural choice where Supercomputer analyzes actual image and video frames directly rather than working from a text description of them, reading hooks, pacing, and performance patterns. Apply: Feed the agent raw video or image assets rather than descriptions so it can extract what makes a clip perform and reuse those patterns across new content.
Model Orchestrator — A routing layer that automatically selects the best-suited model for each subtask (e.g., Gemini for visual analysis, Claude for planning, Soul for characters) instead of the most expensive one, and also handles parameter tuning, asset routing, format choices, and retry logic. Apply: Describe only the task and let Model Orchestrator choose models and settings automatically, which the video claims cuts cost roughly 15x versus using one overpowered model for every step.
Picsel Elements (reference layer) — A reference layer sitting between the prompt and the generation model that pins characters, environments, props, and brand assets so they can be reused across every shot in a session. Apply: Pin a character, location, or brand asset once as an "element," then reuse that same reference across a multi-clip video to keep it consistent.
Sol ID — A trained identity model, described as going further than general reference-pinning for faces, that holds facial likeness across all generations. Apply: Use Sol ID when a project needs a specific character's face to stay consistent across many separately generated shots.
Cinematic Flow skill — The named skill Supercomputer ran to produce the long-form cinematic demo: generate characters/locations in Sol Cinema and save them as elements, generate each clip in Cidase 2.0 using Cinema Studio tools while maintaining scene continuity, then stitch everything into one sequence. Apply: Run the Cinematic Flow skill with a single command to auto-produce a multi-clip cinematic video that holds consistent characters and environments across the sequence.
Connector layer — A set of 27+ integrations (Telegram, Gmail, Slack, Notion, Superbase, GitHub, Docs) that give the same agent, memory, and skills across different surfaces, expanding weekly. Apply: Trigger a Supercomputer session from a connector such as Telegram on mobile, then open the desktop app to find the same synced session and continue the work.
The skills architecture is explicitly framed as compensating for a stated limitation of LLMs themselves — "LLMs alone don't know how to produce nonstop creative work. They guess at it" — so skills function as an injected production playbook rather than something the model reasons out on its own.
Model Orchestrator is presented as a direct response to a specific user complaint ("burning credits on overpowered models running trivial steps"), turning cost-efficiency into an explicit competitive claim (a stated 15x reduction) rather than a side effect of the routing design.
Higgsfield frames the friction of switching between competing AI agent platforms as itself a product opportunity: the MCP-based two-click import is pitched squarely at users already invested in Claude, ChatGPT, Hermes, or Open Claude, turning migration cost into an acquisition hook.
In the website-build demo, the stated order of operations — research first, then generate assets, then assemble the site around those assets, then deploy — is offered as evidence of how the agent "thinks," i.e., content-first rather than template-first.
The long-form video segment concedes, in the same breath as the pitch, that character/environment consistency across multi-clip sequences is "the hardest problem in this category" that "nobody in the industry has fully cracked," positioning Picsel Elements/Sol ID as having gotten "closer than anyone else" rather than as a solved problem.
«This is the first agent that does research, generates visual content, and runs marketing end-to-end.»
— 00:27
«Migration between agent platforms has always been a mess.»
— 00:57
«You describe the task. Supercomputer handles everything downstream.»
— 02:56
«LLMs alone don't know how to produce nonstop creative work. They guess at it.»
— 03:35
«One pattern we kept seeing in user feedback was cost. Users were burning credits on overpowered models running trivial steps.»
— 04:26
«The result is a 15x reduction in cost across the same workload.»
— 04:38
«Supercomputer isn't a tool. It's an entire marketing department running on one chat.»
— 05:08
«And that came from one simple prompt.»
— 06:20
«Long form has always been the hardest problem in this category.»
— 06:23
«This is the operating system for a new era of AI content.»
— 10:09
Reception
Audience recognizes the product's capability but is frustrated by high costs, lack of pricing transparency, and technical limitations like consistency issues.
This is a company-produced feature-launch video — Higgsfield's own head of product walking through three curated demos — so its capability and cost claims (100 UGC videos, a full deployed website, a 15x cost cut) are presented as fact but are not independently verified within the source itself.

10:44