This chapter collects the parts of the toolkit that live outside Photoshop and outside a local Flux setup: Midjourney's Omni Reference for holding a character's look steady across generations, Adobe Firefly's split between a style reference and a composition reference — and the manual Photoshop trick needed to feed that composition slot a shape Firefly can actually read as 3D — plus one pure prompting method, the juxtaposition technique, that makes a model invent a story premise instead of illustrating one you already wrote. Each is presented with the limitation its source actually reported: Midjourney accepts only one reference image at a time, and Firefly's 3D-render pipeline does not preserve a real person's facial likeness. The chapter is deliberately narrow, and the closing section says plainly which tools named in the subject's goal it never reaches.
The only technique here that belongs to no particular tool is the Juxtaposition Technique (Story Generation): put two opposite or incompatible character archetypes into a single image prompt — a punk and a sophisticated old woman, a monster and a nun, a vampire and a priest — and ask the generative model to supply the title, the description, and the narrative connection between them. You are not authoring a premise and handing it over to be illustrated. You are staging a collision and letting the model do the writing.
The reason it works, as the technique is documented, is that genuine opposition is a concrete creative constraint. The model has to resolve or explain the incongruity it has been handed, and that pressure tends to produce a richer, more specific premise than an open-ended "invent a story" prompt, which gives the model nothing to push against. The specificity comes from the constraint, not from the instruction to be creative.
A juxtaposition is also a seed rather than an endpoint. It can be expanded into a large grid of narrative beats, which makes visible how a story arc rises over time, or condensed the other way into a small storyboard grid of filmable scenes. This form of the technique was first documented via Olivio Sarikas's ChatGPT Image 2 + Seedance 2.0 prompting guide, filed 2026-08-06.
Worth being clear about the scope: everything documented here is about generating story premises. Nothing in the material extends the juxtaposition idea to editing an existing photograph, which is what the rest of this book is mostly about.
Midjourney Omni Reference (Character Consistency) is Midjourney's character-reference feature, and the mechanics are as direct as they sound: drag a reference image straight into the prompt bar — no download-and-re-upload round trip — write a new prompt, and Midjourney generates a new scene in which the character resembles the reference. Consistency across generations is the whole point.
The technique that makes it behave has less to do with the reference image than with how the character was designed in the first place. Unique-look prompting means giving the character one or two distinguishing, describable traits — pink hair, a specific jacket color, a particular body shape — so the model has something concrete to reproduce. A character defined by nothing in particular gives the model nothing in particular to hold onto.
Two practical habits follow from that. For full-body shots, explicitly prompt for footwear and add "standing" to the prompt: left to itself, Midjourney tends to seat the character or simply omit shoes. And for wardrobe changes across a story, feed the same Omni Reference image back in while prompting the new outfit — a wedding dress, say — rather than re-describing the character from scratch each time. The reference carries the identity; the prompt carries only what changed.
This material comes from a single source, the video 'Achieve Perfect Character Consistency in Midjourney with These Tips'.
The hard ceiling on Midjourney Omni Reference (Character Consistency) is that Midjourney, as of this writing, supports only one Omni Reference image per generation. That is a real problem the moment a story needs two established characters in the same frame.
The workaround is unofficial and it works imperfectly: combine both characters into a single reference image first, then use that composite as the one allowed reference. Details leak between the two characters — one character's leather shoes bleeding onto the other is the reported example. The same combined-reference trick can also be aimed at one character's clothing inside a two-character image, and it lands in roughly the same place: usable but not exact, with a designed skirt coming out longer than intended.
A separate lever is worth knowing because it has nothing to do with what you wrote. Aspect-ratio switching — regenerating the same prompt at a different aspect ratio — acts as a general-purpose fix for broken composition, wrong elements, and anatomical or proportion errors, and is reported as especially effective on non-human subjects. It is a fix independent of prompt content, which makes it the first thing to try when re-wording has stopped helping.
There is also a chaining move that treats Midjourney as one station rather than the whole shop. Rough Midjourney character design sheets are sketch-quality and valued as their own aesthetic, but they can be cropped, described via ChatGPT, and re-rendered with more surface detail through Krea's hosted Flux model — explicitly not aiming to preserve 100% fidelity — before coming back into Midjourney as a refreshed reference for further variation. Note that this is Flux used as a hosted detail pass on someone else's terms; running Flux yourself, with control over samplers and LoRAs, is the subject of Flux & ComfyUI: Local Model Tuning.
Adobe Firefly exposes two reference inputs that can be set from two entirely different images, and Firefly Style Reference & Composition Reference is really about what that independence buys you: the ability to explore material, lighting, and mood freely without ever touching the shape you have committed to.
The style reference is set by marking a previously generated or uploaded image via its pencil icon and choosing "use as style reference"; it then governs material, lighting, and mood for subsequent generations. The recommended order of operations matters here. Generate a handful of options from a pure text prompt that describes only the desired material, environment, and lighting — no shape involved at all — then pick the best one and lock it in as the style reference before you go anywhere near composition. Style gets settled first, on its own terms.
The composition reference is where a shape or silhouette image gets uploaded so the generated output follows that structure. Alongside it sits a strength slider controlling how closely the output must match what you uploaded. Push it all the way up when the result has to stay faithful to an exact shape — readable text, a specific silhouette — and lower it when some AI-driven drift away from the shape is acceptable.
Two operational facts round this out. Results can be regenerated indefinitely, in the source's phrasing, by clicking generate after generate after generate; and any result can be upscaled to higher resolution at the cost of one additional credit. There is one detection quirk that shapes the entire next section: Firefly's shape/composition detection appears to treat touching same-colored regions as a single merged blob.
The merged-blob behaviour has a direct consequence. A flat, single-color 2D shape — which is exactly what typed text and letterforms are — produces a flat-looking Firefly result, because adjoining same-colored regions get read as one shape. The Photoshop Fake-3D Silhouette Technique is the manual Photoshop answer: fake the depth information in the reference image itself so the detector has something dimensional to find. Simple shapes like emoji don't need it and can be uploaded directly.
The steps are short and mechanical:
The conceptual inversion is the part worth carrying away. Those contrasting colors are shape metadata for a detector, not an aesthetic choice. This is Photoshop used to make an image machine-readable rather than to make it look good, which is the opposite of everything the Photoshop chapters of this book are about.
Two related fixes travel with the technique. If Firefly generates a texture — sand, in the reported case — that fails to visually merge with the shape, painting a solid color block over that area in Photoshop before re-exporting helps the material apply correctly, implying Firefly's material application can fail to saturate a shape unless the reference is solidly filled in that zone. And background elements such as a horizon or floor are built with the Rectangular Marquee tool plus a Solid Color adjustment layer above the background layer, so the floor color stays editable at any point — the same non-destructive habit that Photoshop Foundations: Non-Destructive Editing & Selection establishes for editing work generally.
Assembled, those pieces form the Photoshop-to-Firefly 3D Render Pipeline: a way of turning any 2D shape — an emoji, typed text, or a photograph — into a realistic, dramatically-lit 3D-looking render. The defining choice is that the shape gets built in Photoshop first rather than uploaded to Firefly directly, because doing the shape-build yourself preserves full control over framing, styling, and environment, all of which a direct upload gives away.
The chain runs: build or export the shape as a flat PNG in Photoshop → set the Firefly style reference independently, from a pure text prompt about material, lighting, and mood with no shape involved → upload the Photoshop shape as the composition reference → push the strength slider up to lock the output to that shape → generate repeatedly, and optionally upscale the best result for one credit. Simple flat shapes like emoji go in as-is; text and letterforms need the fake-3D silhouette pass first, or they come back flat.
The pipeline was also tested on an actual portrait photo to produce a "stone statue" effect, and this is where it stops. It does not reliably preserve facial likeness on real people. It works well for generic human or animal statues and for objects where exact likeness isn't the goal — and this is reported directly as a limitation, not as an edge case someone could engineer around. If your subject is a specific person whose face has to survive the edit, this is the wrong tool, and the manual approaches in Manual Compositing Craft: Light, Color & Depth or Photoshop's own generative stack in Generative Fill & AI Compositing Toolkit are the places to look instead.
One access note that lowers the barrier considerably: even free versions of Photoshop grant 25 free Firefly credits, which is enough to run the full loop — generate, pick a style, build the shape, composite, upscale — without a paid plan.
This chapter is thin in places, and it's more useful to say so than to imply coverage that isn't there.
The subject's stated goal names ChatGPT's image model, Lovart, and mobile editors among the tools worth tracking. None of their mechanics appear here. ChatGPT surfaces twice, and only as a helper: describing a cropped Midjourney design sheet inside the Midjourney Omni Reference (Character Consistency) chaining workflow, and as the host of the guide where the Juxtaposition Technique (Story Generation) was documented. It is never given a prompting technique of its own. Lovart does not appear at all. Seedance 2.0 is named once, purely as provenance for that same guide, with nothing said about how to use it.
The tool coverage that does exist is narrow by source. Midjourney here is one video's worth of material, entirely about Omni Reference and character consistency — nothing on the rest of the tool. Firefly here is entirely the 3D-render pipeline built on Firefly Style Reference & Composition Reference; there is nothing about using Firefly for ordinary photo editing.
And the "prompting techniques" half of this chapter's title rests on a single technique, one that generates story premises rather than editing photographs. Readers looking for prompting craft aimed at the phone-photo pipeline this book is otherwise built around will find more of it in the tool-specific chapters — the local generation side in Flux & ComfyUI: Local Model Tuning, the everyday retouching side starting at Mobile Baseline: iPhone Photos App & Lensa — than here.