Lore

Generative Fill & AI Compositing Toolkit

Из Read: AI Photo Editing

This chapter is about the AI compositing path in Photoshop: swapping a background with Generative Fill, matching the subject to it with Harmonize, and pushing resolution with Generative Upscale — plus the selection strategies (precise mask, whole-image conversational prompt, or a marquee-preserved patch of real photography) that decide what the AI is allowed to touch. The organising claim is that these tools get a composite most of the way but not to the finish line: edges, background resolution, and texture mismatch remain, and the chapter's escalation order and finishing stack — tilt-shift blur, a three-layer grading stack, and Camera Raw grain — are what close the gap. It also covers the hard pixel ceilings on AI output and the two ways around them.

Reach for the free path before the metered one

Before any of the compositing craft in this chapter, there's a habit worth installing: when Photoshop offers both a classical and a generative route to the same result, try the classical one first. The Remove Tool is the clearest case. It has an AI toggle, and with generative AI switched off it falls back to traditional content-aware fill — which costs nothing. Only when that result isn't good enough do you flip the toggle on, and then you're spending: one credit under Adobe's Firefly-model pricing as of 2026-08, returning a single variation on a new layer. That's the whole of Remove Tool: Non-AI-First Escalation (Content-Aware Fill Before AI), and its value doesn't depend on the price staying at one credit, since Adobe flags its own credit pricing as subject to change.

The underlying reason is structural rather than commercial. Detection and selection operations — finding a subject, building a mask — are algorithmic and cheap, so they're free. Generating pixels is expensive, so it's metered. Once you can feel that distinction, the escalation order stops being a rule to memorise and becomes obvious: exhaust the analysis-shaped tools, then pay for the generation-shaped ones.

Generative Fill has its own removal mode, and it sits one rung up the same ladder. Lasso just outside a leftover remnant or a messy patch, open Generative Fill, leave the prompt field blank, and click Generate — Photoshop reads the empty prompt as removal intent and fills the area contextually. That's Generative Fill with Blank Prompt (Content-Aware Erase). It isn't free the way content-aware fill is; cost depends on which model you've chosen. What it buys over the Remove Tool is context: it fills against whatever now surrounds the area, which matters most when the surroundings are themselves AI-generated rather than original photograph. That makes it the natural cleanup pass after a background swap, not a first resort.

The rough first draft: what a Generative Fill background swap really delivers

The spine of this chapter is Generative Fill Background Compositing Workflow, and its core claim is deliberately two-sided. Generative Fill gets a background swap most of the way — it matches lighting, color, and perspective to the subject on its own. But it botches edge quality and background resolution. So the output is a rough first draft, not a finished composite, and everything else in this chapter exists to take that draft the rest of the way.

The sequence runs: isolate the subject with Select Subject, then invert the selection using the Contextual Taskbar's invert button so the background is what's selected. Prompt Generative Fill with the new scene — the worked example is as plain as "Green Mountains" — generate multiple variants, and pick one. Clean up leftover subject remnants with the blank-prompt erase described above. Then hide the AI background's low resolution with Tilt-Shift Blur to Mask Low-Resolution AI Backgrounds, re-mask the subject cleanly (Select Subject or Quick Selection into the Mask button, refined with a soft brush or the Pen Tool, or Load Selection to reuse a precise selection you saved), apply the grading stack in Post-Generative-Fill Color Grading Stack (Light Flare + LUT + Clipped Curves), and finish with Camera Raw Grain to Unify AI Composite Texture.

The last step of the workflow isn't a tool at all: step away from the image and come back with fresh eyes as a deliberate QC pass. Unnatural highlights and artifacts that were invisible during the build become obvious on return, and they need manual correction. Treating this as a numbered step rather than a nicety is the point — the AI passes don't flag their own failures.

Two limits are stated flatly and are worth knowing before you promise anyone a result. Generative Fill cannot reproduce a specific real place — ask for Machu Picchu or the Taj Mahal and it invents something unique instead. And generating an entire background in one pass is currently low quality; the practical workaround is generating box by box rather than all at once. The selection-to-mask and non-destructive plumbing this workflow leans on is covered in Photoshop Foundations: Non-Destructive Editing & Selection, and the fully manual counterpart — building a composite back-to-front and hand-matching light at every layer — is Manual Compositing Craft: Light, Color & Depth.

Three ways to decide what the AI is allowed to touch

The workflow above assumes you select precisely and then prompt. That's one of three approaches in this chapter, and they differ mainly in where the boundary between original and generated pixels gets drawn.

The second is Whole-Image Select + Conversational Prompt (Generative Fill Partner Models): skip per-area selection entirely, hit Cmd/Ctrl+A to select the whole image, and describe every intended edit — removals, additions, a time-of-day or season shift — in one natural-language prompt. This is enabled by Photoshop's partner models, Google's Gemini 2.5 Nano Banana and Black Forest Labs' Flux Context Pro, which sit alongside Adobe's own Firefly Image Model 3 and Image Model 1 in the Generative Fill model picker and can be swapped from the Properties panel at any time without redoing the selection. Partner models appear to reason over a whole-image selection well enough to infer which regions to touch from the prompt alone, which is what removes the manual brushing step that Adobe's native Firefly models still benefit from.

The practical notes on that technique are more interesting than the technique itself. Partner models render one variation per click; Firefly Image Model 3 renders three per generation, which changes the arithmetic on credit spend. Quality isn't uniformly won by one model — in one comparison Nano Banana most often preserved original framing and scene fidelity, while Flux Context Pro won outright on the most complex multi-edit prompt and was sometimes judged more photorealistic. The recommended practice is therefore to run the identical prompt through multiple partner models and pick, not to settle on a favourite. Even the stronger model can leave artifacts on complex prompts — one first pass left some wires behind and needed a second generation. Every generation consumes paid credits whether or not the result is usable, and Photoshop's UI doesn't surface that cost. Finally, an explicit boundary: this is a creative and compositing technique, not a documentation-accuracy tool. Don't use it where representation carries liability, such as removing permanent fixtures from real-estate listing photos.

The third approach draws the line by hand. Selective AI Background Replacement (Marquee Mask-Fill Preservation) keeps part of the original background as real photography while the rest gets generated. Run Remove Background to cut the subject onto a layer mask, duplicate that layer and re-run Remove Background on the duplicate, then use the Rectangular Marquee to select the region that must stay literal — the floor a subject is standing on, say. Target the layer mask and fill that selection with white, restoring the original pixels there. Now generate a new background by prompt: only the unselected area receives AI content. Optionally run Upscale afterwards to sharpen the generated portion. What this gives you is a controllable, visible boundary between real and generated pixels inside a single mask, which is exactly what you want whenever a product, a floor, or a pair of hands has to stay literal while the scene around it is reimagined.

Harmonize, and why the order of operations matters

When the subject comes from a different photograph rather than staying put while the background changes, the matching problem gets harder — and Harmonize: AI Lighting/Color Match for Composites (Firefly Model 3) is the one-click AI answer. Powered by Adobe Firefly Model 3, it matches the lighting and color of a composited subject layer to its new background, returning three variations per generation and consuming generative credits.

It appears automatically in the contextual taskbar for any layer that has a layer mask, and it's also reachable via Layer > Harmonize or by searching 'harmonize' in Photoshop's Help. The mask requirement shapes the workflow: drag the subject onto the background, run Remove Background to auto-mask it — which is what unlocks Harmonize in the first place — position and size the subject, then run Harmonize. As of Aug 2026 the tool was newly out of beta, and single-shot output isn't always convincing, so the guidance is to generate one or two extra passes before accepting a result.

The part worth internalising is the chaining behaviour. Harmonize's lighting match persists through later edits. Run a Select All plus Generative Fill with a partner model afterwards to change clothing or content, and the matched lighting survives — provided you keep the order mask → harmonize → content edit. Reverse that and you're asking a content edit to preserve a match that hadn't been made yet. This is the kind of sequencing constraint that doesn't announce itself in the UI, and getting it wrong costs a full round of credits.

Blur, grade, grain: the three passes that sell the composite

Everything so far produces an image where a sharp, real subject sits against a generated background that is lower-resolution, smoother, and lit approximately right. Three finishing passes close that gap, and each one works by leaning into a limitation instead of fighting it.

First, blur. Tilt-Shift Blur to Mask Low-Resolution AI Backgrounds uses Blur Gallery – Tilt-Shift to hide the AI background's weak resolution and messy detail, on the reasoning that a background behind a subject was usually meant to be out of focus anyway — and the source photo was often already slightly soft there. Stamp visible layers into a new one with Ctrl/Cmd+Alt+Shift+E, Convert for Smart Filters so the blur stays editable, then position the solid in-focus band around the subject and drag the dashed lines out over the background. Around 40px of blur is enough to obscure the low resolution without reading as artificial. The elegance here is that one limitation becomes cover for another rather than something you have to eliminate.

Second, the grade. Post-Generative-Fill Color Grading Stack (Light Flare + LUT + Clipped Curves) is three layers, in order. A Radial Gradient adjustment layer acting as a light flare — white or bright-yellow core fading through reddish/orange to transparent, blend mode Screen so it only ever brightens — repositioned and scaled to sit where the scene's light source is. A Color Lookup adjustment layer on the "Crisp Warm" LUT preset for the overall grade, dialed back to roughly 50% opacity and further tamed by Alt/Option-clicking to split the underlying-layer tonal-range slider so the grade doesn't crush the darks. And a Curves adjustment layer clipped to the subject layer only, darkening, adding contrast, and pushing blue and green down toward warm/yellow wherever the subject reads unnaturally bright — brushed in with a soft white brush, typically to fix what the fresh-eyes pass caught. The stated attitude toward that last layer is worth quoting in spirit: it isn't about restoring strict lighting accuracy, because a fabricated adjustment sometimes reads as more realistic than the literal, unedited truth.

Third, grain. Camera Raw Grain to Unify AI Composite Texture adds photographic grain through Camera Raw Filter's Effects panel — Amount, Size, Roughness — on a final merged or stamped layer, to make the sharp real subject and the smooth AI background share a texture. The mechanism is a happy side effect: raising grain size in Camera Raw also introduces a slight blur, which would normally be a defect but is exactly how grain behaves in old and vintage photographs, so it naturalises the fidelity mismatch rather than exposing it. This step isn't an AI workaround at all, and the chapter says so explicitly: PHLEARN's remade panorama-compositing course — fully manual, zero AI — teaches the same mandatory last pass, matching resolution and grain across every layer of a dozens-of-photographs composite so the result "looks like it was shot on one camera." The hand-built version of that discipline is the subject of Manual Compositing Craft: Light, Color & Depth.

Working against the pixel ceiling

AI passes in this toolkit have output limits, and two concepts here are entirely about routing around them.

Generative Upscale 6144px Output Cap (Downsize-First Workaround) covers Photoshop's Generative Upscale (Image > Generative Upscale, powered by the Firefly Upscaler), which enlarges 2x or 4x and invents plausible missing detail rather than stretching pixels — markedly better than a plain non-AI resize. The hard ceiling is 6144px on the long edge. Attempt 2x on a 4,000×6,000 source, which would land at 8,000×12,000, and it throws 'output is too large.' The mechanical workaround is to shrink the source first via Image > Image Size so the scaled result lands under the cap — downsize to roughly 1,000×1,500px before a 4x pass, for instance — with an ideal starting long edge somewhere around 2,000–4,000px, though smaller works; you back-calculate from 6144 against the scale factor you want. The caveat is emphatic and easy to miss: that downsize-then-upscale demonstration exists to illustrate the limit, not as a recommended workflow. Don't shrink a large image just to blow it back up. The genuine use case is upscaling small originals — scanned or photographed old family photos. One more gap: partner models aren't yet available for Generative Upscale, unlike Generative Fill.

The second is Crop-and-Recombine AI Retouch Realignment (Difference Blend + Masking), for when an external generator's output resolution sits below your source photo's native resolution. Crop the photo down to just the region needing AI work — ideally to a square 1:1 crop, since many generators, Nano Banana among them, behave most consistently on square inputs — run the edit on that crop, then recombine with the full-resolution original. It also lowers per-generation cost on credit-metered tools, since effective resolution drops.

The recombination is the clever part. Bring the AI-edited crop in as a smart object above the original and set its blend mode to Difference: misaligned areas glow, perfectly aligned areas go black, so alignment becomes a visual signal rather than a guess. Free Transform (Ctrl/Cmd+T) with an Alt/Option-clicked anchor point, scale and nudge until the overlap goes black, then switch back to Normal. Edit > Auto-Align Layers is faster but requires rasterizing first, sacrificing the smart object. Finally, Alt/Option-click the layer mask button to add a black hiding mask and paint with a soft white brush to reveal only the AI regions you actually want.

The reasoning behind all that masking care generalises well beyond this one technique: generative models produce an entirely new set of pixels rather than editing existing ones, so a result can look convincing and still be subtly wrong — an eye that no longer reads as the same eye. Because the mask is selective and reversible, one black brush stroke returns any specific wrong region to the original, with no need to discard or regenerate the whole thing. Note the contrast with the blur approach earlier: tilt-shift solves the same underlying problem — AI content that doesn't match the original's fidelity — by hiding it, while this solves it by realigning and selectively revealing.

What this chapter doesn't tell you

Read together, the material here is opinionated about order — cheap before metered, mask before harmonize before content edit, blur before grade before grain — and much vaguer about judgment. The recurring instruction at every decision point is some version of "if it isn't good enough, escalate" or "generate an extra pass or two," with no stated criterion for good enough. That's not an oversight to paper over; it's where the craft lives, and this chapter locates it in the fresh-eyes QC pass rather than in any threshold you could write down.

A few things are genuinely absent. The credit arithmetic is partial: one credit for an AI Remove Tool pass as of 2026-08, three variations from Firefly Image Model 3 versus one from a partner model, and an explicit note that Photoshop doesn't surface generation cost in its own UI — but no worked budget for a full composite. The 6144px ceiling is documented for Generative Upscale, but nothing here states what resolution a Generative Fill background actually comes out at, which is odd given that its low resolution is the stated reason the tilt-shift blur exists at all.

The scope is also narrower than "AI compositing" in general. Everything above is Photoshop plus its partner models. The node-based, local side of generation — LoRAs, samplers, tuning a Flux generation deliberately — is a separate track in Flux & ComfyUI: Local Model Tuning, and Midjourney character consistency, Firefly's 3D-style reference tricks, and prompting technique proper live in Beyond Photoshop: Other Tools & Prompting Techniques.

Открытые вопросы

Концепты

Источники