Lore

AI image generation

Nano Banana INSANE Product Photos (TUTORIAL)

In this hands-on first-look review, Thomas Lundström argues that Nano Banana (Google's Gemini 2.5 Flash image model, free at gemini.google.com) is a major leap for AI-generated product photography, because it can reproduce product labels almost pixel-perfectly across radically different scenes, generate camera angles not present in the source images, and be iteratively refined through plain conversational requests instead of manual editing.

Thomas Lundström · 2025-08-27 · English

Key ideas

  1. Nano Banana is Google's new AI image generation model, also known as Gemini 2.5 Flash, free to use at gemini.google.com.

  2. The test method: take existing product photos, remove the background, then feed the bare product plus a reference image and a text prompt to the model.

  3. Products were deliberately chosen with varying label difficulty (graphical labels vs. the Prime bottle's simple text) to stress-test how faithfully the model reproduces labels.

  4. Test 1 (hot sauce flat lay among chicken wings): label kept intact; rated 'seven or eight out of 10,' with lighting flagged as could be more dynamic.

  5. Test 2 (hot sauce bottle half-submerged in orange sauce, based on a reference of a tube floating in water): label again intact, and the model added realistic sauce residue on the bottle cap unprompted; called something 'absolutely impossible to do in Photoshop or anywhere else.'

  6. Test 3 (novel camera perspective — a hand lifting the bottle out of the sauce at surface level, an angle not present in any input image): called the best single AI-generated image the creator has ever produced, with label near-100% intact and 'perfect' studio lighting.

  7. The model doubles as a conversational image editor: an unwanted splash under the hand was removed simply by asking, without any manual editing.

  8. Test 4 (coffee bag): initial result judged too flat; lighting was changed to 'more dramatic' purely by text prompt, framed as a game changer for people without editing skills.

  9. Test 5 (Prime bottle, a stress test): a chain of sequential edits — dramatic lighting, then softened back down, then blue raspberries added on the glass, then raspberries added inside the ice — to see how far repeated changes could push the model before it broke.

  10. Test 6 (YoPro bottle): matched a reference image's tilted camera angle and swapped the background to blueberries and blueberry bushes; rated 'almost 100% perfect' aside from slightly warped label text.

  11. Test 7 (Crisp beer can): kept the exact style/composition of an existing image while swapping in a different product and changing the background to dark blue.

  12. Head-to-head comparison: the same task given to ChatGPT failed to reproduce the label correctly, while Nano Banana's result was near-perfect, described as 'completely night and day.'

Insights

The model can synthesize a genuinely new camera perspective (e.g., a surface-level shot with a hand pulling the bottle from the liquid) rather than just recombining elements already visible in the source images — the creator specifically sought this out after seeing it discussed online.

It adds physically plausible incidental details unprompted, such as sauce residue on the bottle cap, which reinforces the illusion that the product was actually submerged rather than composited.

Because changes can be requested in plain language, the tool collapses tasks that normally require Photoshop skill (relighting, removing unwanted elements, mood changes) into a single conversational turn.

Deliberately chaining several edits in a row (the Prime bottle test) reveals a degradation risk: after multiple rounds, 'some of the details of the image is a bit distorted' and a resulting harsh highlight was judged 'maybe not optimal,' indicating edits are not fully lossless over iterations.

Label/text fidelity is used as the specific axis for the competitive claim against ChatGPT rather than composition, lighting, or realism — it's the one dimension singled out where the gap is called 'night and day.'

«And is it the game changer that everyone is talking about?»

— 00:07

«For a first try, I think this is really impressive.»

— 01:34

«And this would be absolutely impossible to do in Photoshop or anywhere else.»

— 02:31

«Uh, this I think is the best single image I've ever generated with AI considering it's keeping the label almost 100% intact.»

— 03:14

«Mind blown.»

— 03:33

«The fact that you can change the lighting in this way by simply talking to the model that is so powerful for people that don't have the editing skills of editing images.»

— 04:16

«Now, this one again had my jaw dropping because it's almost 100% perfect.»

— 06:33

«So, to summarize my initial reaction, I'm very very impressed by this new image model.»

— 07:16

«Chad GPT doesn't even get the label right on the image. And Nanobanana is almost perfect.»

— 07:44

Reception

Strong positive reception with enthusiastic praise for specific examples and creative applications, tempered by some technical frustrations with tool limitations in text handling and image modifications.

The video is an enthusiastic, informally-scored hands-on demo rather than a rigorous benchmark — judgments like 'seven or eight out of 10' are impressionistic and the competitor comparison rests on a single ChatGPT example — but its label-fidelity and iterative-editing claims are each backed by a specific before/after example shown on screen.

08:13

↳ Thomas Lundström · YouTube

Watch original