Lore

AI image generation

Can ChatGPT 4o Image Generation REALLY Generate Product Photo Ads?

ChatGPT 4o's new image generation model is genuinely capable at product photography tasks — placing a product into a reference style/scene, replicating an AI-background compositing workflow, and generating graphical commercial templates — but it consistently lands at roughly 90-95% fidelity to the real product and its text/label, never a perfect 1:1 replica, so usability depends on the product's complexity and how much tweaking the user is willing to do afterward.

Thomas Lundström · 2025-04-03 · English

Key ideas

  1. The test is deliberately done with a smartphone rather than professional gear, to keep it comparable to what a small business owner could realistically do.

  2. Technique 1: give the model a reference photo (a styled scene, e.g. a yellow studio wall with a paper-breakthrough hand holding a can) plus a photo of your product, and ask it to place the product into that reference scene — tested on a yellow chocolate bar with a result the presenter called 'quite satisfied' with, text mostly accurate.

  3. Technique 2: try to replicate the presenter's own established workflow (create an AI background, then composite a real product photo into it) entirely inside ChatGPT by describing the scene directly — tested with a shampoo bottle in an underwater scene with bubbles, producing a 'quite impressive' first result that is described as close enough to finish by either taking your own photo or graphically replacing the label.

  4. Technique 3: feed it multiple photos of a product (an old shoe) and ask for a graphical commercial style with doodles and text around it — result wasn't a 1:1 match to the shoe but works as a background/graphic template into which the real product can later be placed.

  5. Harder test cases were chosen deliberately: products with more text and advanced graphical elements, since that's where the model tends to mess up.

  6. Tested a newly released Finnish energy drink with a Finnish-language label — the model got most of the text completely right, with only some details around 90% there.

  7. Tested a wine bottle specifically for its intricate label detail, expecting it not to be 100% accurate — result was close but with some label details changed.

  8. Overall conclusion: text, label, and product fidelity land around 90-95% of the original, never quite exact, and whether that's 'good enough' depends on the product and whether the user can tolerate doing further adjustments/edits.

  9. The presenter frames this video as a follow-up to an earlier 'first impressions' video, and teases an upcoming video testing the tool with professional gear on a real client photo shoot.

Insights

The most effective way to use the tool right now, per the presenter, isn't as a final-output generator but as a 95%-there template/base image for brainstorming, which is then finished by compositing a real product photo or graphically fixing the label.

Text/label accuracy appears to scale inversely with product complexity: simpler labels came back close to perfect while the wine bottle, chosen specifically for its intricate detail, showed more label detail changes.

Even 'imperfect' results are framed as practically useful because getting a near-finished base image removes most of the manual background-creation work, leaving only minor touch-ups instead of building a scene from scratch.

The model handled a non-English (Finnish) label with what the presenter called text 'completely right' in most places, suggesting the text-rendering strength isn't limited to English source material.

«You're going to notice in this video, usually it's like 95% there.»

— 01:46

«I have to say the first result is quite impressive.»

— 02:41

«I wouldn't take the results from Chat GPT as 100% usable yet, but they are like so so close to a finished product that you just need to tweak a couple of things and then you're pretty much there.»

— 03:08

«So I think using JGPT in this way for brainstorming and maybe creating like a 95% template will be probably the most effective way so far.»

— 04:01

«And I'm really really impressed that it got most of the text completely right.»

— 04:57

«And I especially wanted to see if it could handle this wine bottle because it has a lot of intricate detail.»

— 05:14

«So the conclusion we can make from this video is that Chad GPT's new image generation model is a really good but has some flaws.»

— 05:40

«the text and label and the product itself is 90 to 95% the same as the original but not quite there.»

— 05:47

Reception

The audience appreciates the helpful tutorial and tutorial approach, but raises valid concerns about AI accuracy limitations, product detail preservation, and rapid technological obsolescence.

A practical, non-hyped hands-on test that maps ChatGPT 4o's image generation capabilities and limits for product photography using only a smartphone, consistently landing on the same finding across five different products: impressive but not pixel-perfect, roughly 90-95% accurate.

06:58

↳ Thomas Lundström · YouTube

Watch original