The short version
- Compositing a screenshot onto a generated phone produced 0 usable frames. Generating the same shot in one pass produced 3 of 3.
- A composite fails on light rather than on craft: no glow on the hand, no reflection, two colour temperatures, one sharp edge sitting on grain.
- Composite only where the scene has nothing to give the screen, which means a straight-on device mock and never a lifestyle frame.
- When the model renders your product wrong, change the composition instead of the prompt. Fighting one prior cost 5 attempts on the same shoot.
Generate the screen into the scene. On an Australian taxi payments account we needed a product screen visible on a phone inside a vehicle. Compositing the real screenshot onto a generated phone produced 0 usable frames. Generating the whole shot in one pass produced 3 of 3.
Why do composited screens read as fake?
The pasted layer arrives with no relationship to the light around it. A screen inside a scene throws colour onto the fingers holding it, catches a reflection off the window behind it, distorts with the same lens as everything else in frame, and carries the same grain. A pasted rectangle has none of that.
Nobody itemises any of it. A more careful paste does not rescue it.
The same shot, a product screen on a phone inside a vehicle, produced both ways:
| What the frame needs | Generated in one pass | Composited afterwards |
|---|---|---|
| Screen glow on the hand and interior | Rendered by the model | Absent unless painted in by hand |
| Reflection and glare on the glass | Rendered | Flat, or a stock overlay that fits nothing |
| Colour temperature | One light source for the whole frame | Two, and the eye finds them |
| Perspective and lens distortion | Matches the virtual lens | Warped by hand, approximately |
| Grain and edge softness | Continuous across the frame | A sharp rectangle on a soft plate |
| Usable frames in our run | 3 of 3 | 0 |
What does one-pass generation give you?
The screen becomes an object in the light rather than a layer above it. The model renders the glow and the reflection because it is producing one photograph rather than two things stuck to each other, and the lens distortion comes along with it.
The trade is that you accept the model’s version of your interface, and it will be approximately your product. Teams balk at that, which confuses two different assets. A lifestyle frame exists to prove somebody uses the thing, in a moment a buyer recognises. Product accuracy is the job of a separate clean shot, and nothing is gained by asking one image to carry both.
When is compositing still correct?
When the pixels have to be exact and legible: regulated copy, or a pricing table a customer is meant to read line by line. At that point stop asking a lifestyle frame to carry it.
Shoot a clean device mock instead, straight on against a neutral background with no hand and no window in it. The composite holds up there because the scene has nothing to give the screen in the first place.
What do you do when the model renders the product wrong?
Change the composition rather than the prompt. Crop so the screen sits small in the frame, or angle it far enough off the lens that nobody tries to read it.
Fighting the model is expensive. Forcing a right-hand-drive interior on that same account took 5 attempts, because the prior for a car interior is left-hand drive. Every attempt is a full call.
A composite reads as a sticker because nothing in the scene lights it.
How do you decide before you burn a day?
Ask what the frame is for before anyone opens a tool. A feeling, a driver getting paid at the end of a shift, gets generated whole. Anything that has to be read gets shot on its own, clean, with no scene around it to match.
Cost points the same way. A generation call runs roughly 25k to 28k tokens with about 5k per extra image, so three images in one call cost 38.7k against 28.7k for one. Compositing spends a person on each frame instead, one frame at a time, and that cost does not fall when the set gets bigger.
The two neighbouring rules from the same production run are the style-reference test that took an ad set to 3 of 3 and the five attempts it took to beat one model prior.
Frequently asked questions
Should you composite a real screenshot into an AI-generated image?
Only when the screen content has to be exact and legible. On a lifestyle frame, compositing a real screenshot onto a generated phone produced 0 usable images for us, while generating the screen into the scene in one pass produced 3 of 3.
Why do composited screens look fake?
The pasted layer carries none of the scene's light. No glow on the hand, no reflection on the glass, no matching colour temperature, no lens distortion, and a sharp edge against grain.
When is compositing still the right method?
When the pixels must be readable: regulated copy, or a chart with numbers a customer has to check. Use a clean straight-on device mock rather than a lifestyle frame, because a composite only holds up where the scene has no light to match in the first place.
How do you get a product interface into an AI-generated scene?
Generate it in the scene, accept an approximate rendering, and control legibility through composition instead. Crop the screen small, or angle it off the lens so nobody tries to read it. Product accuracy belongs in a separate clean product shot.
How many attempts does one-pass generation take?
One-pass generation of a device screen inside a scene produced 3 of 3 usable frames on that account. The expensive case is fighting a model's prior: forcing a right-hand-drive car interior took 5 attempts on the same shoot.
Is generated creative good enough for paid campaigns?
Yes, when it is produced in one pass with a style reference and reviewed frame by frame against the brief. One generated set from that account ran across three markets, the US, the UK and Australia.
Sources
- Chua Network delivery data across 8 client accounts (internal fact bank)
- Chua Network engagement records, anonymized (internal experience bank)