The short version
- One style-reference image took an ad set from 2 of 3 on-style to 3 of 3, with no change to the prompt wording.
- Adjectives leave the model a wide region of style space to land in, and it lands somewhere different on every call.
- The reference carries the treatment and a different subject. Put your subject in it and the model copies the photograph you handed it.
- A re-run is a full call, roughly 25k to 28k tokens of overhead before any image exists. Removing one re-run pays for the reference immediately.
Attach a style-reference image to every generation call. Measured in our own production runs, prompts with no reference produced 2 of 3 on-style images. With one reference attached, 3 of 3. The prompt wording did not change between the two runs. The reference did the work the adjectives could not.
What is a style reference, and what does it replace?
A real image passed into the call alongside the prompt, carrying palette, film grain, lens character, lighting direction and framing. The model reads the look off the picture instead of off your adjectives.
It replaces the adjective pile. Cinematic, warm, filmic, natural light, shallow depth of field: each of those words covers an enormous spread of images in the training data, and the model lands somewhere different inside that spread on every call. The spread is why a set of four ads can look like four campaigns.
One ad set, same subject, same prompt language, run twice:
| Measure | No style reference | One style reference |
|---|---|---|
| On-style on the first pass | 2 of 3 | 3 of 3 |
| Images needing a re-run | 1 | 0 |
| Extra generation calls | 1 | 0 |
| What the reviewer does next | Works out which frame is off and why | Approves the set |
Why does one image beat a paragraph of adjectives?
Naming a specific film stock in the prompt does not rescue it. The name is still a word, and an attached frame pins the look to something that already exists.
What belongs in the reference image?
The treatment, and a subject other than yours. Hand the model a reference containing your subject and it copies the subject, which is how a campaign ends up as one photograph with three different captions. Pick a frame with the light and the colour you want in it, showing something else entirely.
Use one reference per set rather than one per image. That is what holds a set together across markets. A single generated set from that account ran in three markets, the US, the UK and Australia, and the consistency came from the shared reference rather than from anyone matching frames by eye afterwards.
How do you know whether it worked?
Somebody counts. Skip that step and the report is that the images look pretty good, which is not a report.
Grade every frame on the attributes the reference was supposed to carry: palette, grain, lens character, direction of light, and whether the subject does what the brief asked for. A beautiful frame that misses the brief is a failure and gets logged as one.
The numbers 2 of 3 and 3 of 3 exist because each frame was reviewed against the brief and the misses were recorded with a reason attached.
What does a re-run actually cost?
More than one image. Each call carries a fixed overhead of roughly 25k to 28k tokens before a picture exists, and each extra image in the same call adds about 5k. One image costs 28.7k. Three cost 38.7k.
Almost all of that cost belongs to the call rather than to the images inside it. Generate in batches rather than one frame at a time, and treat every re-run as a call you paid full price for. A reference that removes one re-run in three has already paid for itself.
The two production rules that sit next to this one are generating a device screen into the scene instead of pasting it on and what it costs to fight a model’s prior.
Frequently asked questions
What is a style reference image in AI image generation?
An existing image passed into the generation call alongside the prompt, so the model reads palette, grain, lens character and lighting from a picture rather than from adjectives. The reference sets the look. Subject and action stay in the prompt text.
Do reference images improve consistency across a set of AI images?
Yes, measurably. On one ad set run twice with identical prompt language, no reference produced 2 of 3 on-style images and one reference produced 3 of 3. The reference is also what holds a set together when the same creative runs across several markets.
What should you use as a style reference?
A frame that carries the treatment you want and a different subject from the one you are generating. If the reference contains your subject, the model copies the subject and the campaign becomes variations of a single photograph.
How many reference images should you attach?
One per set. Extra references dilute the target, and each additional image in a call adds roughly 5k tokens on top of the 25k to 28k fixed overhead. Consistency comes from every frame in the set sharing the same single reference.
Why do AI-generated images look inconsistent across a campaign?
Because adjectives describe a wide region of the model's style space and each call lands somewhere different inside it. Words like cinematic or filmic map to an enormous spread of training images. A shared reference image collapses that spread to one point.
How should generated images be reviewed?
Frame by frame against the brief, recording which frames miss and on which attribute: palette, grain, lens character, direction of light, or subject action. Counting is what turns a review into a number you can act on next run.
Sources
- Chua Network delivery data across 8 client accounts (internal fact bank)
- Chua Network engagement records, anonymized (internal experience bank)