How do I use reference images in ChatGPT?
Updated July 2026
Ask an AI about this guide
Open this article in your assistant with a ready-made question about doing the same job on your own API keys.
Reference images are how you get ChatGPT to generate something specific - your product, your character, your visual style - instead of a generic interpretation. The app supports this well for single images; this guide covers how to do it and where it gets awkward at volume.
Uploading and prompting with a reference
Attach one or more images to your message using the paperclip (or drag and drop), then describe what you want in relation to them. GPT Image models genuinely use the attached image as visual input - this is not just caption-and-regenerate - so identity, layout, and style transfer work reasonably well.
The phrasing matters. "Make this product on a beach at sunset, keep the label exactly as shown" outperforms "beach version of this". Be explicit about what must stay fixed (logo, colors, face) and what can change (background, lighting, angle).
- Use one clear, well-lit reference rather than five mediocre ones when identity matters.
- For style transfer, attach the style example and say "in the style of the attached image" plus a text description of the style - the text reinforces the image.
- For edits, ask ChatGPT to change only the named element; unprompted regions can still drift slightly.
What works well and what drifts
Product shots, style matching, and background swaps are strong. Faces and fine text are weaker: run the same reference ten times and you will get subtle identity drift and occasional mangled label text. Each generation is independent, so consistency across a series is a matter of luck plus careful prompting, not a guaranteed feature.
The limits: one conversation, one image at a time
The bigger constraint is workflow. Every variation is a new message in a chat, generated sequentially, subject to the same rate limits as everything else (users report roughly 50 images per 3-hour window on Plus). There is no way to say "this reference plus these 40 prompt variations" and get a grid back. If you are producing a catalog or a campaign from one reference, you will spend most of your time scrolling a conversation and re-attaching context.