AI Image Prompts
The difference between a generic image and a stunning one is the prompt. Here's the anatomy of a winning prompt — works in Midjourney, Stable Diffusion and every tool.
The anatomy of a winning prompt
A good image prompt is built from layers. The more specific layers you add, the more control you have over the result. The components:
- Subject: what's in the image. Specific — not "a dog" but "a young golden retriever."
- Action/context: what's happening. "running on a beach at sunset."
- Style: photography, oil painting, anime, 3D render, minimalist.
- Composition: close-up, wide shot, from above, rule of thirds.
- Lighting: golden hour, soft light, dramatic, neon — a huge influence on the feel.
- Camera/lens: shot on 85mm, shallow depth of field, cinematic.
- Color & mood: warm tones, moody, pastel.
- Quality: highly detailed, 4k, sharp focus.
Put the important things first — the model gives them more weight. Subject → style → lighting → details.
Example: before and after
Weak prompt:
a coffee shop
Strong prompt:
cozy specialty coffee shop interior, warm morning light through
large windows, wooden tables, plants, shallow depth of field,
shot on 35mm, cinematic, highly detailed, warm color palette
Same subject — a completely different result. The second prompt gives the model all the context: atmosphere, lighting, photographic style and color palette.
Negative prompts
In Stable Diffusion (and many tools) you can define what not to include — a negative prompt. Useful for removing common distortions:
Negative: blurry, low quality, distorted hands, extra fingers,
watermark, text, deformed, ugly
In Midjourney you use the --no parameter (e.g. --no text). It cleans up a lot of problems at once.
Parameters & aspect ratio
- Aspect ratio: in Midjourney
--ar 16:9(video/banner),--ar 9:16(story/Reel),1:1(post). Match it to the destination. - Weights: you can emphasize a word — in SD with parentheses
(word:1.3), in Midjourney with::. - Seed: save the seed to reproduce, or to vary a result you liked.
- Reference image: you can supply a reference image for style/composition.
Iteration — the real secret
Nobody gets the perfect image on the first try. The process: generate → pick the closest → change one element → repeat. Change one thing at a time (only the lighting, only the style) to understand what affects what. Save prompts that worked — they become your own templates.
What the model is actually doing with your words
Before the techniques, a mental model that makes the rest predictable — and that explains why some prompts fail in ways that feel unfair.
These models learned from enormous numbers of images paired with captions. So when you write a prompt, you are not issuing instructions. You are writing a caption for an image that does not exist yet, and the model is reconstructing what an image with that caption tends to look like.
Three consequences follow immediately, and together they account for most beginner frustration:
- Describe, do not instruct. "A cat sitting on a windowsill in morning light" works. "Make me a picture of a cat and put it on a windowsill" contains words that were never in a caption, and they dilute the ones that were.
- Negation is unreliable. "A street with no cars" contains the word cars, and captions saying what is absent are rare. You will often get cars. This is what the separate negative-prompt field exists for.
- Words that appear in real captions work best. The vocabulary of photography, art history and stock descriptions is what the model has seen most of — which is why "shallow depth of field" outperforms "blurry background".
Words that no longer do much
A lot of prompt advice is inherited from earlier, weaker models and has quietly stopped applying.
Strings like highly detailed, 4k, 8k, masterpiece, award-winning, sharp focus, trending on artstation were genuinely useful when models needed pushing towards competent output. Current models mostly produce competent output by default, and these terms now do one of three things: nothing, a mild stylistic shove towards a particular kind of digital art, or active harm by taking up space and attention that your actual subject needed.
The test is cheap and worth running once on whatever tool you use: generate with your full prompt, then again with the quality boilerplate removed, and compare honestly. Most people find the second is the same or better — and it is shorter, which makes it easier to iterate on.
Keep the words that describe something specific. Drop the ones that are just saying "please be good".
Why the red hat ends up on the wrong person
The most common failure with anything complex, and the one worth designing around because it is structural rather than fixable by better wording.
Ask for "a woman in a red dress next to a man in a blue suit" and you will regularly get a red suit, a blue dress, or both colours on one person. The model is matching a bag of concepts to an image and its grip on which attribute belongs to which object is weak. The same weakness produces wrong counts — "three cats" gives you two or five — and unreliable spatial relations, where "above", "behind" and "to the left of" are suggestions rather than instructions.
Four ways round it, in order of how well they work:
- Simplify the scene. One subject, one set of attributes. Compose multiple elements afterwards in an editor, which is also more controllable.
- Generate elements separately and combine them. Two clean images beat one muddled one.
- Use a reference image for composition, which communicates spatial arrangement far better than any sentence.
- Inpaint the correction — generate the scene, then mask the wrong element and regenerate only that region.
Rewriting the prompt a tenth time is the approach that feels like it should work and reliably does not.
Prompting for where the image will actually be used
A prompt written to produce a beautiful picture and a prompt written to produce a usable asset are different prompts, and this is where most generated images fail in practice.
An image destined for a banner, a thumbnail or a slide has to accommodate something else — text, a logo, a crop. If you generate a perfectly balanced composition that fills the frame, you will end up either covering the subject or shrinking it until the composition collapses.
So describe the negative space deliberately:
- "Subject positioned on the left third, plain uncluttered background on the right" — gives you somewhere for the headline.
- "Generous empty space above the subject" — for a title across the top.
- "Centred subject, wide margins" — survives being cropped square, wide and vertical from the same file.
- "Simple dark background" — because pale text on a busy photograph is unreadable at any size.
Generate at the aspect ratio you need rather than cropping later. Cropping a square composition to a wide banner removes the subject's head with impressive consistency, and a model asked for 16:9 composes for 16:9.
One more habit worth building: check the generation at the size it will be seen before accepting it. An image that is gorgeous at full resolution and unreadable as a thumbnail has failed at the only job it was given.
The controls that beat prompting
Most people stay in the text box far longer than they should. The larger gains are in the features that do not involve describing anything.
- Image-to-image. Supply a starting image and a strength setting. Low strength keeps the composition and changes the style; high strength keeps the idea and reinvents everything. This is the fastest route to "like that, but…".
- Style reference. Point at an image whose look you want and get that look applied to a new subject — far more precise than trying to describe a style in words.
- Inpainting. Mask a region and regenerate only it. Fixing a hand, removing an object, changing one item of clothing. This is how professional output actually gets made, and it is the single most underused feature in these tools.
- Outpainting. Extend beyond the original frame — useful when you need a wide crop from a square generation.
- Structural conditioning (ControlNet and its equivalents). Supply an edge map, a depth map or a pose and the output follows that structure exactly. This is how you get a specific composition rather than a lucky one.
- Seeds. The number that determines the starting noise. Same prompt plus same seed gives the same image, so keeping the seed and changing one word isolates what that word did.
A useful rule: if you have regenerated more than four or five times, stop prompting and start editing. The remaining distance is usually a mask away.
Getting the same look twice
For anything beyond a one-off image — a set of posts, a series of thumbnails, a brand — consistency matters more than any individual result, and prompts alone will not give it to you.
What works, from cheapest to most involved:
- A fixed style block. The same sentence describing lighting, palette, medium and framing, pasted verbatim into every prompt. Only the subject changes. Most of practical consistency is just this discipline.
- A style reference image reused across the set.
- Keeping seeds when you want variations that stay related.
- Character references, where the tool supports them, for a recurring person or mascot.
- A trained LoRA on your own images — the only method that reliably reproduces a specific character, product or house style, and worth the effort only when you will generate hundreds of images.
The things it still cannot do
Knowing the boundary saves an afternoon of trying.
- Text inside images. Better than it was and still unreliable, especially for longer words and any script the model saw less of. Do not let the generator set the words on anything you publish — add type in an editor.
- Hands, teeth and counting. Improved, not solved. Nothing is tracking how many fingers there should be.
- Precise likeness of a specific real person, which most tools also restrict on purpose.
- Exact brand colours. You get near, and near is not a brand colour. Apply real hex values in an editor.
- Technical accuracy — a wiring diagram, a real product's actual buttons, a working interface. It produces something that looks like one.
The tools read prompts differently
Prompts are not fully portable, and a prompt that works beautifully in one tool can come out flat in another.
Broadly: the more curated hosted tools apply a strong aesthetic of their own, so short prompts give attractive results and fighting that house style is work. Models that follow instructions more literally reward longer, more explicit descriptions and give you a plainer starting point. Open models run locally give the most control — negative prompts, weighting, structural conditioning, custom fine-tunes — and expect you to do more.
The practical advice: pick one and learn its behaviour rather than collecting prompts for all of them. Most of the skill here is knowing how a particular tool responds, and that knowledge does not transfer as cleanly as prompt lists imply. See the tool comparison for choosing.
Build your own library, structured
Saved prompts are the compounding asset here, and a flat list of them is much less useful than it looks.
Store them in pieces instead: a style block per project or brand, a set of lighting descriptions that produce the moods you use, composition phrases that reliably leave room for text, and negative prompts per tool. A new image becomes assembling blocks rather than writing from nothing.
Save the result alongside the prompt, plus the seed and the tool version. A prompt without its image is hard to evaluate six months later, and tools change — a prompt that worked in one version sometimes needs adjusting in the next, and you will only notice if you can compare. The prompt library is a starting point for the blocks themselves.
Ethics and rights, briefly
- Living artists' names in prompts are restricted by several tools' terms and are ethically contested. Describing the qualities you want — the palette, the brushwork, the era — gets you there without invoking a specific person whose work was used without their say.
- Real people's likenesses need their permission, and most tools prohibit generating identifiable public figures anyway.
- Commercial use depends on your plan, not on the tool. Check before publishing, not after.
- Copyright in purely generated images is unsettled or limited in several jurisdictions — relevant the moment a client wants to own something exclusively.
- Keep your prompts and dates for anything commercial. It is the only provenance record you will have.
A session that ends with something you can use
The usual pattern — type, look, retype, repeat, lose an hour — wastes the speed advantage. This takes about twenty minutes and produces a finished asset rather than a folder of near-misses.
- Write down what the image is for in one sentence, including where it will appear and what goes on top of it. This decides the aspect ratio and the negative space before you touch the prompt.
- Start short. Subject, style, lighting. Four or five elements, no quality boilerplate. See what the tool's default looks like before you start steering.
- Generate a batch and pick the closest, not the best. Closest to the composition you need beats prettiest, because composition is the hard part to change.
- Lock the seed and change one element. Lighting, then palette, then framing — one at a time, so you learn what each word did rather than which combination happened to work.
- Switch to editing once you are close. Inpaint the wrong hand, remove the stray object, extend the frame. This is where the last twenty per cent comes from and it is faster than regenerating.
- Finish in a design tool. Exact brand colours, type, logo placement. Never in the generator.
- Save the prompt, the seed and the output together, and add the reusable parts to your blocks.
The step people skip is the fourth, and skipping it is why prompting feels like gambling. Changing five things and getting a better image teaches you nothing you can repeat; changing one thing teaches you something you will use for a year.
When prompting is the wrong tool
- You need a logo. Logos need vector, and generators produce pixels. Explore concepts here, then have it redrawn.
- You need a real product photographed. No description produces your actual product.
- You need it exactly right. Diagrams, interfaces, anything technical — build it in the appropriate tool.
- One licensed stock photo would do. Sometimes the fifteen minutes of iteration costs more than the image.
Quick tips
- Write prompts in English whichever language you think in — these models saw far more English captions, and prompts in other languages are translated internally with losses you cannot see.
- Less is often more. Contradictory or padded prompts dilute. Cut before you add.
- Reverse-engineer images you like. Name the lighting, the lens, the palette, the medium — then reuse those words.
- Change one thing at a time, keeping the seed, so you learn what each word actually does.
- Start from ready-made prompts and adapt rather than writing from scratch.
Next step
Pick a tool and start creating. Compare the tools or dive into a specific guide.