How to Prompt an AI Photo Editor
Most people write editing prompts as if they were generation prompts, and that single mistake explains almost every bad result. A generation prompt describes the whole image you want. An editing prompt describes one change and explicitly protects everything else. They are opposite jobs: in generation, more description gives you more control; in editing, more description gives the model more permission to redraw things that were already correct. The practical form of an edit prompt is two halves — the change, then the preservation clause — and the second half is the one almost nobody writes.
Generation describes a frame. Editing describes a delta.
When you generate, the canvas is empty. Nothing is right yet, so every noun you add — the subject, the light, the lens, the mood — buys you a decision you would otherwise leave to the model. That is why generation guides tell you to be specific and why a three-line art-director brief beats a pile of keywords.
When you edit, the canvas is already right — except for the one thing you want changed. The photo has a real person in it with a real face, real hair, and real light falling on them from a real direction. Your job is not to describe that person — it is to leave them alone. Every noun you mention is a noun the model now considers in play. Say "professional woman in a bright modern office" over a photo of an actual woman and you have not asked for a background swap; you have asked for a new picture that happens to resemble the old one.
The inversion also shows up in length. A generation prompt keeps earning its keep as it grows, because every clause you add is one more decision taken away from the model. An edit prompt stops earning at roughly two sentences: the change, then what to protect. If yours runs to five, the extra three are usually either a second pass in disguise or a description of things that were already correct.
Every edit prompt has two halves
Write the change first, in one clause: a verb, a target, and an outcome. "Replace the background with a grey studio wall." "Change the shoe upper from blue to red." "Remove the folding chair on the left." One verb. One target. If you need the word "and" twice, you are writing two prompts and should send two passes.
Then write the preservation clause: the specific things that must survive the edit, named as nouns. Not "keep the rest the same" — the actual list. The reason this half feels redundant and is not: the model has no idea which pixels are load-bearing. It cannot tell that the freckles are the point and the lamp in the corner is not. You are the only one who knows, so you have to say.
| What you want | Weak prompt (reads like generation) | Strong prompt (reads like an edit) |
|---|---|---|
| Swap the background | Professional woman in a modern open-plan office, natural light, 50mm | Replace only the background with a bright open-plan office. Keep the woman, her pose, hair, clothing and the direction of light falling on her exactly as they are. |
| Change a product colour | Red running shoe on a white background, studio product photo | Change the shoe upper from blue to red. Keep the mesh weave, stitching, sole, laces, logo placement and contact shadow unchanged. |
| Remove an object | Clean empty living room, minimal, tidy | Remove the folding chair on the left. Continue the hardwood floor and skirting board already visible behind it. Leave the sofa, rug, window and lighting untouched. |
| Retouch a portrait | Flawless skin, glamour retouch, beauty magazine finish | Reduce the shine on the forehead and remove the three blemishes on the left cheek. Keep skin texture, pores, freckles, moles and the existing makeup. |
| Relight a shot | Golden hour, cinematic lighting, warm and moody | Warm the key light and move it to camera left. Update the shadows and highlights on the subject to match that direction. Keep the subject, framing and background as they are. |
| Edit text on a label | Coffee bag packaging design, 500g | Change the label text from 250g to 500g. Keep the same typeface, weight, colour, letter spacing and every other element of the label identical. |
Read the right-hand column again and notice how boring it is. No adjectives about quality, no "hyper-detailed", no "8k". Edit prompts are administrative documents. They are supposed to read like instructions to a retoucher who is fast, literal, and has never seen your brand.
Name what to keep, by name
"Keep everything else unchanged" is a useful backstop and a bad strategy. The problem is that "everything else" is not a visual concept — it does not correspond to anything the model can attend to. "Her face, hairline, pose and jacket" does. Named nouns anchor; abstractions drift.
You do not have to invent the list each time. It is mostly fixed per subject type:
- People — face and identity, expression, pose, hairline and hairstyle, skin texture, body shape, clothing and its folds, glasses and jewellery, and the direction of the light on the face.
- Products — silhouette, material and finish, logo and label text, seams and stitching, every colour you are not changing, and the contact shadow that grounds it.
- Interiors and property — architecture (windows, doors, mouldings, floor), perspective and camera height, the existing daylight direction, and any fixed fixtures.
- Any photo — framing and crop, aspect ratio, and the grain or noise character of the original, so the edited region does not read as cleaner than the rest of the frame.
Pick three to six invariants, not twenty. A preservation clause that lists every noun in the picture starts to read like a description again — and you are back to the original problem. Protect the things a viewer would notice if they changed.
Change one thing per pass
Compound edits fail in a specific and recognisable way: the model does about one and a half of your three requests. Ask it to swap the background, warm the colour grade, and remove the coffee cup, and you will get a new background, a hint of warmth, and a coffee cup that is still there. Worse, you now cannot tell which instruction it ignored versus which one it attempted and botched.
So chain instead. Send one edit, look at the result, then feed that file into the next pass. Three sequential prompts almost always beat one prompt with three clauses, and they cost you nothing extra in judgement because each result is either right or wrong on one axis.
The corollary matters just as much: if pass two undoes pass one, the fix is not to add "and keep the background from before" to the prompt. That is re-litigating with words something you should be doing with files. Go back to the pass-one output and re-run pass two against it. The same logic kills the re-roll habit — sending the identical prompt again and hoping for better luck is not iteration. If the output was wrong, the prompt was wrong. Change a word before you spend another render.
Write an edit prompt in five steps
- Say what you are starting from, in one clause. "In this photo of a woman at a desk…" gives the instruction an anchor and stops the model treating your text as a scene description.
- Name exactly one change. One verb, one target, one outcome. If the sentence needs a second "and", split it into a second pass.
- List the invariants by name. Three to six specific nouns that must survive — face, pose, clothing, logo, floor, shadow — rather than a blanket "keep the rest".
- Add the physical constraints. Light direction, perspective, scale and the edge where old meets new. This is what stops a pasted-in background from sitting at the wrong angle behind a correctly-preserved subject.
- Render, compare to the original, then change one thing. Put the two files side by side and ask what moved that should not have. Fix that one thing in the next prompt, on the edited file, not the original.
Six failure modes and what actually causes them
| What you see | What went wrong | The fix |
|---|---|---|
| The background swap changed the person too | Your prompt described a scene, so the model rebuilt the scene — subject included | Name the identity invariants: face, hair, pose, clothing, and the light direction on the subject |
| The colour change also changed the material | Colour words carry material baggage — asking for "red leather" quietly swaps suede for leather | Split them: change the colour, and explicitly hold the material, weave, finish and texture |
| Object removal invented new furniture | You said what to take out but never what belongs behind it, so the model guessed | Describe the fill: the continuing floor, the wall, or the surface already visible on both sides of the object |
| The retouch flattened skin into plastic | Words like "flawless" and "perfect" are unbounded — there is no point at which they are satisfied | Bound it: name the specific blemishes to remove, and list pores, texture and freckles as things to keep |
| The second edit undid the first | You prompted against the original file again instead of the edited one | Chain the passes — pass one output becomes pass two input |
| Nothing changed at all | The instruction was subjective ("make it pop", "clean it up") with no visual target | Convert it into something measurable: what specifically gets brighter, warmer, straighter, or removed |
Watch the edges, not the middle. Most edits that "look off" are technically correct in the changed region and wrong at the boundary — a hairline that has gone soft, a shadow falling one way while the new light comes from another, a horizon sitting at the wrong height behind the subject. Zoom to 100% along the seam before you accept a render.
When the agent writes the preservation clause for you
All of the above is what you do when you are talking to a general-purpose editor through a text box. The alternative is to use a studio that already knows what it is protecting. ReelWand’s Background Swap Lab is built around exactly one edit — cut the object out, drop it into a new scene — so the invariants live in the agent instead of in a sentence you have to remember to write: the silhouette, the material and finish, the colours you are not changing, and a cast shadow generated to match the new surface so the object sits on it rather than floating above it. It is aimed at product shots, not portraits. You upload the photo and describe the new scene, which is the only half of the prompt that is genuinely yours to decide.
The same split applies across the rest of the catalogue: Object Eraser holds the "reconstruct what was behind it" rule, Glow Retouch holds the "keep pores and texture" rule. Where an agent exists for your edit, use it — the preservation clause is the part people get wrong, and it is the part the agent has already written. Where one does not, write the two halves yourself.
Honest limits: instruction-following editors are much better at bounded, nameable changes than at open-ended taste. "Remove the chair" is a good instruction. "Make this feel more premium" is a brief for a human. And no prompt recovers detail that was never captured — if the file is soft or blown out, you are looking at a reshoot or an upscale, not a prompt problem.
Upload a product shot and describe only the new scene — the silhouette, material and grounding shadow are handled for you. Guests get one free watermarked render a day, no signup.
Try an edit with the invariants lockedFrequently asked questions
What is the difference between an AI editing prompt and a generation prompt?
A generation prompt describes the entire image you want, because nothing exists yet and every detail you leave out is a detail the model decides for you. An editing prompt describes a single change to an image that is already mostly correct, and explicitly names what must not change. The practical consequence is that detail works in opposite directions: describing the subject helps you in generation and hurts you in editing, because every noun you mention becomes something the model feels licensed to redraw.
Why does the AI change my subject when I only asked to change the background?
Almost always because the prompt described a scene instead of an instruction. If you write "professional woman in a modern office" over a photo of a real person, you have asked for a picture matching that description, not for a background replacement — so the model regenerates the person to fit. Fix it by writing the edit as a delta with a preservation clause: replace only the background with X, and keep the subject, pose, hair, clothing and the direction of the light on them unchanged.
Should I put multiple edits in one prompt?
No. Compound edit prompts fail partially rather than cleanly — the model typically completes one instruction, half-completes another, and silently ignores the third, and you cannot tell which is which from the output. Send one change per pass and feed each result into the next pass. Three sequential prompts cost three renders — the same as three re-rolls of one overloaded prompt, except each of the three actually banks a result you can keep.
Do I still need negative prompts when editing photos?
Much less than when generating. On instruction-following editors, a well-written preservation clause does the work a negative prompt used to do — saying "keep the skin texture and freckles" is more effective than listing "no plastic skin, no over-smoothing" as things to avoid. Negatives remain useful on Stable-Diffusion-style engines where they are a real sampling control; see our primer on negative prompts for where the line falls.
What should I do if repeating the prompt gives the same bad result?
Stop repeating it. Re-rolling an identical prompt is not iteration — if the output was wrong, the instruction was wrong, and the same instruction will keep producing the same class of error. Change one element: make a subjective word measurable, split a compound request into two passes, or add the invariant that keeps getting broken. If several rewrites all fail on the same edit, the change may be out of scope for prompting and better handled by a dedicated tool or a manual mask.
Put it into practice
Specialized image agents carry the craft this guide describes. Pick an available agent and start creating.
Part of ReelWand's AI Product Photography & Photo Editing tools.