How to preserve images during edits?

I’ve been using ChatGPT’s image generation/editing pretty heavily for a long-term project, and I’ve hit a problem that has become so bad that I actually requested and received a refund because the image tool was no longer doing what I needed it to do.

The frustrating part is that this isn’t just a problem with reference images.

There are really two related problems:

1. Reference-image fidelity

I can provide an exact screenshot of an environment and ask ChatGPT to recreate that environment at a higher visual quality while preserving the camera angle, geometry, proportions, furniture placement, etc.

Instead, it frequently decides to “interpret” the image.

For example, I can show it an exact screenshot of a specific area of a game and say, essentially:

“Recreate this exact environment, make it look better, morning lighting, and don’t change the layout.”

(Note: I will get the ai to give me a prompt that asks for everything I want, not just a small sentence, but I explain what I need and then it gives me the command or prompt that will work.. after I confirm its ok and says everything I wanted.)

And it will generate a completely different room that merely resembles the concept.

A cafeteria becomes a different cafeteria. An atrium becomes a giant greenhouse. The camera moves. Furniture moves. Architecture changes. New railings, signs, balconies, windows, etc. appear.

At that point it isn’t really recreating the reference anymore.

2. Editing an existing image without destroying what is already correct

This is even more frustrating.

I can have an image that is basically finished and ask for something extremely simple:

“Add one person here.”

Or:

“Move this person slightly to the right.”

Or:

“Add this specific element while leaving everything else unchanged.”

Instead, the model may alter the background, move existing objects, change lighting, change the character’s position, alter proportions, redesign the environment, or otherwise modify things that I explicitly did not ask it to touch.

And this is the part that really bothers me because the whole point of an editing tool is that I should be able to preserve the things that are already correct.

I’ve tried being extremely explicit.

I’ve tried:

  • “Do not change anything else.”

  • “Preserve the existing image exactly.”

  • “Only add this one element.”

  • Providing the exact reference image.

  • Providing annotated images showing exactly where something should go.

  • Breaking the process into smaller steps.

  • Building prompts incrementally and only locking in things that worked.

  • Removing unnecessary references so there would be less ambiguity.

It still drifts.

What makes this especially frustrating is that I’m not asking for subjective artistic interpretation. I’m working with environments where spatial accuracy matters.

If a chair moves, that’s wrong.

If the camera moves, that’s wrong.

If the atrium changes shape, that’s wrong.

If I add one character and the entire background changes, that’s wrong.

I don’t need the AI to be more creative. I need it to be more faithful.

The workflow I actually want is pretty simple:

Existing image → preserve it → make one requested change → preserve everything else.

Or:

Reference image → recreate the same scene → improve visual quality → preserve the geometry and composition.

Instead, it often feels like:

Existing image → “I see what you were going for” → generates an entirely new interpretation.

And I’ve had generations where I spent a significant amount of time carefully getting an image almost perfect, only to lose that progress because the next edit changed unrelated parts of the image.

That is incredibly frustrating when image generations are limited and you are paying for access to the tool.

Eventually I reached the point where I felt like I was paying for a feature that I couldn’t reliably use for the workflow I needed, so I requested a refund — and I did receive one, atleast I think (We will know in 10 days). I’ll probably cave and subscribe again since I dunno how to even finish my projects now.

I’m posting this because I genuinely want to know:

Has anyone else noticed a decline in image preservation/editing fidelity?

Is this a known limitation of the current image model? A regression? Something about how reference images are being passed to the model? Or is there a particular workflow that people have found that actually allows you to make precise edits without the model redesigning everything around the edit?

I’m not expecting pixel-perfect editing from a generative model.

But if I say “change this one thing and leave everything else alone,” I feel like the tool should at least have a reasonable chance of doing exactly that.

Right now, that feels increasingly unreliable.

Hi and welcome to the community!

While I am not sure how to resolve the issue with image editing, maybe our @Image_Enthusiasts can share some insights.

It would be great if you could share a few examples so we have something to work with.

Hi, welcome to the forum!

Here’s follow-up. What kind of edits are you doing? Conversational or using the edit tools through the user interface? Edits, do you do them through https://chatgpt.com/images or through https://chatgpt.com/library, or directly in the same chat?

Depending on which kind of approach you take there are different things you can do.

Currently I would not encourage edits through the images or the library, this because the hand-off to a new chat is wonky. I’d stick to edits in an on-going chat.

The approach with the UI Tools have to make the model aware that there is a selection/an interaction directly with the image, otherwise there’s a chance the model will reinterpret the user’s input as a request for a new image.

Personally I don’t do edits from the images or the library. The creation of the chat, at least for me, has an empty reference.

The conversational approach is pretty much straightforward, you tell GTP the changes you want to make in the image by telling what you want to do. This is where I’ve seen more success with the edits compared to the other methods.

I noticed you do lots of constraints and negations in the statements you do. Something I’ve noticed is that using factual positive or neutral statements work much better for images than resorting to negatives and constraints. Use the constraints and negations only if really needed.

For the images with ChatGPT assign an ID to the images within the chat, that way the model is aware of which image you are referring to.

One idea I like a lot is the small iterative steps. Helps a lot when it comes to edits. Still sometimes the text model falls in love with one element of the image and won’t let go.

You mentioned that you work with references of images. In this case, instead of uploading it to the instanced chat, save it in the library or in a specific project’s knowledge. That way you create a second chat to do parallel work with the edits. Like having a chat talking about the other chat. You can use in both instances the @<file> to reference the image you are using as reference.

Still there is something that is currently happening… The text model is working fine but the image model drifts after the hand-off. In this case you can branch out the chat instance but clicking the … on the previous interaction and choose

Another idea that comes to mind is to start doing images with ChatGPT Work instead of the Classic version. Might be a perception bias here, but workflows go smoother this way.

Finally, always start a chat with your intention, telling the model what you are going to do with images. It also helps if you assign GPT a role/persona like Art Director, Eccentric Studio Photography Expert… and much more. Use an adequate persona for the purpose of the chat.

Hope this helps.

Dys

Ah yes, we got two big threads with images, if you want, you can share your images with us:

and

Thanks — this is actually really helpful, because I can explain exactly what I’ve been doing rather than just saying “image editing doesn’t work for me.”

I’ve been working on a fairly complicated long-term image project in ChatGPT throughout August. I started on August 1, and as of August 25 I still have not been able to get one of the final images completely finished, despite getting individual components extremely close.

My workflow has primarily been inside ongoing project chats, rather than taking an image from the Images page and starting a completely unrelated conversation with it.

The project is organized into separate chats/masters because I am building the final image in components.

For example, I have an Environment Master where I developed the cafeteria/environment, a Face Master where I am developing Sephiroth’s face, and a separate workflow for his body/uniform. The idea is eventually to combine the finished components into the final image.

The reason I separated them is that the final image is complicated enough that trying to solve the face, costume, environment, lighting, pose, etc. simultaneously creates even more opportunities for drift.

Environment Master

The cafeteria was developed in its own project/chat.

I spent a substantial amount of time getting the environment itself correct — architecture, layout, furniture, lighting, colors, composition, and the overall look of the cafeteria.

Eventually I had a version that I considered the correct environment.

The problem came when I tried to add a relatively small missing element, such as a cashier at the register.

I used the conversational editing workflow and also tried the UI editing/selection approach.

For example, I would provide the existing cafeteria image and ask for a cashier to be added at the register. When the cashier was generated incorrectly — for example, facing the wrong direction or not standing correctly — I tried to make a subsequent localized correction such as having the cashier stand straight and face the correct direction.

Instead of reliably modifying only the cashier, the system sometimes produced what was effectively a new interpretation of the entire cafeteria.

The architecture, furniture arrangement, proportions, lighting, or other previously solved details could change.

This created a loop:

Finished environment → add cashier → cashier wrong → correct cashier → environment changes → rebuild environment → add cashier again.

That is one of the major reasons I have been unable to finish the project.

Face Master

I eventually realized that the exact same thing was happening with Sephiroth’s face.

I had a face that was extremely close to what I wanted. I then tried to make small changes — particularly to the eyes, pupils, eyelids, and expression.

For example:

Face is correct → fix eye → face changes.

Then:

Rebuild face → fix expression → eye geometry changes.

Then:

Fix eye → identity/expression drifts again.

At one point I was trying to get the expression from a reference image I call the “Such a Puppy” reference. I also have separate eye references, including files I call Eye_C and Eye_O, because I was trying to establish the correct iris/pupil appearance.

The files are kept with the project/reference material rather than being treated as disposable generations.

Some of the reference files currently associated with the project include:

  • Sephiroth_Eye_C_HD.png

  • Sephiroth_Eye_O_HD.png

  • Sephiroth_Super_Master_02_Costume_Materials_v2.png

  • Face Model.zip

  • several dated image-generation/reference files such as image_2026-08-24_015929029.png and image_2026-08-24_020305031.png

  • another reference image stored as fc1cbe94-0f92-44b1-8d68-ef1d02357bb5.png

The project also contains the generated master/reference images from the various stages of the work.

I have tried the “edit instead of regenerate” approach

This is important because I don’t think my problem is simply that I am asking ChatGPT to generate a new image every time.

I have used the existing image and attempted to edit it through the interface.

I have also used conversational instructions in the same ongoing chat.

I have tried selecting the area that needs changing and describing the requested correction.

And I have tried extremely explicit instructions such as:

  • change only the eye

  • don’t change the face

  • preserve the environment

  • don’t alter the cafeteria

  • only add the cashier

  • keep everything else identical

Those instructions sometimes appear to work, but they are not reliably preserving the finished portions.

Eventually the system can produce a fresh generation that is merely “inspired by” the previous image rather than actually behaving like a localized edit.

I also tried changing the prompting strategy

I originally used a lot of constraints because I was trying to prevent drift.

For example, I would explicitly say things like:

“Do not change the face.”

“Do not change the background.”

“Only change the eye.”

“Keep the cafeteria exactly the same.”

I have since experimented with the opposite approach: describing the successful features positively instead of relying heavily on negative instructions.

I developed a version-control style workflow for the prompts.

The idea was:

Successful prompt → lock the successful wording → add only the next correction → generate → evaluate → keep successful wording → modify only what failed.

So rather than treating the approved image as the thing that is locked, I treated the successful prompt + reference set as the locked production recipe.

That helped conceptually, but it did not solve the fundamental problem.

The image generator can still reinterpret the whole image during a fresh generation.

The reference-image problem

I’ve also experimented with assigning different references different jobs.

For example:

Identity reference: controls Sephiroth’s facial identity.

Expression reference: controls the expression.

Eye reference: controls iris/pupil appearance.

Hair reference: controls the hairline and hair construction.

Environment reference: controls the cafeteria.

The theory was that this would let us build the image systematically.

However, when several references are involved, the generator can also start blending them rather than treating them as strict technical authorities.

That was particularly obvious with Sephiroth.

We could get a beautiful expression from one reference, but then the identity would drift.

Or the identity would be correct but the eyes would change.

Or the eye would improve while the expression changed.

The important distinction I have discovered

The problem isn’t that ChatGPT cannot generate any of these things.

It can.

In fact, that is what makes this so frustrating.

I’ve had individual generations where:

  • the Sephiroth identity was excellent

  • the expression was excellent

  • the eyes were close

  • the cafeteria was excellent

  • the costume was excellent

The problem is combining those successful pieces and then making a small subsequent correction without losing the previous success.

The generator seems capable of creating the desired result, but I have not been able to make it reliably preserve a finished component while generating a correction to another component.

So my workflow repeatedly becomes:

Get 95% correct → attempt the final 5% → lose 30% of the previous work → rebuild → repeat.

That is why this has been so exhausting.

Why I am asking about this now

I’ve been working on this essentially the entire month of August.

I started August 1.

It is now August 25.

I’ve spent the month trying to finish these images and have not produced one completely finished final image.

I actually ended up getting a refund for my ChatGPT subscription because I was consuming the available image generation trying to get these projects over the finish line without actually getting a finished result.

And that is what I am trying to understand.

I don’t necessarily need advice on how to write a better prompt.

I need to understand whether there is a specific ChatGPT image-editing workflow that I am missing which gives the image model a much stronger preservation signal.

In other words:

Is there a way to tell ChatGPT, technically rather than merely linguistically, that an existing image is the source image and that only a particular selected region should be regenerated?

And if there is, is there a particular way I need to initiate the edit so that the image model receives the image and selection as an actual edit operation rather than the text model interpreting my request as a new image-generation request?

Because I suspect this is where my workflow is breaking.

The text/conversation side of ChatGPT understands exactly what I am asking.

The problem seems to occur at the point where the request is handed off to the image-generation system.

The text model can understand:

“The cafeteria is finished. The cashier is the only unfinished part.”

But the resulting image can behave as though the instruction was:

“Generate a cafeteria based on this previous cafeteria and include a cashier.”

Those are obviously very different operations.

The same thing happens with the face:

“This face is finished. Correct the pupil.”

can turn into:

“Generate a Sephiroth face inspired by this face, with a different pupil.”

That is the behavior I am trying to figure out how to prevent.

So if anyone knows of a specific workflow involving the current ChatGPT image editor, image selection/masking, image IDs, projects, branching, or another way of forcing the image to be treated as an edit of the supplied source rather than a new generation inspired by it, that is exactly what I am looking for.

I am already keeping the reference images and successful generations organized inside the project, and I am working in ongoing chats rather than simply pulling random images from the library and starting over.

At this point I am trying to determine whether there is a missing technical step in my workflow — or whether what I’m experiencing is simply a limitation of the current image model.

Also:
Even after branching the chat specifically to prevent the drift, the same failure happened at roughly 95% completion. We successfully mapped the actual Eye_C texture onto Panel 7 and created a side-by-side comparison. But when I asked for the final depth/shading pass, the generation stopped treating Eye_C as the literal texture being edited and instead interpreted it as inspiration, generating a completely new eye. So the frustrating part is that the workflow worked until the final refinement stage, where it reverted to the exact behavior the branch was supposed to prevent.

So I think I get it. Before, you could have a chat, and you could use that chat to generate images and go to the next step and so on and so forth. But now, because of DALL-E going away, what they want you to do is actually go to the image page and generate an image there. So you select an image that was already generated, you click it, you make comments, you add attachments for what you want changed, and then the image will generate from there. Now, saying that, I am not sure if that means that it’s a guaranteed fix to preserving an image and just changing one thing. But I might try this later on at some point and see. Does that sound correct?

I am having this issue that started today for me as well… I’ve used it to edit images all the time and this has worked fine retains a lot of the original and only edits what I ask - but today it generates images completely unrelated to the image I’m telling it to edit… entirely new images. What’s going on?

You noticed that too? I just made a thread about it and I noticed the change within the past hour or so, even though it was working normally for me. Basically, I upload a picture and tell chatgpt: ‘Use this picture, change the pose and camera to x…’, and it was doing that and had always done that. Suddenly I noticed it was generating brand new characters or not using the image I just uploaded. At the same time I noticed that picture files in chat are now shown differently with a file name next to them.

yeah I don’t think it works in chat anymore.. you have to actually go to the image page.. choose the image.. or choose the image in your library and then work from there.

Yes! I also noticed that, too. Chatgpt recommended that to me, itself. I noticed that after going to the library, uploading the image there and then asking for updates made the image upload like before in the chat: So, one big image instead of showing me a tiny image with the file name next to it.

Edit: Okay…wait, I just uploaded a picture without going through that library process and the picture uploaded like before. I need to test if the model anchors to the picture.
Edit: I just tested it in regular chat, just uploading the picture in there and asking for an edit…now it seems to be working like before, again. The model referenced the uploaded picture and generated based on my instructions instead of making a brand new character. I do see new options for editing the picture when I click on it directly.

Nice! I was just guessing but I think that might be the new way to do it.. which sucks cause it means a bunch of new chats per edit I am assuming?

Yeah, if they make us go through the library for it, then that’s more work.
Could you do me a favor and see what happens when you try upload a picture file in a regular chat now, please? See if it works like before.

I will not be able to.. I ended my subscription so cannot use those chats anymore because they have attachments. But I will test when I can!

Oh, sorry about that. No worries then!

Yeah I regret doing it now that I figured this out. But I hope its helpful!

I’m experiencing exactly the same regression.

Before, ChatGPT Images was extremely faithful when editing existing images: it would return essentially the exact same image, with only the requested micro-correction. The result was almost pixel-for-pixel aligned with the original.

Since around June 2026 (I’m not certain of the exact date), this has changed dramatically. It now zooms, changes framing, moves the subject, alters the resolution, and modifies unrelated elements.

This makes precise image editing practically unusable compared with how it worked before.

(This is absolutely not resolved. Why is this topic marked as “Solved” when the underlying issue is still clearly happening?)