Collection of GPT-image-generator 2.0 issues, bugs, and work-around tips (check first post)

Do you currently use ChatGPT-5.5 thinking? I’m pretty happy with working with thinking when I prompt for images​:relieved_face:

(Small note: speaking of languages, English isn’t my native language (no surprise there), so if my typos/grammar is unbearable/unreadable, please tell me. I usually try to read it through what I’ve written and correct myself if I notice my mistakes)

What will exhibit “the pattern” more highly is image input, which can be done on the edits endpoint. Lets send your two images for a new amalgam to gpt-image-2, at medium quality (and mandatory “high” input fidelity):

The mood and subject of a new image reflects these past images. However, there is no sunset - the image is just well lit by an overcast and gloomy sky. Sparse pine trees are seen. The castle keep is more a ruins, with tattered flags and chinks in its brickwork.

The output:

Now lets keep on iterating.

This aged ruins is transformed back in time into the same castle keep in its heyday, with noble knights on the path and the castle reconstructed to a noble fortress in its original form, without signs of death or gloomyness.

What this model cannot help but add is sunsets…

This noble fortress is transformed forward in time to a gloomy shade of its former self oozing death and defeat and destruction. The castle keep is more a ruins, with tattered flags and chinks in its brickwork, and the road is wet and rutty with evidence of battles past.

All requests contained the same second image, while iterating on the first.

After just a few passes, we see that the symptom is there: mottled clouds, highly-textured mud, and overlay of a tight blotchy pattern to everything drawn. The API is not immune, even when you control everything sent in input size, quality, output resolution. You trying on virtual clothes will be just as damaged.

Is this intentional, so that output images become unworkable before you can see after several rounds, the model has transformed your identity and ethnicity from an input picture into something unrecognizable, as with prior models making social media rounds?

Prompting 'a quirky mousy girl', with details

Absolutely agree here. I’ve noticed the same symptom even when editing through the API.

I still feel like API outputs can sometimes be a little less noisy than ChatGPT outputs, but the pattern is definitely still there. It may just be less obvious depending on the image/prompt.

I also wonder if fantasy images are especially affected by this, since they often include fog, mud, clouds, battle damage, fabric, vegetation, armor, ruins and other high-detail textures where that blotchy/noisy pattern can hide or amplify​:thinking: Because I can’t really see that in cinematic or photorealistic images.

Did you check the video that was linked earlier in the thread? I took a screenshot of this part because it reminded me of the same issue.

Maybe this is partly “default fantasy texture behavior” or maybe fantasy just makes the symptom easier to see.

FYI

Haha nice one with quirky mousy girl…thank god she at least was visible unlike the cow​:winking_face_with_tongue:

I use thinking standard. But it is actually not good for testing. If you really want to see what the generator does with prompts and triggers, you should be able to talk with the generator directly, without any changes. Otherwise you will never really know what a good and bad prompt is, and what triggers are doing.

Before GPTs prompt style was horrible! I wrote all prompts my self. But now its is way better, but not perfect.

If you just want generate good images, you can use the help of GPT. it still puts bad words like “scene” in the prompt, but the prompting style is way better now.

I even think that the system uses again in the background a improvement system, just as a part of the conversation context and the multi modal structure of the generator now. This will make testing way way more difficult for me.

(i use sometimes my dyslexia phrase: “yu ar alowed to kip al erors you find” :nerd_face:)

For Developers Only!

For Developers Only!

All examples shown here come from the DallE 3 system, not from 2.0.
Read to the end to get the point.

This is directed at developers, not users. (You probably know far more than I do, but I am writing this anyway because 2.0 is broken for fantasy images, and the developers obviously did not notice it before release.)

I am showing a few images here that demonstrate what toxic input data can do to weights. AI systems are not truly intelligent, they are pattern-analytical and transformative systems. That also means: garbage in, garbage out.

At first I assumed this was stuff that OpenAI itself inserted into the images. But it could also be a reaction caused by poisoned training data from other images and other companies inserting stuff into images.

Here are a few examples of toxic effects that in certain situations, led to damaged weights and results.

Errors Around Hair

Artifacts around hair. It is difficult to say exactly what triggered this. Possibly compression artifacts from poor cameras or badly compressed images.

Blue Shift in Shadows

In this image, the same issue appears, plus an additional blue shift in black shadows, typical of yellowed or poor-quality photographs. I have seen many images that had blue and red shifts in the shadows.

Over-Sharpening Artifacts

Here is the problem of excessive sharpening. It repeatedly produced a background motif: a horribly overemphasized galaxy and stars so strong that they only looked like blind noise.

Photo Camera Quality (image generator 1.0)

compression noise and grain

Over-Sharpening

“Flicker Confetti Noise”

Here is a similar image, something I called “flicker confetti noise.” It looks like dirty or damaged photographs. The system may have inserted it because it matched a learned pattern, namely small dust particles in sunlight.


Me, when i get text or patterns in my images. (Confetti)

Sharpening Damage

Here is an image with oversharpening. These are typical defects from poor sharpening algorithms that unfortunately are still widespread everywhere. They create edges that are too black and too white on the edges. The system learned this defect.

Chromatic Aberration Shift

Here is an image showing chromatic aberration shift, a typical effect of poor camera lenses. It is a prism effect that shifts colors away from the image center.

“Bird Shit Moon”

On of my “favorites” in DallE.
This is what I called the “bird shit moon” effect. There must also have been images in the training data that trained this horrible moon into the system. I observed it suddenly appearing after a system update, and it never disappeared again. (And the horrible galaxy and “starts” noise.)

And not in all systems! The Anakin people get the better quality :pensive_face:

Pattern Dissolution / Possible Poisoning

And here is the most important point.

I speculated that this could originate from Nightshade and Glaze images. When triggered, there was about a 10% chance of receiving such an image. Except during a very severe phase where almost every image was broken, these images later appeared only sporadically, but they are almost certainly still present in the DallE 3 weights.

The images dissolve into a pattern. If this is not Nightshade poisoning, then there must be some other kind of stuff that negatively affected the training data.

Why Testing Matters

When testing an image generator, it is important to test as many styles and motifs as possible, and therefore also the full spectrum of training data for errors.

Someone who never generates manga images will never see manga-related training damage (for example me). Someone who never creates fantasy images will not see the damage in that sector.

Since a prompt does not activate all weighting data equally, it is possible for broken and correct images to be produced simultaneously by the same engine. That now also makes me suspect that this may not be stuff inserted into the latent space, but instead may originate directly from the training data.

Speculation

Now the speculation.

If OpenAI is not inserting stuff into latent spaces (in that case, sorry for the accusation), then these defects could come from stuff inserted by other groups. The AI learned these patterns and is now reproducing them.

Both the images I observed in DallE and the new pattern could also originate from the training data if it was poisoned with stuff. So it does not necessarily have to be Nightshade, it could also be another learned pattern.

The reason why this appears more frequently in fantasy images, or in patterns derived from fantasy images within the training data, is that almost all of these images were generated by image generators. They are not photos with typical photographic weaknesses. They are artificially created patterns.

If this is true, then the training data needs a better filter. One that can either remove such patterns or completely filter the affected images out of the dataset.

I know what that means… New training… Possibly very expensive…

If it was not this guy…
the nightshade monster
…then it may be stuff inserted by companies.

That is why I wrote the critical text above. If you poison the infosphere with patterns in text, images, and sound, they can not only be detected and distinguished, they may also damage training data.

I have seen patterns that could match those used in 2.0. (No visual analysis, I am not a professional, but they are visible.) Companies should communicate what kind of stuff they are embedding into images, because that could help in developing filters.

I will wait for the first published papers to see whether this speculation turns out to be correct…

Yeah I totally get that.

For me, using ChatGPT -5.5 thinking can be useful because my goal is often to get a better image result and to see which words in the prompt carries more weight when the gen is generating images. So I can follow in realtime what the possible output might be and what gets ignored. If I’m not happy with the output, I can be clearer in my next prompt, that I don’t want anything changed or ask what needs to be changed in my prompt to get the result I’m striving for​:relieved_face:

But for testing the image generator itself, I agree that any extra interpretation layer makes it harder to know what is actually causing the result. Then it becomes more like testing the whole pipeline: user prompt → ChatGPT interpretation → image output, instead of only the generator.

So in that sense, the API is probably the more fitting place to test gpt-image-2 directly, because it gives a cleaner path to the model and makes it easier to control prompt, endpoint, quality, image input and repeated edits.

Now I know you don’t use API, but I actually do test both and I have to say, my outputs at least, doesn’t differ much using gpt-image-2 in API or using it through 5.5 thinking in ChatGPT. If anything…I think the outputs are almost scarily similar, compared to, for example gpt-image-1.5. With that model (in my own experience and opinion) outputs had clearer differences using the gen model in API and in ChatGPT.

I don’t know if anyone else can relate to that?

I’m trying to absorb your post, even a day later.

I guess I’m missing the part concerning how you’re iterating those images?

It’s confusing because I have a really strong read beforehand, when I’ll get the pattern.

I thought I already showed that when you took the main descriptor of style offered in the original prompt, which the system claims isn’t a known style - the pattern goes away.

Am I getting confused? It looks like you’re iterating from the image itself

That is sending the generated output as the new input image to the API edits endpoint. it is not as a conversational context where a chat model can trigger a tool. On the edits endpoint, the prior output is used as a prompt’s reference image for the AI to interpret as instructed. I then just show a few successive alternations between an initial refinement, a reversal of the gloomy theme to happy, and back. With a degrading input image, the symptoms pile up.

I’m just dropping in here (even when no one asked, sorry for that) just because I have thoughts, that apparently needs to get out​:woman_facepalming:

But the way I’m perceiving this, I think these might be slightly different layers of the same issue.

One part may be prompt/style wording that triggers the pattern. But another part seems to be repeated image editing itself. Even without memory/context in the ChatGPT sense (in API), if the previous output becomes the next input, the same artifacts can carry forward and get stronger over several passes.

So maybe the prompt can trigger it, while iteration makes it worse or locks it in.

Gotcha.

Sometimes it takes me a couple of runs to figure out your vector.

Thanks :clinking_beer_mugs:

Do you think even the API uses an prompt improvement layer?

Because this lets me think that prompt improvement is now fixed in the system, even in API, because jeffvpace says he used API for his picture and used only a short prompt.

GPT says the used prompt sent to the generator was this:

A surrealistic underwater alien city, brightly illuminated and crystal clear, with every figure and object fully visible. Show a vast luminous metropolis beneath the sea with fantastical alien architecture, glowing domes, elegant towers, transparent walkways, coral gardens, floating vehicles, and diverse alien inhabitants moving through plazas and terraces. Include large marine creatures drifting above the city, shimmering water, and radiant light beams filtering from the surface. Use a vivid, dreamlike, surreal visual style with rich detail, strong clarity, and balanced bright lighting so no important elements are lost in shadow.


I think if the weights are responsible for the pattern, they must amplify the pattern over time, they have no other choice. Every time image data is reused, even small parts, it amplifies the effect.

Only reducing details can stop it partially. Details triggers it. The weights find a structure and this triggers the pattern. At least this is my theory so far.

I had a similar thought. One possible way around this might be analogous to how early animated films were built with Cel animation.

Cel animation explained: definition, types, and methods

In traditional cel animation, separate visual elements were drawn or painted on transparent sheets and then layered together to form the final frame. The important idea here is not the historical medium itself, but the separation of elements into layers.

If the problem is that artifacts from one generated image are being fed into the next generation and amplified, then it may be better to isolate the parts of the image that need to change and update only those parts, rather than repeatedly regenerating the whole image.

In other words, instead of treating each pass as a new full-image generation, the image could be built or edited more like a layered composition: background, foreground objects, labels, callouts, shadows, and other elements handled separately, then composited at the end.

I am sure some systems and workflows already do something like this, but I have not seen it discussed much in this topic. The cel-animation analogy came to mind because it captures the idea of preserving stable parts of the image while only changing the layers that actually need revision.

gpt-image-1.5 will bill you for reasoning, able to produce language before the transition to image tokens. gpt-image-2? no discrete bill, just higher costs.

The model itself is likely a combination of training and prompt construction within instructions, with no way of receiving any other output than an image. Just like gpt-translate models are multimodal and are tuned and likely have a containerized version of the input as your “prompt”, but occasionally can go wrong and follow instructions instead of producing the destination language.

So, not “prompt improvement layer” before an actual model, but a transformer language model with understanding.

You can use this understanding of its ability to reflect on top-down image content as a generative sequence, to ensure adherence.

@EricGT Yes, that makes a lot of sense.

The cel-animation analogy is a really good way to explain it. Instead of throwing the whole image back into the blender every time, the more stable parts could stay untouched and only the parts that actually need revision would change.

That would probably make it harder for the same artifacts/noise to get carried forward and amplified over several passes.

So maybe part of the issue is not only that the model creates the pattern, but that the workflow keeps giving the pattern another chance to survive and multiply​:woman_shrugging:

@Daller I think _j answered that pretty well.

But I was thinking that, maybe the next interesting test would be to compare the same subject in two versions: one very low-detail/simple scene and one highly detailed/fantasy-textured scene, then run both through a few edit iterations and see where the pattern starts showing first​:thinking: (someone willing?)

🐙

Great image example :winking_face_with_tongue::finland:

For Developers only

Not in 2.0

I have been experimenting with the Cel technique, but I now think the AI may have hallucinated much of the process rather than actually following the intended method.

It seems I will need to work with the AI to establish a shared language that we both understand before moving forward. This is a technique I learned years ago with early generative LLMs. They did not know everything, but when I learned to describe problems in terms they clearly understood, it often opened successful paths to solving them.

In addition to what you said, here is another example using the API:

Original Image:

Image Edit:

Prompt

Change the squid mask to an elephant mask.

Notice the degradation of the image (use two-click zoom). The more complex the image, the more noticeable the degradation

Here is less complex example:

Original Image:

Image Edit:

Prompt

Modify the image: Small stick figures represent the players and must be correctly positioned on the field as per the official rules of baseball.

The degradation is still there, but not as noticeable.

Any subsequent edits will result in increased degradation.

Therefore any use case involving incrementsl edit workflow, either with ChatGPT or the API, is problematic.

I agree with you.

And this actually got me thinking about how I work with persona SKILL.md’s too.

For me, the point is not only giving the model instructions, but building enough structure and language that the tone becomes believable and consistent, instead of just imitating the surface of a style.

Image prompting feels similar in that way. It is not only about adding “better-looking” words, but about creating a structure and language the model can actually follow.

Sharing my Cel attempt conversation, the last image is lol.


Thinking I may have to talk about paint-by-numbers to setup the starting areas that a Cel can use, then move on from those.


Should have noted this other related topic, my mistake.

In my conversation about Cel animation, ChatGPT mentioned

Note: the generated “transparent” cel images were actually RGB images with a checkerboard background, not true alpha PNGs. I removed the checkerboard by approximation and then composited the layers. That is why some light edge artifacts remain around the fort, fence, and flags.

I saw what you meant about hallucination​:smirking_face:

I’ll actually need to test that Cel animation too.

And transparent background is by default chessboard, but there are some tips and tricks to get transparent background and pros in that.

I would love to see more images that you’ve prompted​:raising_hands: