How to prevent the API returning multiple outputs?

[quote=“david.captur, post:41, topic:1251365, full:true”]
The proposed work-around by OpenAI does not work, I’ve attempted various permutations of it to force the model to return a single response but it keeps happening at the same rate.

We’re running into the same issue using the Responses API with gpt4.1 and using file search tool. Leading to a confusing, sub-par user experience, so hopefully this can get resolved soon! :crossed_fingers:

Hello @OpenAI_Support , do you intend to address this issue or not? Are there any updates regarding this matter? I find it absurd that you continue to introduce new features aimed at ChatGPT users, who do not pay for the service, while completely disregarding the persistent issues affecting the API, which impact paying businesses.

This happens to me as well but it seems only on the responses API. When I go back to completions it appears to work flawlessly for gpt 4.1 mini, which is my primary agent model.

Still an issue for us using Responses API (and Agents SDK for TypeScript as well). Would be great if @OpenAI_Support would address.

+1 here. I get multiple versions of the output. Varies between 1-15 versions of it.

Building a voice agent with gpt-4.1-mini + streaming + structured outputs seems to have the right amount of brains and speed, but I cannot use responses API, just trash experience with it. Other models aren’t quite as reliable with tool calls or aren’t fast enough for naturally flowing voice conversations. I really wanted to use responses API, but looks like it only works with chat completions reliably. For now that is the direction I am going with, but I am a little concerned that support for this won’t be long lived since OpenAI is really pushing responses api. @OpenAI_Support - would live to hear that fixes are on the way or that gpt-4.1-mini on chat completions will have support for a long while.

I am experiencing this. Looks like it has not been fixed. I am using 4.1-mini

I believe we have a fix for this now, which we’re rolling out to a select group of customers to measure impact before we roll it out more broadly. If you’d like early access to the fix and are able to help us validate the fix, please send me a message and I’ll add you to the test group.

Thanks again for all your patience while we investigate this issue!

Hi! Do you have an estimation about when a solution could be broadly available? I think I’m hitting this very problem, when I switched to JSON schema specifications. Strangely enough, in the multiple-output responses, I get a very high (~30000) number of input tokens used. Thanks!

Please, we need a solution ASAP.

Will you return the money?
This bug has caused a significant increase in our consumption, and our clients are making claims for refunds.

hi @Carlos_Lombardi hoping to fully roll out by end of week. Sorry to hear about the high input tokens, this fix won’t impact that, but if you are seeing duplicate (or very similar) messages repeated, it hopefully addresses that.

Hi @ferbotmaker, sorry to hear about that, working to get the fix tested and fully rolled out this week.

I think you misunderstand.
It would be from all the developers across months of this issue being billed for OUTPUTs, where generation continues non-stop in repeating loops until hitting the maximum the AI model is allowed to produce, and getting parse errors instead of consumables out of OpenAI’s SDKs used as they are documented.

  • No special token emitted or caught stops the loop.
  • No end of JSON stops the loop.
  • No "stop" sequence the developer can offer on the Responses API can stop the loop because OpenAI removed that parameter.
  • No logit bias on the Responses API can tune up the start of another loop, nor penalty against repeating forever, because OpenAI removed those parameters and also blocks special token numbers from being employed.
  • Multiple symptoms which also include loops of trained “pretty” JSON whitespace tabs and linefeeds taking over when the CFG grammar is released upon entering a string and the AI generates maximum output.

Hi @jasondouglas thanks for the prompt answer. I admit I’m kind of worried, or confused, about what you wrote. I tried several times the same conversation pattern involving the same prompts using Responses API, one with JSON schema and other without. The conversation with JSON schema generates multiple outputs with slight variations in the text, and ~30K input tokens. The conversation without JSON schema does generate just one output, and with ~500 input tokens. Hence I expect that when the problem gets solve, I will notice a decrease in the quantity of input tokens related with the response. I hope this expectations to be realistic. Thanks for reading!

Yes, correct, assuming that the issue you’re seeing with similar multiple outputs is covered by this fix, and not a different model behavior issue.

Jason - I’d like to help test. We are seeing similar problems on a fine tuned 4.1-mini model (happens on non min too) where the agent sends multiple responses after being instructed to only send one.

Hi @HA_Adam, sent you a message!

I’m also facing this issue (using structured output and tools, with gpt-4.1). Is there any estimate for when this problem will be fixed?

I am also facing this issue with the Response API (using tools with gpt-4.1). I keep getting numerous Response Outputs with a mixture of similar responses and a function call that leads to separate pathway of the logic flow. It would be really helpful to know the estimate for when this problem will be fixed as well.

I’m also facing this issue when using GPT-5, the responses API, and streaming. After the first completion finishes, I immediately get another with a different message id (item id). My code errors out at this point because it is not expecting a second message. Therefore, I don’t know if just a second message is sent or several more. It seems to happen about one out of twenty requests.