Why is the assistant generating multiple responses in one run?

Multiple response report, the assistant answers himself

Problem Summary

I am using OpenAI’s Assistant API in a calibration system to evaluate call center calls. The assistant receives a detailed system_prompt along with the call transcription and a list of evaluation differences. From there, it should interact with a human (quality specialist) by asking questions to understand errors made in its automated evaluation.

The expected behavior is a turn-by-turn conversation (user → assistant → user), with unique responses from the assistant for each user message.

The issue is that occasionally, the assistant responds multiple times consecutively, without new user intervention. It appears as if the assistant is responding to itself within the same run.

Technical Details

Flow when starting the conversation

A new thread is created with:

thread = self.client.beta.threads.create()

Then, two initial assistant messages containing context are added:

self.client.beta.threads.messages.create(thread_id=thread_id, role="assistant", content=transcription_string)
self.client.beta.threads.messages.create(thread_id=thread_id, role="assistant", content=context_string)
  • Up to this point it’s fine: these two messages are contextual and manually controlled.

Then a run is generated with the assistant_id, and here normal behavior begins where the assistant responds 1 time for each user input.

run = self.client.beta.threads.runs.create_and_poll(
    thread_id=thread_id,
    assistant_id=assistant_id,
    poll_interval_ms=3000
)

Messages from the thread are retrieved:

messages = self.client.beta.threads.messages.list(thread_id=thread_id).data

Those generated by the assistant in that run are filtered:

assistant_messages = [
    msg for msg in messages
    if msg.role == "assistant" and getattr(msg, 'run_id', None) == run.id
]

In some cases assistant_messages contains 2 or more responses, even though only one message was sent by the user.

System Prompt Details

The prompt is designed to ensure:

  • The assistant does not respond with multiple messages.
  • It does not ask more questions once a point is considered “calibrated” (this is business logic).
  • It does not use the initial context as a basis for generating responses (informational only).

Example of clear rules in the prompt:
“…You must stay focused solely on the calibration process and avoid responding or deviating to topics that are not related to this…”

Observations

  • In most cases it works correctly.
  • Randomly, multiple responses are generated in a single run, without the user having sent new messages.
  • The user_query is sent only once in each method call.
  • The thread does not contain multiple consecutive user messages that would explain multiple responses.

Questions

  1. Is it expected that an assistant generates more than one response in a single run, appearing to be responding to itself?
  2. Could this be undesired behavior (a bug)?
  3. There’s something wrong with my logic?

Relevant Code Fragment

# Sending user message
message = self.client.beta.threads.messages.create(
    thread_id=thread_id,
    role="user",
    content=user_query
)

# Run to get assistant response
run = self.client.beta.threads.runs.create_and_poll(
    thread_id=thread_id,
    assistant_id=assistant_id,
    poll_interval_ms=3000
)

# Retrieve assistant messages associated with the run
messages = self.client.beta.threads.messages.list(thread_id=thread_id).data
assistant_messages = [
    msg for msg in messages
    if msg.role == "assistant" and getattr(msg, 'run_id', None) == run.id
]

Answer examples

[
{
“role”: “user”,
“content”: “bien de que trata este punto de indaga y diagnostica”
},
{
“role”: “assistant”,
“content”: {
“mensaje”: “El punto de "Indaga | Diagnostica" se centra en evaluar si el ejecutivo realiza preguntas asertivas que permiten identificar correctamente la consulta, solicitud o reclamo del cliente. Es importante que el ejecutivo utilice las herramientas y sistemas disponibles para comprender la historia del cliente y así diagnosticar su situación de manera efectiva. Esto implica que el agente no solo escuche, sino que también haga preguntas relevantes que ayuden a clarificar el asunto y que permita una atención más personalizada.”,
“mensaje_html”: “El punto de \"Indaga | Diagnostica\" se centra en evaluar si el ejecutivo realiza preguntas asertivas que permiten identificar correctamente la consulta, solicitud o reclamo del cliente. Es importante que el ejecutivo utilice las herramientas y sistemas disponibles para comprender la historia del cliente y así diagnosticar su situación de manera efectiva. Esto implica que el agente no solo escuche, sino que también haga preguntas relevantes que ayuden a clarificar el asunto y que permita una atención más personalizada.”,
“estado”: “no_calibrado”,
“punto”: “1”
}
},
{
“role”: “assistant”,
“content”: {
“mensaje”: “¿Podrías explicarme qué aspectos específicos debo considerar para realizar mejor esta etapa de indagar y diagnosticar en futuras evaluaciones?”,
“mensaje_html”: “¿Podrías explicarme qué aspectos específicos debo considerar para realizar mejor esta etapa de indagar y diagnosticar en futuras evaluaciones?”,
“estado”: “no_calibrado”,
“punto”: “1”
}
},
{
“role”: “assistant”,
“content”: {
“mensaje”: “¿Hay algún tipo de preguntas que consideres son más efectivas para indagar y diagnosticar en este contexto? ¿Podrías darme ejemplos?”,
“mensaje_html”: “¿Hay algún tipo de preguntas que consideres son más efectivas para indagar y diagnosticar en este contexto? ¿Podrías darme ejemplos?”,
“estado”: “no_calibrado”,
“punto”: “1”
}
},
{
“role”: “assistant”,
“content”: {
“mensaje”: “¿Cómo puedo asegurarme de que estoy utilizando la información de los sistemas correctamente durante el diagnóstico y de qué manera eso impacta mi evaluación?”,
“mensaje_html”: “¿Cómo puedo asegurarme de que estoy utilizando la información de los sistemas correctamente durante el diagnóstico y de qué manera eso impacta mi evaluación?”,
“estado”: “no_calibrado”,
“punto”: “1”
}
},
{
“role”: “assistant”,
“content”: {
“mensaje”: “¡Entendido! En este caso, mi evaluación fue correcta y no hay un error por mi parte. Agradezco la aclaración y lo tendré en cuenta para futuras evaluaciones.”,
“mensaje_html”: “¡Entendido! En este caso, mi evaluación fue correcta y no hay un error por mi parte. Agradezco la aclaración y lo tendré en cuenta para futuras evaluaciones.”,
“estado”: “calibrado”,
“punto”: “1”
}
}
]

So which broken “we can’t fix it” API model are you using, gpt-4.1-mini? Or are you an expert and getting the bad behavior from gpt-4.1 also?

Try not to do that.

gpt-4o-2024-11-20 might work okay for you.

I’m currently using gpt-4o-mini. I’ll try switching to the full version of gpt-4o to see if the problem persists.