Gpt-6-luna occasionally returns an additional message item for structured outputs

Environment

  • Surface: API
  • Endpoint or feature: Responses API (POST /v1/responses) with structured outputs
  • Model: gpt-6-luna
  • SDK and version, if applicable: openai-python 3.20.0 (also seen with 2.44.0)

Bug

What happened, and what did you expect?

gpt-6-luna occasionally returns an additional message output item for a structured output request. Only one item holds the schema-valid JSON. The additional item is empty, holds free-text working notes (“We need output JSON schema. classify stance. …”), or repeats the JSON with stray text such as </assistant:commentary> after the first value. Expected: one message item whose output_text is the JSON value.

Because of the additional item, response.output_text is not one JSON value and client.responses.parse(..., text_format=Model) raises Invalid JSON.

The response format name decides how often it happens. With the request below, remark_filing reproduces within a few calls; RemarkFiling once in 50; Answer, answer, and filing_answer never in 50 to 100 calls.

Reproduction

Minimal code, request, or steps to reproduce:

"""gpt-6-luna occasionally returns an additional message item for structured outputs.
pip install openai pydantic; set OPENAI_API_KEY."""

from typing import Any, Literal

from openai import OpenAI
from pydantic import BaseModel, Field

SYSTEM = """\
You file remarks a summarizer wrote about one client into the client's stored service notes.
Each remark states one specific plus the client's own stance on it.
For each remark answer three things: whether it carries the client's stance, which node it belongs under, and a summary of that node when the node is new or its tendency changed.

- C: a concrete thing plus the client's own evaluation, reason, comparison, preference, or avoidance.
- T: an evaluation is present but thin. The default when unsure.
- N: no evaluation of the client's own.

A node is a category or a subcategory under it. Reuse an existing subcategory's exact name when it fits; otherwise open a new one.
Write a summary for a new subcategory; for a listed one, null unless the remark changes its tendency.
Refer to the client only as "the client".

<catalog> lists the nodes: `##` is a category, `###` is a subcategory with its summary.
<remarks> holds the remarks, numbered; cover every remark once, in input order.
"""

USER = """\
<catalog>
## delivery
### speed — Settled criteria; rarely changes them.
### packaging — Settled criteria; rarely changes them.
### pickup — Settled criteria; rarely changes them.
## support
### response time — Settled criteria; rarely changes them.
## app
### notifications — Settled criteria; rarely changes them.
### search — Settled criteria; rarely changes them.
## product
### photos — Settled criteria; rarely changes them.
## deals
### coupons — Settled criteria; rarely changes them.
</catalog>

<remarks>
[1] Prefers delivery within two days; needs the items for the weekend.
[2] Finds support replies that take over a day frustrating.
[3] Prefers thick padding in packages; glass items keep breaking.
[4] Keeps only sale notifications turned on.
[5] Finds delivery easier than store pickup on weekends.
[6] Prefers handling exchanges in the store right away.
[7] Skips coupons with complicated conditions and pays full price.
[8] Trusts review photos more than product photos; they show the real color.
[9] Finds ordering in the app during the morning commute most convenient.
[10] Goes to the store near work in the morning because evenings are too crowded.
</remarks>
"""


class Entry(BaseModel):
    """One remark's filing entry, in the order the filing call generated them."""

    id: int
    stance: Literal["C", "T", "N"] = Field(
        description="Whether a remark carries the client's own stance: a clear one, a thin one, or none at all."
    )
    category: str
    subcategory: str
    summary: str | None


class Answer(BaseModel):
    """The filing call's whole answer: one entry per remark."""

    entries: list[Entry]


def strict_schema(model: type[BaseModel]) -> dict[str, Any]:
    """`model_json_schema()` with `additionalProperties: false` on every object, as strict mode requires."""

    def walk(node: Any) -> Any:
        if isinstance(node, dict):
            out = {key: walk(value) for key, value in node.items()}
            if out.get("type") == "object":
                out["additionalProperties"] = False
            return out
        if isinstance(node, list):
            return [walk(item) for item in node]
        return node

    return walk(model.model_json_schema())


client = OpenAI()
  text_format = {"type": "json_schema", "name": "remark_filing", "strict": True, "schema": strict_schema(Answer)}

for call in range(1, 51):
    response = client.responses.create(
        model="gpt-6-luna",
        input=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": USER}],
        text={"format": text_format},
        reasoning={"effort": "none"},
        temperature=0.2,
        max_output_tokens=4096,
        store=False,
    )
    messages = [item for item in response.output if item.type == "message"]
    if len(messages) == 1:
        print(f"call {call}: ok")
        continue
    print(f"call {call}: {len(messages)} message items, id={response.id}")
    for index, item in enumerate(messages):
        print(f"--- output_text[{index}] ---")
        print(item.content[0].text)
    break

This reproduced on call 5 (10 remarks) and on call 1 with 15 remarks. Changing only "name": "remark_filing" to "name": "Answer" did not reproduce in 100 calls.

Diagnostics

  • Exact error and HTTP status:
    HTTP 200, status: "completed", reasoning_tokens: 0.
    When the extra item is free text, client.responses.parse raises pydantic_core.ValidationError: Invalid JSON.
  • Request ID, if available:
    resp_0f93f0115237ff4b016abb8035d4e087d0a727d5b126237500
  • When it happened (with timezone) and how often:
    continuously since 2026-09-28 (KST, UTC+9), when we moved to gpt-6-luna, at about 3 to 5 in 100 calls on our production request.
    The request IDs above are from 2026-09-29, 16:12, 18:08, and 18:09 KST.
    Also seen with streaming and with temperature 1.0.

I like what it returned when I ran this as a RESTful call…A complete reasoning commentary channel dump to the output - at reasoning.effort:“none” confirmed in the response echo, completely bypassing the strict schema enforcement.

DevDay: no encryption? Rather, another poor prediction where the AI wanted to reason somewhere anyway.

358 tokens of message 1, 248 tokens of message 2 - a usage of 622 output tokens, none classified as reasoning and apparently 16 of channel overhead given to the gods.

We need output JSON schema. classify each. 1 C delivery speed existing tendency maybe settled criteria rarely changes, no change null. 2 T? frustration evaluation, concrete support response time, own evaluation yes thin? C definition evaluation reason etc. "frustrating" evaluation, likely T as thin evaluation. 3 C packaging. 4 N? keeps only sale notifications turned on is preference/behavior, own stance? clear preference implied, C. app notifications existing settled criteria, no change. 5 C delivery pickup? category delivery subcategory pickup likely preference delivery over pickup weekends. Existing pickup summary settled criteria, no change. 6 C delivery pickup? store exchanges, handling exchanges in store right away. Could category support? delivery pickup node is pickup logistics, but this is exchange preference in store. New category? likely delivery/pickup? Node taxonomy categories given, choose delivery pickup perhaps. Existing pickup exact fits store pickup? This is exchanges, not pickup. Could new category "returns" subcategory "exchanges" summary. 7 C deals coupons existing. 8 C product photos existing. 9 C app new subcategory ordering? convenience app ordering morning commute. New subcategory "ordering" summary. 10 C delivery pickup? store near work morning due crowds, likely pickup? This is store visit timing, new category? delivery pickup perhaps pickup node, but no pickup order stated. Could category delivery pickup as store pickup? Remark says goes to store, likely shopping behavior, not delivery. New category "store" subcategory "visit timing". Need categorize sensible. 6 exchanges in store right away perhaps support / exchanges new subcategory. 10 store visit timing new category. Ensure summaries tendency. 2 T. 4 C. 9 C.

API response

x-request-id: req_46447f9c6b5a446e89984d5aef12663a

{
  "id": "resp_07f3e2e44586534f016abb87d62bfc87d08116f69e785359d5",
  "object": "response",
  "created_at": 1790674902,
  "status": "completed",
  "access_programs": null,
  "background": false,
  "billing": {
    "payer": "openai"
  },
  "completed_at": 1790674909,
  "error": null,
  "frequency_penalty": 0,
  "incomplete_details": null,
  "instructions": null,
  "max_output_tokens": 4096,
  "max_tool_calls": null,
  "model": "gpt-6-luna",
  "moderation": {
    "input": {
      "type": "moderation_result",
      "categories": {
        "harassment": false,
        "harassment/threatening": false,
        "sexual": false,
        "hate": false,
        "hate/threatening": false,
        "illicit": false,
        "illicit/violent": false,
        "self-harm/intent": false,
        "self-harm/instructions": false,
        "self-harm": false,
        "sexual/minors": false,
        "violence": false,
        "violence/graphic": false
      },
      "category_applied_input_types": {
        "harassment": [
          "text"
        ],
        "harassment/threatening": [
          "text"
        ],
        "sexual": [
          "text"
        ],
        "hate": [
          "text"
        ],
        "hate/threatening": [
          "text"
        ],
        "illicit": [
          "text"
        ],
        "illicit/violent": [
          "text"
        ],
        "self-harm/intent": [
          "text"
        ],
        "self-harm/instructions": [
          "text"
        ],
        "self-harm": [
          "text"
        ],
        "sexual/minors": [
          "text"
        ],
        "violence": [
          "text"
        ],
        "violence/graphic": [
          "text"
        ]
      },
      "category_scores": {
        "harassment": 0.0000011125607488784412,
        "harassment/threatening": 1.4593762210768862e-7,
        "sexual": 0.0000025071593847983196,
        "hate": 5.422218970016178e-7,
        "hate/threatening": 4.450850519411503e-8,
        "illicit": 4.495181578513688e-7,
        "illicit/violent": 3.089494448007139e-7,
        "self-harm/intent": 4.356879983467376e-7,
        "self-harm/instructions": 4.8883027733610114e-8,
        "self-harm": 5.338155705389253e-7,
        "sexual/minors": 1.2098656835274183e-7,
        "violence": 9.223470110117277e-7,
        "violence/graphic": 5.255395707074638e-7
      },
      "flagged": false,
      "model": "omni-moderation-latest"
    },
    "output": {
      "type": "moderation_result",
      "categories": {
        "harassment": false,
        "harassment/threatening": false,
        "sexual": false,
        "hate": false,
        "hate/threatening": false,
        "illicit": false,
        "illicit/violent": false,
        "self-harm/intent": false,
        "self-harm/instructions": false,
        "self-harm": false,
        "sexual/minors": false,
        "violence": false,
        "violence/graphic": false
      },
      "category_applied_input_types": {
        "harassment": [
          "text"
        ],
        "harassment/threatening": [
          "text"
        ],
        "sexual": [
          "text"
        ],
        "hate": [
          "text"
        ],
        "hate/threatening": [
          "text"
        ],
        "illicit": [
          "text"
        ],
        "illicit/violent": [
          "text"
        ],
        "self-harm/intent": [
          "text"
        ],
        "self-harm/instructions": [
          "text"
        ],
        "self-harm": [
          "text"
        ],
        "sexual/minors": [
          "text"
        ],
        "violence": [
          "text"
        ],
        "violence/graphic": [
          "text"
        ]
      },
      "category_scores": {
        "harassment": 2.1907815825757866e-7,
        "harassment/threatening": 4.181187231316626e-8,
        "sexual": 3.3931448129766124e-7,
        "hate": 1.5534977545950122e-7,
        "hate/threatening": 4.450850519411503e-8,
        "illicit": 3.138146753575985e-7,
        "illicit/violent": 1.4144759985825914e-7,
        "self-harm/intent": 6.681516500435412e-8,
        "self-harm/instructions": 4.5921356162867124e-8,
        "self-harm": 1.0348531401872454e-7,
        "sexual/minors": 8.059446204294678e-8,
        "violence": 4.356879983467376e-7,
        "violence/graphic": 2.2603242979035749e-7
      },
      "flagged": false,
      "model": "omni-moderation-latest"
    }
  },
  "output": [
    {
      "id": "msg_07f3e2e44586534f016abb87d71cf087d0a6d197a9b60aca72",
      "type": "message",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "annotations": [],
          "logprobs": [],
          "text": "We need output JSON schema. classify each. 1 C delivery speed existing tendency maybe settled criteria rarely changes, no change null. 2 T? frustration evaluation, concrete support response time, own evaluation yes thin? C definition evaluation reason etc. \"frustrating\" evaluation, likely T as thin evaluation. 3 C packaging. 4 N? keeps only sale notifications turned on is preference/behavior, own stance? clear preference implied, C. app notifications existing settled criteria, no change. 5 C delivery pickup? category delivery subcategory pickup likely preference delivery over pickup weekends. Existing pickup summary settled criteria, no change. 6 C delivery pickup? store exchanges, handling exchanges in store right away. Could category support? delivery pickup node is pickup logistics, but this is exchange preference in store. New category? likely delivery/pickup? Node taxonomy categories given, choose delivery pickup perhaps. Existing pickup exact fits store pickup? This is exchanges, not pickup. Could new category \"returns\" subcategory \"exchanges\" summary. 7 C deals coupons existing. 8 C product photos existing. 9 C app new subcategory ordering? convenience app ordering morning commute. New subcategory \"ordering\" summary. 10 C delivery pickup? store near work morning due crowds, likely pickup? This is store visit timing, new category? delivery pickup perhaps pickup node, but no pickup order stated. Could category delivery pickup as store pickup? Remark says goes to store, likely shopping behavior, not delivery. New category \"store\" subcategory \"visit timing\". Need categorize sensible. 6 exchanges in store right away perhaps support / exchanges new subcategory. 10 store visit timing new category. Ensure summaries tendency. 2 T. 4 C. 9 C."
        }
      ],
      "phase": "commentary",
      "role": "assistant"
    },
    {
      "id": "msg_07f3e2e44586534f016abb87dcfa7887d0a3dc11af0b2b2350",
      "type": "message",
      "status": "completed",
      "content": [
        {
          "type": "output_text",
          "annotations": [],
          "logprobs": [],
          "text": "{\"entries\":[{\"id\":1,\"stance\":\"C\",\"category\":\"delivery\",\"subcategory\":\"speed\",\"summary\":null},{\"id\":2,\"stance\":\"T\",\"category\":\"support\",\"subcategory\":\"response time\",\"summary\":null},{\"id\":3,\"stance\":\"C\",\"category\":\"delivery\",\"subcategory\":\"packaging\",\"summary\":null},{\"id\":4,\"stance\":\"C\",\"category\":\"app\",\"subcategory\":\"notifications\",\"summary\":null},{\"id\":5,\"stance\":\"C\",\"category\":\"delivery\",\"subcategory\":\"pickup\",\"summary\":null},{\"id\":6,\"stance\":\"C\",\"category\":\"support\",\"subcategory\":\"exchanges\",\"summary\":\"The client prefers handling exchanges in the store right away.\"},{\"id\":7,\"stance\":\"C\",\"category\":\"deals\",\"subcategory\":\"coupons\",\"summary\":null},{\"id\":8,\"stance\":\"C\",\"category\":\"product\",\"subcategory\":\"photos\",\"summary\":null},{\"id\":9,\"stance\":\"C\",\"category\":\"app\",\"subcategory\":\"ordering\",\"summary\":\"The client finds ordering in the app during the morning commute most convenient.\"},{\"id\":10,\"stance\":\"C\",\"category\":\"store\",\"subcategory\":\"visit timing\",\"summary\":\"The client visits the store near work in the morning to avoid crowded evenings.\"}]}"
        }
      ],
      "phase": "final_answer",
      "role": "assistant"
    }
  ],
  "parallel_tool_calls": true,
  "presence_penalty": 0,
  "previous_response_id": null,
  "prompt_cache_key": "repro-demo",
  "prompt_cache_retention": "24h",
  "reasoning": {
    "context": "all_turns",
    "effort": "none",
    "mode": "standard",
    "summary": "detailed"
  },
  "safety_identifier": "repro-demo",
  "service_tier": "default",
  "store": false,
  "temperature": 1,
  "text": {
    "format": {
      "type": "json_schema",
      "description": null,
      "name": "remark_filing",
      "schema": {
        "$defs": {
          "Entry": {
            "description": "One remark's filing entry, in the order the filing call generated them.",
            "properties": {
              "id": {
                "title": "Id",
                "type": "integer"
              },
              "stance": {
                "description": "Whether a remark carries the client's own stance: a clear one, a thin one, or none at all.",
                "enum": [
                  "C",
                  "T",
                  "N"
                ],
                "title": "Stance",
                "type": "string"
              },
              "category": {
                "title": "Category",
                "type": "string"
              },
              "subcategory": {
                "title": "Subcategory",
                "type": "string"
              },
              "summary": {
                "anyOf": [
                  {
                    "type": "string"
                  },
                  {
                    "type": "null"
                  }
                ],
                "title": "Summary"
              }
            },
            "required": [
              "id",
              "stance",
              "category",
              "subcategory",
              "summary"
            ],
            "title": "Entry",
            "type": "object",
            "additionalProperties": false
          }
        },
        "description": "The filing call's whole answer: one entry per remark.",
        "properties": {
          "entries": {
            "items": {
              "$ref": "#/$defs/Entry"
            },
            "title": "Entries",
            "type": "array"
          }
        },
        "required": [
          "entries"
        ],
        "title": "Answer",
        "type": "object",
        "additionalProperties": false
      },
      "strict": true
    },
    "verbosity": "medium"
  },
  "tool_choice": "auto",
  "tool_usage": {
    "image_gen": {
      "input_tokens": 0,
      "input_tokens_details": {
        "image_tokens": 0,
        "text_tokens": 0
      },
      "output_tokens": 0,
      "output_tokens_details": {
        "image_tokens": 0,
        "text_tokens": 0
      },
      "total_tokens": 0
    },
    "web_search": {
      "num_requests": 0
    }
  },
  "tools": [],
  "top_logprobs": 0,
  "top_p": 0.001,
  "truncation": "disabled",
  "usage": {
    "input_tokens": 678,
    "input_tokens_details": {
      "cache_write_tokens": 0,
      "cached_tokens": 0
    },
    "output_tokens": 622,
    "output_tokens_details": {
      "reasoning_tokens": 0
    },
    "total_tokens": 1300
  },
  "user": null,
  "metadata": {}
}

Repro and its output

You’ll find that I ran this at top_p:0.001, less variance in sampling than temperature adjustment, which then comes almost entirely from model non-determinism. Less trials needed if the symptom does come about immediately.

MyPlayground Serverless preset share

I can state that the prompting itself is poor. “You file remarks a summarizer wrote about one client into the client’s stored service notes.”?

You’ll have much better performance if the system message is a role and a job position for the AI, and the user message is describing the task. Then you can provide the data. Talk normally to the AI. Here you are treating the AI model like it is fine-tuned to “process” by putting only data as a user message.

Thank you for your comment. Please note that this code snippet was generated using AI for reproduction purposes, as I don’t want disclose our internal production prompts.

Try: let the model reason. For that, don’t send untolerated sampling parameters such as a modification of temperature.

Allowing reasoning.effort of “none” on new mini models might be an API validation fluke, as that is not tolerated against gpt-6-astra.