Environment
- Surface: API
- Endpoint or feature: Responses API (
POST /v1/responses) with structured outputs - Model: gpt-6-luna
- SDK and version, if applicable: openai-python 3.20.0 (also seen with 2.44.0)
Bug
What happened, and what did you expect?
gpt-6-luna occasionally returns an additional message output item for a structured output request. Only one item holds the schema-valid JSON. The additional item is empty, holds free-text working notes (“We need output JSON schema. classify stance. …”), or repeats the JSON with stray text such as </assistant:commentary> after the first value. Expected: one message item whose output_text is the JSON value.
Because of the additional item, response.output_text is not one JSON value and client.responses.parse(..., text_format=Model) raises Invalid JSON.
The response format name decides how often it happens. With the request below, remark_filing reproduces within a few calls; RemarkFiling once in 50; Answer, answer, and filing_answer never in 50 to 100 calls.
Reproduction
Minimal code, request, or steps to reproduce:
"""gpt-6-luna occasionally returns an additional message item for structured outputs.
pip install openai pydantic; set OPENAI_API_KEY."""
from typing import Any, Literal
from openai import OpenAI
from pydantic import BaseModel, Field
SYSTEM = """\
You file remarks a summarizer wrote about one client into the client's stored service notes.
Each remark states one specific plus the client's own stance on it.
For each remark answer three things: whether it carries the client's stance, which node it belongs under, and a summary of that node when the node is new or its tendency changed.
- C: a concrete thing plus the client's own evaluation, reason, comparison, preference, or avoidance.
- T: an evaluation is present but thin. The default when unsure.
- N: no evaluation of the client's own.
A node is a category or a subcategory under it. Reuse an existing subcategory's exact name when it fits; otherwise open a new one.
Write a summary for a new subcategory; for a listed one, null unless the remark changes its tendency.
Refer to the client only as "the client".
<catalog> lists the nodes: `##` is a category, `###` is a subcategory with its summary.
<remarks> holds the remarks, numbered; cover every remark once, in input order.
"""
USER = """\
<catalog>
## delivery
### speed — Settled criteria; rarely changes them.
### packaging — Settled criteria; rarely changes them.
### pickup — Settled criteria; rarely changes them.
## support
### response time — Settled criteria; rarely changes them.
## app
### notifications — Settled criteria; rarely changes them.
### search — Settled criteria; rarely changes them.
## product
### photos — Settled criteria; rarely changes them.
## deals
### coupons — Settled criteria; rarely changes them.
</catalog>
<remarks>
[1] Prefers delivery within two days; needs the items for the weekend.
[2] Finds support replies that take over a day frustrating.
[3] Prefers thick padding in packages; glass items keep breaking.
[4] Keeps only sale notifications turned on.
[5] Finds delivery easier than store pickup on weekends.
[6] Prefers handling exchanges in the store right away.
[7] Skips coupons with complicated conditions and pays full price.
[8] Trusts review photos more than product photos; they show the real color.
[9] Finds ordering in the app during the morning commute most convenient.
[10] Goes to the store near work in the morning because evenings are too crowded.
</remarks>
"""
class Entry(BaseModel):
"""One remark's filing entry, in the order the filing call generated them."""
id: int
stance: Literal["C", "T", "N"] = Field(
description="Whether a remark carries the client's own stance: a clear one, a thin one, or none at all."
)
category: str
subcategory: str
summary: str | None
class Answer(BaseModel):
"""The filing call's whole answer: one entry per remark."""
entries: list[Entry]
def strict_schema(model: type[BaseModel]) -> dict[str, Any]:
"""`model_json_schema()` with `additionalProperties: false` on every object, as strict mode requires."""
def walk(node: Any) -> Any:
if isinstance(node, dict):
out = {key: walk(value) for key, value in node.items()}
if out.get("type") == "object":
out["additionalProperties"] = False
return out
if isinstance(node, list):
return [walk(item) for item in node]
return node
return walk(model.model_json_schema())
client = OpenAI()
text_format = {"type": "json_schema", "name": "remark_filing", "strict": True, "schema": strict_schema(Answer)}
for call in range(1, 51):
response = client.responses.create(
model="gpt-6-luna",
input=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": USER}],
text={"format": text_format},
reasoning={"effort": "none"},
temperature=0.2,
max_output_tokens=4096,
store=False,
)
messages = [item for item in response.output if item.type == "message"]
if len(messages) == 1:
print(f"call {call}: ok")
continue
print(f"call {call}: {len(messages)} message items, id={response.id}")
for index, item in enumerate(messages):
print(f"--- output_text[{index}] ---")
print(item.content[0].text)
break
This reproduced on call 5 (10 remarks) and on call 1 with 15 remarks. Changing only "name": "remark_filing" to "name": "Answer" did not reproduce in 100 calls.
Diagnostics
- Exact error and HTTP status:
HTTP 200,status: "completed",reasoning_tokens: 0.
When the extra item is free text,client.responses.parseraisespydantic_core.ValidationError: Invalid JSON. - Request ID, if available:
resp_0f93f0115237ff4b016abb8035d4e087d0a727d5b126237500 - When it happened (with timezone) and how often:
continuously since 2026-09-28 (KST, UTC+9), when we moved to gpt-6-luna, at about 3 to 5 in 100 calls on our production request.
The request IDs above are from 2026-09-29, 16:12, 18:08, and 18:09 KST.
Also seen with streaming and with temperature 1.0.