Gpt-5.6-luna: Tool-call bypass & internal reasoning leaking into tool arguments in production

TL;DR: We’re running a production agentic system using gpt-5.6-luna with structured tool calling (strict: true). We’ve hit two distinct failure modes where the model either (1) leaks its own internal reasoning/self-instructions into tool call argument values, or (2) completely bypasses the tool call and dumps the intended tool arguments as raw text in the response content — sometimes prefixed with garbage CJK punctuation characters. Both issues resulted in malformed content being delivered to end users.


Our Setup

  • Model: gpt-5.6-luna
  • Framework: LangGraph + LangChain (Python)
  • Tool binding: model.bind_tools(tools, strict=True)
  • Use case: Customer-facing sales agent on Facebook Messenger. The agent has multiple tools (knowledge search, order capture, etc.) and a mandatory final_response tool that must be called as the last action to deliver a structured reply to the customer.

The final_response tool schema (simplified):

{
  "name": "final_response",
  "parameters": {
    "messages": [
      {
        "type": enum("text", "image", "file"),
        "message": "string | None — the text to send to the customer"
        "url": str | None
      }
    ],
    "is_reply_need": true
  }
}

The system prompt explicitly states:

“You MUST ALWAYS call the final_response tool to complete your turn — even when no reply is needed. NEVER output raw text tokens directly.”


Issue 1: Internal reasoning leaked into tool argument value

What happened

The model correctly called the final_response tool, but concatenated its own internal meta-commentary / self-instruction into the message text field:

User message: “আসসালামু আলাইকুম ভাই আপনাদের ডিভাইসটি মূল্য কত” (Bengali: “Assalamu Alaikum brother, what is the price of your device?”)

What the tool call contained (and what the customer received):

ওয়ালাইকুম আসসালাম, স্যার। কোন ডিভাইসটির দাম জানতে চান? নাম বা ছবি পাঠালে জানিয়ে দিচ্ছি। 😊‌এখনই final_response কল করুন? (Note: already calling)

Translation of the leaked portion:

“Should I call final_response now? (Note: already calling)”

This is clearly the model’s inner monologue about whether to invoke the tool — it somehow ended up as part of the tool’s argument value.

Observations

  • The tools_used field confirms final_response was called (not bypassed).
  • The model reasoning block shows: “I think it’s important to start with a warm greeting… Once I get that clarification, I can proceed with a final response tailored to their needs.”
  • The meta-text is half Bengali, half English — mixing the agent’s configured language with what appears to be internal reasoning.
  • Token count was low (5,922 input / 112 output) — this was a simple greeting exchange, not a complex multi-tool scenario.

Issue 2: Tool call completely bypassed — raw JSON emitted as text with garbage prefix

What happened

The model did not call any tools at all (tools_used: []) and instead emitted the entire final_response JSON structure as plain text in the AIMessage.content, prepended with CJK punctuation characters:

User message: “Hi”

What the model output as raw text content (and what the customer received):

、】【{"messages":[{"type":"text","message":"Hello Ma'am, welcome to GOOD ONE! How may I help you today?"},{"type":"text","message":"Are you looking for a specific product or something to solve an everyday problem?"}],"is_reply_need":true}

Observations

  • tools_used is empty — the model generated zero tool calls despite strict: true binding.
  • The JSON structure inside the text is perfectly valid final_response schema — the model clearly “knew” what to output but chose to emit it as text instead of a tool call.
  • The prefix 、】【 consists of CJK punctuation characters (ideographic comma, right/left corner brackets). These appear to be hallucinated artifacts — they have no relation to the conversation content, the user’s language (Bengali/English), or any system prompt content.
  • This was also a low-token, simple exchange (5,906 input / 61 output).
  • Both issues occurred on the same agent configuration, same page, within minutes of each other (same day).

Impact

Both issues resulted in malformed content being sent to real customers on Facebook Messenger. We’ve built multiple layers of fallback parsing (JSON extraction from text, plain-text chunking, empty-literal detection) but these specific patterns slipped through our defenses.

Questions for the Community

  1. How can I solve this two issue?
  2. Has anyone else observed gpt-5.6-luna leaking its reasoning/self-instructions into tool call argument values? This seems distinct from the reasoning block — it’s the model conflating its “thinking about what to do” with the actual content it produces.
  3. Has anyone seen the model emit tool-call-shaped JSON as plain text content instead of making an actual tool call? Especially with strict: true — we expected this to enforce that the model always uses the tool calling mechanism rather than outputting JSON text.
  4. What’s with the CJK punctuation prefix (、】【)? This is particularly puzzling. The conversation is in Bengali/English, the system prompt is in English, and there’s no CJK content anywhere in the context. Has anyone encountered similar hallucinated character prefixes?
  5. Are there any recommended mitigations beyond post-processing? We’ve added heuristic filters, but fundamentally these are model-level failures that shouldn’t need application-layer workarounds.

Environment Details

  • Model: gpt-5.6-luna (via OpenAI API)

  • API version: Latest as of September 2026

  • Framework: LangChain 0.3.x + LangGraph 0.4.x

  • Tool calling mode: strict: true (structured outputs)

  • Date of incidents: September 26, 2026

  • Frequency: Intermittent — most responses work correctly; but 1% response is like this

    Any insights or similar experiences would be greatly appreciated. Happy to share additional details if helpful.

1 Like