# Gpt-5.6-luna: Tool-call bypass & internal reasoning leaking into tool arguments in production

**URL:** <https://community.openai.com/t/gpt-5-6-luna-tool-call-bypass-internal-reasoning-leaking-into-tool-arguments-in-production/1401196>\
**Category:** API\
**Tags:** chatgpt, function-calling, structured-output\
**Created:** [September 27, 2026, 5:23am UTC](https://community.openai.com/t/gpt-5-6-luna-tool-call-bypass-internal-reasoning-leaking-into-tool-arguments-in-production/1401196 "2026-09-27T05:23:58Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![meharaz733](https://avatars.discourse-cdn.com/v4/letter/m/48db29/32.png) [@meharaz733](https://community.openai.com/u/meharaz733)\
**Post date:** [September 27, 2026, 5:23am UTC](https://community.openai.com/t/gpt-5-6-luna-tool-call-bypass-internal-reasoning-leaking-into-tool-arguments-in-production/1401196/1 "2026-09-27T05:23:58Z")

</div>

**TL;DR:** We’re running a production agentic system using gpt-5.6-luna with structured tool calling (`strict: true`). We’ve hit two distinct failure modes where the model either (1) leaks its own internal reasoning/self-instructions into tool call argument values, or (2) completely bypasses the tool call and dumps the intended tool arguments as raw text in the response content — sometimes prefixed with garbage CJK punctuation characters. Both issues resulted in malformed content being delivered to end users.

* * *

## Our Setup

- **Model:** `gpt-5.6-luna`
- **Framework:** LangGraph + LangChain (Python)
- **Tool binding:** `model.bind_tools(tools, strict=True)`
- **Use case:** Customer-facing sales agent on Facebook Messenger. The agent has multiple tools (knowledge search, order capture, etc.) and a mandatory `final_response` tool that must be called as the last action to deliver a structured reply to the customer.

The `final_response` tool schema (simplified):

```json
{
  "name": "final_response",
  "parameters": {
    "messages": [
      {
        "type": enum("text", "image", "file"),
        "message": "string | None — the text to send to the customer"
        "url": str | None
      }
    ],
    "is_reply_need": true
  }
}

```

The system prompt explicitly states:

> _“You MUST ALWAYS call the `final_response` tool to complete your turn — even when no reply is needed. NEVER output raw text tokens directly.”_

* * *

## Issue 1: Internal reasoning leaked into tool argument value

### What happened

The model correctly called the `final_response` tool, but concatenated its own **internal meta-commentary / self-instruction** into the `message` text field:

**User message:** “আসসালামু আলাইকুম ভাই আপনাদের ডিভাইসটি মূল্য কত” _(Bengali: “Assalamu Alaikum brother, what is the price of your device?”)_

**What the tool call contained (and what the customer received):**

```auto
ওয়ালাইকুম আসসালাম, স্যার। কোন ডিভাইসটির দাম জানতে চান? নাম বা ছবি পাঠালে জানিয়ে দিচ্ছি। 😊‌এখনই final_response কল করুন? (Note: already calling)

```

Translation of the leaked portion:

> **“Should I call final\_response now? (Note: already calling)”**

This is clearly the model’s inner monologue about whether to invoke the tool — it somehow ended up as part of the tool’s argument value.

### Observations

- The `tools_used` field confirms `final_response` **was** called (not bypassed).
- The model reasoning block shows: _“I think it’s important to start with a warm greeting… Once I get that clarification, I can proceed with a final response tailored to their needs.”_
- The meta-text is **half Bengali, half English** — mixing the agent’s configured language with what appears to be internal reasoning.
- Token count was low (5,922 input / 112 output) — this was a simple greeting exchange, not a complex multi-tool scenario.

* * *

## Issue 2: Tool call completely bypassed — raw JSON emitted as text with garbage prefix

### What happened

The model **did not call any tools at all** (`tools_used: []`) and instead emitted the entire `final_response` JSON structure as plain text in the `AIMessage.content`, prepended with CJK punctuation characters:

**User message:** “Hi”

**What the model output as raw text content (and what the customer received):**

```auto
、】【{"messages":[{"type":"text","message":"Hello Ma'am, welcome to GOOD ONE! How may I help you today?"},{"type":"text","message":"Are you looking for a specific product or something to solve an everyday problem?"}],"is_reply_need":true}

```

### Observations

- `tools_used` is **empty** — the model generated zero tool calls despite `strict: true` binding.
- The JSON structure inside the text is **perfectly valid** `final_response` schema — the model clearly “knew” what to output but chose to emit it as text instead of a tool call.
- The prefix `、】【` consists of CJK punctuation characters (ideographic comma, right/left corner brackets). These appear to be hallucinated artifacts — they have no relation to the conversation content, the user’s language (Bengali/English), or any system prompt content.
- This was also a low-token, simple exchange (5,906 input / 61 output).
- Both issues occurred on the **same agent configuration, same page, within minutes of each other** (same day).

* * *

## Impact

Both issues resulted in malformed content being sent to real customers on Facebook Messenger. We’ve built multiple layers of fallback parsing (JSON extraction from text, plain-text chunking, empty-literal detection) but these specific patterns slipped through our defenses.

Questions for the Community

1. How can I solve this two issue?
2. **Has anyone else observed gpt-5.6-luna leaking its reasoning/self-instructions into tool call argument values?** This seems distinct from the reasoning block — it’s the model conflating its “thinking about what to do” with the actual content it produces.
3. **Has anyone seen the model emit tool-call-shaped JSON as plain text content instead of making an actual tool call?** Especially with `strict: true` — we expected this to enforce that the model always uses the tool calling mechanism rather than outputting JSON text.
4. **What’s with the CJK punctuation prefix (`、】【`)?** This is particularly puzzling. The conversation is in Bengali/English, the system prompt is in English, and there’s no CJK content anywhere in the context. Has anyone encountered similar hallucinated character prefixes?
5. **Are there any recommended mitigations beyond post-processing?** We’ve added heuristic filters, but fundamentally these are model-level failures that shouldn’t need application-layer workarounds.

* * *

## Environment Details

- **Model:** gpt-5.6-luna (via OpenAI API)

- **API version:** Latest as of September 2026

- **Framework:** LangChain 0.3.x + LangGraph 0.4.x

- **Tool calling mode:** `strict: true` (structured outputs)

- **Date of incidents:** September 26, 2026

- **Frequency:** Intermittent — most responses work correctly; but 1% response is like this
