Thanks for following up on this three-month issue.
Where reproduction yourself takes under a minute, and my reproduction was one minute ago.
Here is what receiving no cache discount looks like on gpt-5.4-nano calls 30 seconds apart - now with the default backport of 24h policy against prior “in-memory” models:
Note the usage report in the assistant header bar - more than enough input to store, no cached tokens in the usage report
(ignore the poor quality of the excessively instructed “improve language” task not treating input messages as discrete input->output cases, I guess…).
x-request-id: req_055e2a935d02408fa37fb0ffa1e2d383
Full API response object
{
"id": "resp_0ddcbb620d1c0713016a54fa4e0b508199bb8ad393a931a89f",
"object": "response",
"created_at": 1783953998,
"status": "completed",
"background": false,
"billing": {
"payer": "openai"
},
"completed_at": 1783953999,
"error": null,
"frequency_penalty": 0,
"incomplete_details": null,
"instructions": null,
"max_output_tokens": 16384,
"max_tool_calls": null,
"model": "gpt-5.4-nano-2026-03-17",
"moderation": null,
"output": [
{
"id": "msg_0ddcbb620d1c0713016a54fa4e7cd481998cc061b16181474e",
"type": "message",
"status": "completed",
"content": [
{
"type": "output_text",
"annotations": [],
"logprobs": [],
"text": "Please run the same test again. Please share 2–3 `x-request-id` values from GPT-5.4 Nano responses. Ideally, those requests should be from the last three days.\n\nAlso, please share one GPT-5.4 Mini `x-request-id` from the same test as a comparison. These `x-request-id` values will help us trace the exact requests. They will also help us investigate cache behavior more deeply. Finally, they will give our engineering team concrete examples to review if needed."
}
],
"phase": "final_answer",
"role": "assistant"
}
],
"parallel_tool_calls": true,
"presence_penalty": 0,
"previous_response_id": null,
"prompt_cache_key": "myplayground",
"prompt_cache_retention": "24h",
"reasoning": {
"context": "current_turn",
"effort": "low",
"mode": "standard",
"summary": null
},
"safety_identifier": null,
"service_tier": "default",
"store": false,
"temperature": 1,
"text": {
"format": {
"type": "text"
},
"verbosity": "low"
},
"tool_choice": "auto",
"tool_usage": {
"image_gen": {
"input_tokens": 0,
"input_tokens_details": {
"image_tokens": 0,
"text_tokens": 0
},
"output_tokens": 0,
"output_tokens_details": {
"image_tokens": 0,
"text_tokens": 0
},
"total_tokens": 0
},
"web_search": {
"num_requests": 0
}
},
"tools": [],
"top_logprobs": 0,
"top_p": 0.98,
"truncation": "disabled",
"usage": {
"input_tokens": 1650,
"input_tokens_details": {
"cache_write_tokens": 0,
"cached_tokens": 0
},
"output_tokens": 109,
"output_tokens_details": {
"reasoning_tokens": 0
},
"total_tokens": 1759
},
"user": null,
"metadata": {}
}
The request for information is for a comparative “mini” call, and its request ID?
Cache discount delivery. Although an under-performing 1280/1544. The uncached/unstored 264 tokens being 2x greater than the granularity expected before.
x-request-id: req_040d258d86434dcaadb344d2f3bb5330