# \#cache

**URL:** https://community.openai.com/tag/cache/822.md

[Latest](https://community.openai.com/latest.md) · [Categories](https://community.openai.com/categories.md) · [Tags](https://community.openai.com/tags.md)

---

## [Does prompt\_cache\_key guarantee that two calls will have distinct caches?](https://community.openai.com/t/does-prompt-cache-key-guarantee-that-two-calls-will-have-distinct-caches/1389870)

<div class="topic-metadata">

**Author:** [@sousuoyinqing](https://community.openai.com/u/sousuoyinqing)\
**Replies:** 2\
**Last updated:** [August 11, 2026, 6:37pm UTC](https://community.openai.com/t/does-prompt-cache-key-guarantee-that-two-calls-will-have-distinct-caches/1389870 "2026-08-11T18:37:43Z")

</div>

I make a first LLM call with more than 1024 input tokens with \`prompt\_cache\_key=abc\`. I made a second LLM call with more than 1024 input tokens with ‘prompt\_cache\_key=abc’. Both LLM calls have strictly the same system\_…

---

## [GPT-5.6 Cache writes and retrieval counter to (lack of) documentation, expectation](https://community.openai.com/t/gpt-5-6-cache-writes-and-retrieval-counter-to-lack-of-documentation-expectation/1387644)

<div class="topic-metadata">

**Author:** [@\_j](https://community.openai.com/u/_j)\
**Replies:** 1\
**Last updated:** [July 21, 2026, 3:52am UTC](https://community.openai.com/t/gpt-5-6-cache-writes-and-retrieval-counter-to-lack-of-documentation-expectation/1387644 "2026-07-21T03:52:12Z")

</div>

Issues with gpt-5.6 model cache writes and cache retrieval discounts broken expectation that any written cache would be retrieved always broken expectation that any “implicit” call would be stored for cache if not a hit …

---

## [Gpt-5.6 prompt caching fails on partial prefixes](https://community.openai.com/t/gpt-5-6-prompt-caching-fails-on-partial-prefixes/1386887)

<div class="topic-metadata">

**Author:** [@borna\_najafi](https://community.openai.com/u/borna_najafi)\
**Replies:** 5\
**Last updated:** [July 15, 2026, 10:26am UTC](https://community.openai.com/t/gpt-5-6-prompt-caching-fails-on-partial-prefixes/1386887 "2026-07-15T10:26:23Z")

</div>

Bug report: gpt-5.6 prompt caching never matches partial prefixes (only exact full-prompt matches) Product: OpenAI API — Prompt Caching Models affected: gpt-5.6-sol, gpt-5.6-luna (both confirmed) Models NOT affected (c…

---

## [Caching rate drop after switching to Responses API](https://community.openai.com/t/caching-rate-drop-after-switching-to-responses-api/1370788)

<div class="topic-metadata">

**Author:** [@Dobo](https://community.openai.com/u/Dobo)\
**Replies:** 2\
**Last updated:** [January 9, 2026, 12:31am UTC](https://community.openai.com/t/caching-rate-drop-after-switching-to-responses-api/1370788 "2026-01-09T00:31:29Z")

</div>

Recently switched app backend to Responses API and migrated from gpt-5 to gpt-5.1. Seeing cache rate (% cached input tokens) dropping from 40% to 0% after this change. What am I doing wrong? Does caching work differen…

---

## [Caching is borked for GPT 5 models](https://community.openai.com/t/caching-is-borked-for-gpt-5-models/1359574)

<div class="topic-metadata">

**Author:** [@stroop](https://community.openai.com/u/stroop)\
**Replies:** 19\
**Last updated:** [January 8, 2026, 8:20pm UTC](https://community.openai.com/t/caching-is-borked-for-gpt-5-models/1359574 "2026-01-08T20:20:10Z")

</div>

Trying to call the GPT5 model with the responses API. My system prompt is consistent and is longer than 1024 tokens but I don’t hit the cache for any of the GPT5 model series input\_tokens\_details=InputTokensDetails(cach…

---

## [Realtime caching between sessions?](https://community.openai.com/t/realtime-caching-between-sessions/1366739)

<div class="topic-metadata">

**Author:** [@andrew\_tomis.tech](https://community.openai.com/u/andrew_tomis.tech)\
**Replies:** 0\
**Last updated:** [November 18, 2025, 7:42pm UTC](https://community.openai.com/t/realtime-caching-between-sessions/1366739 "2025-11-18T19:42:32Z")

</div>

I’m wondering if there is a way to maintain cached tokens between gpt-realtime sessions. For something like instructions or a large system message that is the same for each realtime session, can a new session take advant…

---

## [Can I cache large chunks on gpt-5-nano?, Does each cache-read request reset cache inactive time?, Does large caches affect cache overflow limits?](https://community.openai.com/t/can-i-cache-large-chunks-on-gpt-5-nano-does-each-cache-read-request-reset-cache-inactive-time-does-large-caches-affect-cache-overflow-limits/1363863)

<div class="topic-metadata">

**Author:** [@Tasmay\_Tibrewal](https://community.openai.com/u/Tasmay_Tibrewal)\
**Replies:** 1\
**Last updated:** [October 25, 2025, 11:17am UTC](https://community.openai.com/t/can-i-cache-large-chunks-on-gpt-5-nano-does-each-cache-read-request-reset-cache-inactive-time-does-large-caches-affect-cache-overflow-limits/1363863 "2025-10-25T11:17:23Z")

</div>

Hey there, It is mentioned on the website (https://platform.openai.com/docs/guides/prompt-caching#page-top) that if we are putting in more than say 15 requests per minute, then that would mean that would lead to cache o…

---

## [Input cache not registering with fine-tuned gpt-4o](https://community.openai.com/t/input-cache-not-registering-with-fine-tuned-gpt-4o/1362629)

<div class="topic-metadata">

**Author:** [@Cemal\_Yilmaz](https://community.openai.com/u/Cemal_Yilmaz)\
**Replies:** 0\
**Last updated:** [October 15, 2025, 4:46pm UTC](https://community.openai.com/t/input-cache-not-registering-with-fine-tuned-gpt-4o/1362629 "2025-10-15T16:46:58Z")

</div>

Hi, I’m trying to get my fine-tuned model on the platform to use input caching, but even though I have a static, and long system prompt (around 1200 tokens), it doesn’t use caching and I have to pay the full monies unne…

---

## [How to use cached\_tokens field to calculate cost estimation](https://community.openai.com/t/how-to-use-cached-tokens-field-to-calculate-cost-estimation/1330187)

<div class="topic-metadata">

**Author:** [@Mesut\_Celik](https://community.openai.com/u/Mesut_Celik)\
**Replies:** 1\
**Last updated:** [July 31, 2025, 4:57pm UTC](https://community.openai.com/t/how-to-use-cached-tokens-field-to-calculate-cost-estimation/1330187 "2025-07-31T16:57:40Z")

</div>

Hi, I am trying to make sense usage data in each prompt and how to calculate cost. “cached\_tokens” field is complete mystery to me. I am attaching a series of usage json below. I have 3 questions. How to calculate co…

---

## [Cached Input Tokens in Chat Completions](https://community.openai.com/t/cached-input-tokens-in-chat-completions/1238755)

<div class="topic-metadata">

**Author:** [@ashwinthandu03](https://community.openai.com/u/ashwinthandu03)\
**Replies:** 1\
**Last updated:** [April 30, 2025, 1:15am UTC](https://community.openai.com/t/cached-input-tokens-in-chat-completions/1238755 "2025-04-30T01:15:09Z")

</div>

Here, I recently went through this ‘Input cached tokens’ in my chat completions api… What is the cached tokens exactly? Is the api caching the entire input tokens for retaining the context in every request? or is it cac…

---

## [Prompt Caching in Batching API](https://community.openai.com/t/prompt-caching-in-batching-api/1223191)

<div class="topic-metadata">

**Author:** [@likhitha](https://community.openai.com/u/likhitha)\
**Replies:** 2\
**Last updated:** [April 6, 2025, 10:17am UTC](https://community.openai.com/t/prompt-caching-in-batching-api/1223191 "2025-04-06T10:17:44Z")

</div>

I’m using GPT-4o for image analysis, and my prompt is quite large (approximately 3,000 tokens). Can I cache this prompt and use it in Batch deployment mode rather than sending the same prompt against each task in the inp…

---

## [Stopped to use Cached Tokens](https://community.openai.com/t/stopped-to-use-cached-tokens/1076083)

<div class="topic-metadata">

**Author:** [@zoss](https://community.openai.com/u/zoss)\
**Replies:** 1\
**Last updated:** [January 2, 2025, 2:57am UTC](https://community.openai.com/t/stopped-to-use-cached-tokens/1076083 "2025-01-02T02:57:31Z")

</div>

Stopped to use cached tokens. When i analyze the api response log, i see about 80% of cached tokens. What could be the problem?

---

## [Caching Strategy for Different Projects](https://community.openai.com/t/caching-strategy-for-different-projects/1073799)

<div class="topic-metadata">

**Author:** [@osmanbulut](https://community.openai.com/u/osmanbulut)\
**Replies:** 4\
**Last updated:** [December 28, 2024, 7:36pm UTC](https://community.openai.com/t/caching-strategy-for-different-projects/1073799 "2024-12-28T19:36:55Z")

</div>

I have a question. Let’s say we have different projects. Different projects have different input prompts. For the best practice, should we use different API keys for different projects regarding caching inputs and boos…
