# How to use cached\_tokens field to calculate cost estimation

**URL:** <https://community.openai.com/t/how-to-use-cached-tokens-field-to-calculate-cost-estimation/1330187>\
**Category:** API\
**Tags:** pricing, api-costs, prompt-caching, cache\
**Created:** [July 31, 2025, 3:52pm UTC](https://community.openai.com/t/how-to-use-cached-tokens-field-to-calculate-cost-estimation/1330187 "2025-07-31T15:52:44Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mesut\_Celik](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mesut_celik/32/511673_2.png) [@Mesut\_Celik](https://community.openai.com/u/Mesut_Celik)\
**Post date:** [July 31, 2025, 3:52pm UTC](https://community.openai.com/t/how-to-use-cached-tokens-field-to-calculate-cost-estimation/1330187/1 "2025-07-31T15:52:44Z")

</div>

Hi,

I am trying to make sense usage data in each prompt and how to calculate cost.  
“cached\_tokens” field is complete mystery to me. I am attaching a series of usage json below. I have 3 questions.

1. How to calculate cost for each prompt using input token + cached token pricing.
2. How is it possible that cached\_tokens = 0 in the middle of chat session. (see below)
3. How is it possible that input\_tokens is less than cached\_tokens in one case below. I thought cached\_tokens field is subset of input\_tokens field.

{ **“input\_tokens”** =\> **4856** , **“total\_tokens”** =\> **4887** , **“output\_tokens”** =\> **31** , **“input\_tokens\_details”** =\> {`"cached_tokens" => 0`}, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **4895** , **“total\_tokens”** =\> **4928** , **“output\_tokens”** =\> **33** , **“input\_tokens\_details”** =\> { **“cached\_tokens”** =\> **4776** }, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **4936** , **“total\_tokens”** =\> **4974** , **“output\_tokens”** =\> **38** , **“input\_tokens\_details”** =\> { **“cached\_tokens”** =\> **4904** }, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **4989** , **“total\_tokens”** =\> **5119** , **“output\_tokens”** =\> **130** , **“input\_tokens\_details”** =\> {`"cached_tokens" => 0`}, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **5136** , **“total\_tokens”** =\> **5337** , **“output\_tokens”** =\> **201** , **“input\_tokens\_details”** =\> {`"cached_tokens" => 10192`}, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **5652** , **“total\_tokens”** =\> **5878** , **“output\_tokens”** =\> **226** , **“input\_tokens\_details”** =\> { **“cached\_tokens”** =\> **5288** }, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **6004** , **“total\_tokens”** =\> **6055** , **“output\_tokens”** =\> **51** , **“input\_tokens\_details”** =\> { **“cached\_tokens”** =\> **5800** }, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **6122** , **“total\_tokens”** =\> **6197** , **“output\_tokens”** =\> **75** , **“input\_tokens\_details”** =\> {`"cached_tokens" => 0`}, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

{ **“input\_tokens”** =\> **6491** , **“total\_tokens”** =\> **6537** , **“output\_tokens”** =\> **46** , **“input\_tokens\_details”** =\> { **“cached\_tokens”** =\> **6184** }, **“output\_tokens\_details”** =\> { **“reasoning\_tokens”** =\> **0** }},

{},

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [July 31, 2025, 4:57pm UTC](https://community.openai.com/t/how-to-use-cached-tokens-field-to-calculate-cost-estimation/1330187/2 "2025-07-31T16:57:40Z")

</div>

Let’s just search things I’ve said before…

> [@OpenAI Agent SDK – Token usage and reasoning output](https://community.openai.com/t/openai-agent-sdk-token-usage-and-reasoning-output/1316366/4):
>
> You’ve got it! Prompt tokens and Output tokens are your billing at the model’s rate. Because some output was reasoning output doesn’t change the cost of output tokens. The usage detail can modify the model costs in other categories, though. For example, cached input would provide the portion of the prompt tokens that have been discounted, and audio tokens (chat completions) are the portions billed at the audio rates (approximately 10 times higher).

Some code developed for demonstrating a model call’s final price, by its token costs and the returned usage and its cached\_tokens field (not considering audio token costs).

> [@How to get the cost for each api call?](https://community.openai.com/t/how-to-get-the-cost-for-each-api-call/1227787/3):
>
> The billing is in terms of tokens. You can then directly translate that to the input and output cost of the model you are employing - the cost being divided by one million then multiplied by your usage. Just typing up a request to an API endpoint.. ... import os, httpx \>\>\> def send\_chat\_request(conversation\_messages): ... # Set your API key in OPENAI\_API\_KEY env variable (never hard-code it). ... api\_key = os.environ.get("OPENAI\_API\_KEY") ... if not api\_key: ... raise Value…

Note that “cached” is not ultimately as a percentage, it is a reduced token cost specified per-model, and before discount percentage varied, that code snippet anticipated this by taking the [model pricing fields](https://platform.openai.com/docs/models).

OpenAI makes a “best effort to route to the same server”, but the context window cache can persist for a time as low as five minutes. There isn’t a central cache database, it is per AI model server.

> [@How to save input tokens in Responses API?](https://community.openai.com/t/how-to-save-input-tokens-in-responses-api/1266965/4):
>
> To receive a (non-guaranteed) discount, the entire initial input context needs to be the same. The cache is temporary after the recent call, with a lifetime around 10-60 minutes (someday I’ll classify how long is delivered). If you change initial system message instructions from “you are a helpful ai” to “you are a helpful ai with these user preferences”, then you have broken the caching. That also applies to the list of internal tools and developer-provided functions that are appended to your…

You can read details about how the same server is “found”…

> [@Prompt Cache Routing + the \`user\` Parameter](https://community.openai.com/t/prompt-cache-routing-the-user-parameter/1267103):
>
> We just updated our [prompt caching guide](https://platform.openai.com/docs/guides/prompt-caching#how-it-works) with details of how cache routing works! We route prompts by org, hashing the first ~256 tokens, and spillover to more machines ~15 RPM. If you’ve got many prompts with long shared prefixes, the user parameter can improve request bucketing and boost cache hits
