I had a short CLI conversation using GPT-6-Luna medium reasoning in fast mode.
The statusline shows 1.1 Mio Tokens In, 31.2K Tokens Out and $1,82 estimated-threat-cost.
There are different pricings for cache-writes, cache-reads and uncached input. Even if all tokens were expensive cache-writes in fast mode ($0,25/Mio Tokens with GPT-6-Luna), these costs were way too high…
Did anyone experience that these estimated costs are far higher than the actually should be? I am also asking since I noted that the costs displayed in the API console online are way higher than they should be as models were billed which I definitely did not use in this moment (see also separate topic).