# Not enough tokens error, even though I've paid A LOT (maximum context length error)

**URL:** <https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140>\
**Category:** API\
**Tags:** api\
**Created:** [September 9, 2023, 4:44pm UTC](https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140 "2023-09-09T16:44:19Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![inthemoment421](https://avatars.discourse-cdn.com/v4/letter/i/ce7236/32.png) [@inthemoment421](https://community.openai.com/u/inthemoment421)\
**Post date:** [September 9, 2023, 4:44pm UTC](https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140/1 "2023-09-09T16:44:19Z")

</div>

I’m posting this in the developer forum out of exasperation with OpenAI’s customer service.

My interaction with OpenAI customer support has been beyond terrible. I understand it’s a startup, but I hope they’ll devote some time and attention to making customer service work for their customers.

Also, their customer service bot ends the conversation upon uploading a screenshot. So, it’s actually impossible to attach screenshots or provide documentation (see the below screenshot).

Basically, I’m getting an error message shown in screenshot 2 that I’m lacking about 3k tokens to complete the request. However, the OpenAPI website (screenshot 3) says I have $50+ balance, which is more than enough tokens.

Due to OpenAI’s customer services’ inability to understand the issue, I’m going to try my best to over-communicate and over-document. The problem is that OpenAI fails to deliver the services for which I’ve paid. I’ve tried my best to explain to them that I have paid OpenAI. But, they seem unable to comprehend that despite the fact that OpenAI’s been charging my credit card for weeks (see screenshot 1) and the OpenAI website says I’ve paid screenshot 3.

 ![screenshot2](https://us1.discourse-cdn.com/openai1/original/3X/e/a/ea8757f37c48aa421d4ab5b631f10ebbe2daec03.png)

And my frustration continues, as ‘new users are only able to post one screenshot.’ I don’t understand why OpenAI seems so determined to make it impossible to report issues and bugs!!

---

<div class="post-metadata">

**Author:** ![PaulBellow](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/paulbellow/32/597962_2.png) [@PaulBellow](https://community.openai.com/u/PaulBellow)\
**Post date:** [September 9, 2023, 4:53pm UTC](https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140/2 "2023-09-09T16:53:47Z")

</div>

Welcome to the forum.

For the model you’re using, the prompt + completion can only be 4097 tokens, and you’re sending 7k… either change to a 16k or 32k model if you have access or make your prompt shorter…

Good luck!

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [September 9, 2023, 4:56pm UTC](https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140/3 "2023-09-09T16:56:07Z")

</div>

Resources to understand what you are doing wrong:

> [@Not allowed to have all 8192 tokens](https://community.openai.com/t/not-allowed-to-have-all-8192-tokens/321500/11):
>
> You understood wrong. max\_tokens is the limit of the response you will get back. max\_tokens also reserves space exclusively for this response formation. The context length of a model is first loaded with the input, and then the tokens that the AI generates are added after that, in the remaining space. Language is formed in a transformer language model by continuing the next token that should logically appear one at a time based on previous input and generated response so far. (In an ideal…

> [@Do 'MAX tokens' include the follow up prompts and completion in a single chat session](https://community.openai.com/t/do-max-tokens-include-the-follow-up-prompts-and-completion-in-a-single-chat-session/328955/7):
>
> There are two things that might be conflated here: A model’s context length is the total amount of tokens it can handle at once, a combined count of both the total input that you send and the response you get back. An API call’s max\_tokens parameter reserves a specific amount of the context length for forming an answer, and sets the size of the maximum response that you will receive back from the AI. “Followup responses” to me means that you have a chatbot application (instead of just making…

---

<div class="post-metadata">

**Author:** ![jwatte](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/jwatte/32/127705_2.png) [@jwatte](https://community.openai.com/u/jwatte)\
**Post date:** [September 9, 2023, 7:09pm UTC](https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140/4 "2023-09-09T19:09:31Z")

</div>

The 4097 token length is a technical limitation for the gpt-3.5-turbo model. The prompts plus the response cannot add up to more than that size; there’s not enough space in the model. It doesn’t matter whether you pay money or not.

If you want to make inference with context and answers that are bigger than this, you need to use a model that has more context size. gpt-3.5-turbo-16k is one option; gpt-4 is another. Both of them are big enough for 7442 tokens.

---

<div class="post-metadata">

**Author:** ![m1m2](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/m1m2/32/195733_2.png) [@m1m2](https://community.openai.com/u/m1m2)\
**Post date:** [September 9, 2023, 9:20pm UTC](https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140/5 "2023-09-09T21:20:43Z")

</div>

Thanks for the recommendations in this thread, I too have been looking for a solution to the problem described by the topikstarter for several days. Now it became clear to me how to solve this problem.

You need to use a model with a large context size: gpt-3.5-turbo-16k and write in the code instead of: “max\_tokens = min(max\_tokens, 4096) # OpenAI token limit” a new parameter: “max\_tokens = min(max\_tokens, 16 385)”.

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [September 9, 2023, 11:13pm UTC](https://community.openai.com/t/not-enough-tokens-error-even-though-ive-paid-a-lot-maximum-context-length-error/362140/6 "2023-09-09T23:13:16Z")

</div>

No, incorrect. `max_tokens` is how much response you want to get back.

Or if you are unable to figure it out, just omit the parameter entirely.
