# Getting 400 response with already working code

**URL:** <https://community.openai.com/t/getting-400-response-with-already-working-code/509212>\
**Category:** API\
**Created:** [November 17, 2023, 7:52am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212 "2023-11-17T07:52:04Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![sp241930](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/sp241930/32/251707_2.png) [@sp241930](https://community.openai.com/u/sp241930)\
**Post date:** [November 17, 2023, 7:52am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/1 "2023-11-17T07:52:05Z")

</div>

we were using the same embedding API for a week it was working,no change in input,all of sudden we are getting 400 on the same API & input, we have also tried updating libraries

this is error

/server/node\_modules/openai/src/error.ts:66  
0|server | return new BadRequestError(status, error, message, headers);  
0|server | ^  
0|server | Error: 400 ‘$.input’ is invalid. Please check the API reference: [https://platform.openai.com/docs/api-reference](https://platform.openai.com/docs/api-reference).  
0|server | at Function.generate (/server/node\_modules/openai/src/error.ts:66:14)  
0|server | at OpenAI.makeStatusError (/server/node\_modules/openai/src/core.ts:358:21)  
0|server | at OpenAI.makeRequest (/server/node\_modules/openai/src/core.ts:416:24)  
0|server | at processTicksAndRejections (node:internal/process/task\_queues:95:5)  
0|server | server/node\_modules/langchain/dist/embeddings/openai.cjs:223:29  
0|server | at RetryOperation.\_fn (/server/node\_modules/p-retry/index.js:50:12)

---

<div class="post-metadata">

**Author:** ![rsanchezsilva](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/rsanchezsilva/32/255642_2.png) [@rsanchezsilva](https://community.openai.com/u/rsanchezsilva)\
**Post date:** [November 21, 2023, 6:21pm UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/2 "2023-11-21T18:21:59Z")

</div>

I was getting the same error with the Python client. You need to strip new lines from the beginning and end of the text in the documents. Hope it helps.

---

<div class="post-metadata">

**Author:** ![innova](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/innova/32/271807_2.png) [@innova](https://community.openai.com/u/innova)\
**Post date:** [December 11, 2023, 3:47pm UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/3 "2023-12-11T15:47:37Z")

</div>

I solved using this code in all the texts I passed:

```auto
import json

def sanitize_for_json(text):
    return json.dumps(text)

```

I am using chroma, this solution was provided via discord by the fantastic mod taz

---

<div class="post-metadata">

**Author:** ![innova](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/innova/32/271807_2.png) [@innova](https://community.openai.com/u/innova)\
**Post date:** [December 13, 2023, 3:17am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/4 "2023-12-13T03:17:22Z")

</div>

I realized that my error was for an none text, so you can also use this:

```auto
def sanitize(text):
    if not text:
        return " "
    return text

```

---

<div class="post-metadata">

**Author:** ![mittal.sameer](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mittal.sameer/32/147990_2.png) [@mittal.sameer](https://community.openai.com/u/mittal.sameer)\
**Post date:** [February 29, 2024, 12:34am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/5 "2024-02-29T00:34:26Z")

</div>

I am getting the same error in previously :

BadRequestError: Error code: 400 - {‘error’: {‘message’: “‘$.input’ is invalid. Please check the API reference: [https://platform.openai.com/docs/api-reference.](https://platform.openai.com/docs/api-reference.)”, ‘type’: ‘invalid\_request\_error’, ‘param’: None, ‘code’: None}}

I have tried approaches mentioned in the thread but they don’t work for me. I have openai-1.12.0 installed

Code:

embeddings\_list=

import openai

client = openai.OpenAI()

input\_list = [x.replace(“\n”, " ") for x in input\_list]

txt\_embeddings = client.embeddings.create(input = input\_list  
,model=“text-embedding-3-large”  
,dimensions=512).data

for x in txt\_embeddings:  
embeddings\_list.append(x.embedding)

Interestingly, the code works if I iterate through the list and pass one value at a time to: client.embeddings.create

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [February 29, 2024, 1:38am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/6 "2024-02-29T01:38:14Z")

</div>

There is no “input\_list” being defined in this snippet.

I gave you one, and improved

```auto
import openai
client = openai.Client()

with open("mytext.txt", "r", encoding="utf-8") as file:
    text_string = file.read() # read file, split into paragraph chunks, index
stripped_chunks = [part.strip() for part in text_string.split("\n\n") if part.strip()]
input_list = [f"[{index + 1}] {part}" for index, part in enumerate(stripped_chunks)]

input_list = [x.replace("\n", " ").replace(' ', ' ') for x in input_list]

txt_embeddings = client.embeddings.create(
    model="text-embedding-3-large",
    # input=["a good bot", "accepts lists"],
    input=input_list,
    dimensions=512,
    encoding_format="float",
)
embeddings_list = []
for x in txt_embeddings.data:
    embeddings_list.append(x.embedding)
    print(x.embedding[:4])
print(txt_embeddings.usage.model_dump())

```

counter to old advice above, also working:

```auto

    input=["\na good bot\n", "\n\n accepts lists\n\n"],

```

---

<div class="post-metadata">

**Author:** ![mittal.sameer](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mittal.sameer/32/147990_2.png) [@mittal.sameer](https://community.openai.com/u/mittal.sameer)\
**Post date:** [March 1, 2024, 12:37am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/7 "2024-03-01T00:37:12Z")

</div>

> [@\_j](#):
>
> `.strip()`

Hi, thanks for your answer.  
The input list I have is quite large, so cannot supply it here. It’s a list of strings that are input’s in a search box of an app from different users.

Not sure the fix you are suggesting above, but I tried all below but none works for me, same error:

1. input\_list = [x.replace(“\n”, " ") for x in input\_list]
2. input\_list = [x.replace(“\n”, " “).replace(” “,” ") for x in input\_list]
3. input\_list = [x.replace(“\n”, " “).replace(” “,” ").strip() for x in input\_list]
4. input\_list = [x.replace(“\n”, " “).replace(” “,” ").strip() for x in input\_list if x]

As mentioned earlier, code works if I iterate through input list and call the API for one string at a time

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [March 1, 2024, 1:20am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/8 "2024-03-01T01:20:44Z")

</div>

With the new embeddings model, there really isn’t a need to strip out linefeeds or separators. If an AI could understand the message, the embeddings AI can understand the message. The paragraphs and formatting might ensure even higher understanding.

Comment out the input\_list modification: does it work?

Don’t output to the same list variable as your input, does it work?

---

<div class="post-metadata">

**Author:** ![mittal.sameer](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mittal.sameer/32/147990_2.png) [@mittal.sameer](https://community.openai.com/u/mittal.sameer)\
**Post date:** [March 1, 2024, 1:34am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/9 "2024-03-01T01:34:51Z")

</div>

I commented out:  
input\_list = [x.replace(“\n”, " “).replace(” “,” ").strip() for x in input\_list if x]

and changed input\_list to input\_list\_1

input\_list\_1 = [x.replace(“\n”, " “).replace(” “,” ").strip() for x in input\_list if x]

Both approaches did not work, same error

There is some value in the list that’s erroring, but it should error even when I do it element by element 🤷‍♂️

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [March 1, 2024, 1:59am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/10 "2024-03-01T01:59:29Z")

</div>

Quite peculiar. Send two entities automated. Error?  
Put the text into the commented list strings I put into the code and comment out the other. Error?

Let’s imagine that you are doing chunking in the worst way possible, like splitting Unicode by byte lengths. And then your text is something like OpenAI documentation that has special tokens in the text itself. Let’s try to fix that:

```auto
import re # add to imports

# function: ensure valid characters in input_list
def ensure_utf8(strings):
    cleaned_strings = []
    for string in strings:
        # Decode using UTF-8, replacing invalid bytes
        clean_string = string.encode('utf-8', 'replace').decode('utf-8', 'replace')
        # Remove substrings enclosed in <| and |> - AI special tokens
        clean_string = re.sub(r'<\|.*?\|>', '', clean_string)
        # Optionally, take out ALL the extra space runs, like code indenting
        clean_string = ' '.join(clean_string.split())
        cleaned_strings.append(clean_string)
    return cleaned_strings

# add this where your input_list "enters" the code
input_list = ensure_utf8(input_list)

# The existing replace if you still want it
input_list = [x.replace("\n", " ").replace(' ', ' ') for x in input_list]

```

I also found that 2048 items is the max. 100000 tokens was accepted but I didn’t run it up higher, which might be a model limit of 128000. So you can also do some token-counting on the input.

---

<div class="post-metadata">

**Author:** ![mittal.sameer](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mittal.sameer/32/147990_2.png) [@mittal.sameer](https://community.openai.com/u/mittal.sameer)\
**Post date:** [March 1, 2024, 5:07am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/11 "2024-03-01T05:07:20Z")

</div>

> [@\_j](#):
>
> ```auto
> # add this where your input_list "enters" the code
> input_list = ensure_utf8(input_list)
> 
> ```

tried your function, same error. Strings in my list are max 3000 characters

2048 is the number of strings or characters per string?

Total tokens in my list are ~ 43,000

---

<div class="post-metadata">

**Author:** ![mittal.sameer](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mittal.sameer/32/147990_2.png) [@mittal.sameer](https://community.openai.com/u/mittal.sameer)\
**Post date:** [March 1, 2024, 6:36am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/13 "2024-03-01T06:36:09Z")

</div>

> # function: ensure valid characters in input\_list
> 
> def ensure\_utf8(strings):  
> cleaned\_strings =   
> for string in strings:  
> # Decode using UTF-8, replacing invalid bytes  
> clean\_string = string.encode(‘utf-8’, ‘replace’).decode(‘utf-8’, ‘replace’)  
> # Remove substrings enclosed in \<| and |\> - AI special tokens  
> clean\_string = re.sub(r’\<|.\*?|\>', ‘’, clean\_string)  
> # Optionally, take out ALL the extra space runs, like code indenting  
> clean\_string = ’ ‘.join(clean\_string.split())  
> cleaned\_strings.append(clean\_string)  
> cleaned\_strings = [x.replace(“\n”, " ").replace(’ ', ’ ') for x in cleaned\_strings]  
> return cleaned\_strings

input\_list = ensure\_utf8(input\_list)

Still same error

Max Tokens for a string in my input list: 812  
Total tokens across all strings: 202595

---

<div class="post-metadata">

**Author:** ![200106064](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/200106064/32/327542_2.png) [@200106064](https://community.openai.com/u/200106064)\
**Post date:** [March 7, 2024, 10:46am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/14 "2024-03-07T10:46:16Z")

</div>

Thanks its working but i cant understand how json.data() converts the itmes to string without a JSON file.

---

<div class="post-metadata">

**Author:** ![somil2760](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/somil2760/32/278028_2.png) [@somil2760](https://community.openai.com/u/somil2760)\
**Post date:** [April 25, 2024, 9:13pm UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/15 "2024-04-25T21:13:54Z")

</div>

Hi @mittal.sameer ,  
Were you able to rectify the error. I am also stuck in a similar situation, i.e. when I pass values one by one, it is working but when I am pass in batches then it gives the same error.

---

<div class="post-metadata">

**Author:** ![somil2760](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/somil2760/32/278028_2.png) [@somil2760](https://community.openai.com/u/somil2760)\
**Post date:** [April 25, 2024, 9:18pm UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/16 "2024-04-25T21:18:46Z")

</div>

I suspect is issue is total token count. As mentioned in the api document:

> input text to embed, encoded as a string or array of tokens. To embed multiple inputs in a single request, pass an array of strings or array of token arrays. The input must not exceed the max input tokens for the model (8192 tokens for `text-embedding-ada-002` ), cannot be an empty string, and any array must be 2048 dimensions or less.

[https://platform.openai.com/docs/api-reference/embeddings/create#:~:text=Input%20text%20to,dimensions%20or%20less](https://platform.openai.com/docs/api-reference/embeddings/create#:~:text=Input%20text%20to,dimensions%20or%20less).

From this understand is that when we pass a list then tokens are counted for the sum of tokens of each str in list. If such is the case then you can try batching.

---

<div class="post-metadata">

**Author:** ![mittal.sameer](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mittal.sameer/32/147990_2.png) [@mittal.sameer](https://community.openai.com/u/mittal.sameer)\
**Post date:** [April 26, 2024, 12:55am UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/17 "2024-04-26T00:55:51Z")

</div>

I solved it by batching as follows:

def divide\_list\_into\_batches(lst, batch\_size):  
for i in range(0, len(lst), batch\_size):  
yield lst[i:i+batch\_size]

## input\_list is your list of texts

for input\_list\_tmp in divide\_list\_into\_batches(input\_list, 2048):

---

<div class="post-metadata">

**Author:** ![DiogoR23](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/diogor23/32/437946_2.png) [@DiogoR23](https://community.openai.com/u/DiogoR23)\
**Post date:** [August 6, 2024, 2:22pm UTC](https://community.openai.com/t/getting-400-response-with-already-working-code/509212/19 "2024-08-06T14:22:19Z")

</div>

Hi everyone I am getting this error message:

```auto
Error initializing the system: Error code: 400 - {'error': "'input' field must be a string or an array of strings"}

```

It appears me when I run this function:

```python
def create_retriever_from_cassandra(session, name, description):
    """Create Cassandra Vector Store and transform it into a retriever.
        Return the retriever tool.
    """
    keyspace = KEYSPACE_ARTICLE
    table_name = "articles"

    embedding = OpenAIEmbeddings(
        api_key=OPENAI_API_KEY,
        base_url=BASE_URL,
        model="text-embedding-3-large"
    )

    CassVectorStore = Cassandra(
        session=session,
        keyspace=keyspace,
        table_name=table_name,
        embedding=embedding
    )

    retriever = CassVectorStore.as_retriever(
        search_type="similarity",
        search_kwargs={'k': 4}
    )

    retriever_tool = create_retriever_tool(
        retriever=retriever,
        name=name,
        description=description
    )

    return retriever_tool

```

I am knew in working with langchains and RAG’s, so it’s kinda hard for me to understand what is going on. I know for the fact, that my OpenAIEmbeddings needs to be a string or a an array of strings, like the error message says. How can I fix this.

Thanks everyone.
