# All languages are NOT created (tokenized) equal

**URL:** <https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407>\
**Category:** Community\
**Tags:** token, app, comparison, statistics\
**Created:** [May 18, 2023, 9:43am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407 "2023-05-18T09:43:07Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![EricGT](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ericgt/32/20571_2.png) [@EricGT](https://community.openai.com/u/EricGT)\
**Post date:** [May 18, 2023, 9:43am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/1 "2023-05-18T09:43:07Z")

</div>

FYI

For those with in interest in using LLMs and understanding the tokens with regards to different spoken languages the following blog is of interest.

[All languages are NOT created (tokenized) equal](https://blog.yenniejun.com/p/all-languages-are-not-created-tokenized)

Also see the related online app.

> The purpose of this project is to compare the tokenization length for different languages. For some tokenizers, tokenizing a message in one language may result in 10-20x more tokens than a comparable message in another language (e.g. try English vs. Burmese). This is part of a larger project of measuring inequality in NLP.

> **[Tokenizers Languages - a Hugging Face Space by yenniejun](https://huggingface.co/spaces/yenniejun/tokenizers-languages)**
>
> Discover amazing ML apps made by the community

* * *

Note:

The [OpenAI tokenizer](https://platform.openai.com/tokenizer) only demonstrates the following tokenizers

- GPT-3
- Codex

The app from the blog states it is using OpenAI GPT-4 tokenizer which is available from the GitHub repository for [tiktoken](https://github.com/openai/tiktoken).

---

<div class="post-metadata">

**Author:** ![N2U](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/n2u/32/54438_2.png) [@N2U](https://community.openai.com/u/N2U)\
**Post date:** [May 18, 2023, 10:05am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/2 "2023-05-18T10:05:03Z")

</div>

Interesting!

I’m wondering how much this has to do with how the words translate to English, I’m thinking that the closer you get to English the fewer tokens you have to use, but I’m also curious how this is affected by words with multiple meanings.

Here’s a few examples of Danish words that require extra context/tokens to be properly understood by an LLM:

 ![IMG_20230226_130653](https://us1.discourse-cdn.com/openai1/original/3X/6/b/6b69fdb4335b5b373d0db88d14c75dbb1574d528.jpeg)  
 ![IMG_20230226_124634](https://us1.discourse-cdn.com/openai1/original/3X/9/e/9eb913183cdd047dfa56491a347bf194385651e8.jpeg)

---

<div class="post-metadata">

**Author:** ![EricGT](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ericgt/32/20571_2.png) [@EricGT](https://community.openai.com/u/EricGT)\
**Post date:** [May 18, 2023, 10:08am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/3 "2023-05-18T10:08:45Z")

</div>

Did you see that with the app you can

1. Select different tokenizers

![image](https://us1.discourse-cdn.com/openai1/original/3X/3/3/3300194f09d9b587686ed4c0b37fe16eead64f36.png)

1. Select different languages for comparison

![image](https://us1.discourse-cdn.com/openai1/original/3X/1/1/115c3d1c30c4afd77624529e489e3112efdcc5c6.png)

One downside of the app is that I don’t see a way change the data source. There is an option to randomly sample the data source.

![image](https://us1.discourse-cdn.com/openai1/original/3X/f/2/f2bd517302b7986600fe68db84d956739f806a69.png)

---

<div class="post-metadata">

**Author:** ![N2U](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/n2u/32/54438_2.png) [@N2U](https://community.openai.com/u/N2U)\
**Post date:** [May 18, 2023, 10:34am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/4 "2023-05-18T10:34:24Z")

</div>

No I didn’t see that, but will definitely play with it later ❤

I think it could be interesting to figure out what language results in the least amount of tokens used, although I’m expecting the answer to be English.

---

<div class="post-metadata">

**Author:** ![EricGT](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ericgt/32/20571_2.png) [@EricGT](https://community.openai.com/u/EricGT)\
**Post date:** [May 18, 2023, 10:40am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/5 "2023-05-18T10:40:22Z")

</div>

> [@N2U](#):
>
> I’m expecting the answer to be English

Not so fast.

For most general usage tokenizers where the training data is mostly English I would agree but I suspect either in private, research, non English countries and such that tokenizers might be crafted for a language other than English and give better results.

In the back of my mind I am also asking what a tokenizer for Math would do, how would the vectors work and can such vectors be incorporated into a general LLM based on say English.

---

<div class="post-metadata">

**Author:** ![N2U](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/n2u/32/54438_2.png) [@N2U](https://community.openai.com/u/N2U)\
**Post date:** [May 18, 2023, 11:11am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/6 "2023-05-18T11:11:13Z")

</div>

Agreed,

I’m curious whether the cooperation between the Icelandic government and OpenAI will cause the amount of tokens used to tokenize Icelandic language to go down over time.

The sentence: `The quick brown fox jumps over the lazy dog` tokenized in English is 9 tokens. The Icelandic sentence (`Hinn fljóti brúni refur hoppar yfir lata hundinn`) is currently 22.

This was done using the tokenizer on the OpenAI site (GPT-3).

---

<div class="post-metadata">

**Author:** ![humbroll](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/humbroll/32/177566_2.png) [@humbroll](https://community.openai.com/u/humbroll)\
**Post date:** [August 13, 2023, 5:09am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/7 "2023-08-13T05:09:05Z")

</div>

I’m interested in this topic. Have there been any updates since the last comment? I’m not sure we should charge differently for our services depending on the users’ country.

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [August 13, 2023, 9:22am UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/8 "2023-08-13T09:22:43Z")

</div>

If you have a service that is _unaware of the actual token consumption_ of users and doesn’t bill accordingly, along with billing the value that your service actually adds, then it will indeed be a better value than were the user to use an API key to converse with the AI directly for under-represented languages. Some languages will have significant amplification of the number of tokens per character or per semantically-similar language, Chinese being one of the largest disparities at 2-3 tokens per character.

Korean has a robust dictionary, where we have 1 hangul combination unicode per token, making most words 1-3 tokens. Other languages are varied.

If you are not specializing with your chat application, you are competing with ChatGPT, where the output language and its length is not discriminated against.

---

<div class="post-metadata">

**Author:** ![N2U](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/n2u/32/54438_2.png) [@N2U](https://community.openai.com/u/N2U)\
**Post date:** [August 13, 2023, 5:05pm UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/9 "2023-08-13T17:05:10Z")

</div>

Very interesting observation here:

> [@\_j](#):
>
> Korean has a robust dictionary, where we have 1 hangul combination unicode per token, making most words 1-3 tokens. Other languages are varied.

Seems like money could be saved, I’ll call this “the Korean discount” from now on. 😆

If anyone knows a language that tokenizes to even fever tokens, please tell us.

---

<div class="post-metadata">

**Author:** ![EricGT](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ericgt/32/20571_2.png) [@EricGT](https://community.openai.com/u/EricGT)\
**Post date:** [December 17, 2023, 7:11pm UTC](https://community.openai.com/t/all-languages-are-not-created-tokenized-equal/216407/10 "2023-12-17T19:11:41Z")

</div>


