# Understanding the Chunking Process in a Vector Store for Plain Text Files

**URL:** <https://community.openai.com/t/understanding-the-chunking-process-in-a-vector-store-for-plain-text-files/850242>\
**Category:** API\
**Created:** [July 2, 2024, 7:39am UTC](https://community.openai.com/t/understanding-the-chunking-process-in-a-vector-store-for-plain-text-files/850242 "2024-07-02T07:39:19Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![medi.mhb.w](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/medi.mhb.w/32/448307_2.png) [@medi.mhb.w](https://community.openai.com/u/medi.mhb.w)\
**Post date:** [July 2, 2024, 7:39am UTC](https://community.openai.com/t/understanding-the-chunking-process-in-a-vector-store-for-plain-text-files/850242/1 "2024-07-02T07:39:19Z")

</div>

I have a vector store with many plain text files of different sizes. I would like to understand **how these text files are divided into smaller chunks for storage in the database**. Specifically, I am interested in knowing **how the system ensures that the context and continuity of the original text are preserved when large text files are split into smaller chunks**. Thank you.

---

<div class="post-metadata">

**Author:** ![sergeliatko](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/sergeliatko/32/816_2.png) [@sergeliatko](https://community.openai.com/u/sergeliatko)\
**Post date:** [July 2, 2024, 7:44am UTC](https://community.openai.com/t/understanding-the-chunking-process-in-a-vector-store-for-plain-text-files/850242/2 "2024-07-02T07:44:48Z")

</div>

[Using gpt-4 API to Semantically Chunk Documents](https://community.openai.com/t/using-gpt-4-api-to-semantically-chunk-documents/715689) would be a good starting point

---

<div class="post-metadata">

**Author:** ![\_j](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/_j/32/766292_2.png) [@\_j](https://community.openai.com/u/_j)\
**Post date:** [July 2, 2024, 8:23am UTC](https://community.openai.com/t/understanding-the-chunking-process-in-a-vector-store-for-plain-text-files/850242/3 "2024-07-02T08:23:56Z")

</div>

## Assistants

TL;DR: There is no context understanding for split points. No “ensuring”.

The extracted file is split at 800 token intervals (approximately), then including an overlap into adjacent sections, 400 tokens.

These parameters can now be altered when adding a file to a vector store.

## Your vector database solution

(You can probably do better.)
