# I built an LLM-powered tool that can comprehend any website structure and extract the desired data in the preferred format

**URL:** <https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987>\
**Category:** Community\
**Created:** [January 19, 2023, 4:58pm UTC](https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987 "2023-01-19T16:58:40Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![mangotree](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mangotree/32/20140_2.png) [@mangotree](https://community.openai.com/u/mangotree)\
**Post date:** [January 19, 2023, 4:58pm UTC](https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987/1 "2023-01-19T16:58:40Z")

</div>

I got frustrated with the time and effort required to code and maintain custom web scrapers, so I built a more generic ML-based solution for data extraction from unstructured websites (and potentially other sources).

One of the killer use cases of GPT is reformatting information from any format X to any other format Y, so I leveraged that to understand websites and extract any data in the preferred format:

[Landing Page and Demo](https://kadoa.co)

We’re currently working on fine-tuning the platform and would love to have some early adopters test it out and provide feedback. Would love to hear your thoughts!

[![](https://us1.discourse-cdn.com/openai1/original/3X/7/5/75f134411112324146b502212fa1e4365f126510.jpeg "how kadoa works") ](https://www.youtube.com/watch?v=aa-MP-3J15U)

---

<div class="post-metadata">

**Author:** ![ashish\_vohra](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/ashish_vohra/32/22336_2.png) [@ashish\_vohra](https://community.openai.com/u/ashish_vohra)\
**Post date:** [January 20, 2023, 4:15am UTC](https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987/2 "2023-01-20T04:15:15Z")

</div>

Great idea! My company has been looking for this type of resource for years. I used your early access form and filled out a request. Looking forward to seeing the results you send! Thanks.

---

<div class="post-metadata">

**Author:** ![mangotree](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/mangotree/32/20140_2.png) [@mangotree](https://community.openai.com/u/mangotree)\
**Post date:** [January 24, 2023, 5:11pm UTC](https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987/4 "2023-01-24T17:11:04Z")

</div>

Thanks for the feedback. We’re onboarding users on a rolling basis while still fine-tuning and testing the platform for different use cases. You’ll hear back from us soon.

---

<div class="post-metadata">

**Author:** ![info21](https://avatars.discourse-cdn.com/v4/letter/i/6de8d8/32.png) [@info21](https://community.openai.com/u/info21)\
**Post date:** [January 27, 2023, 7:33pm UTC](https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987/5 "2023-01-27T19:33:32Z")

</div>

Interesting! How do you make sure GPT-3 (or any other LLM) is reliably returning the same output data and not making things up e.g. if a field doesn’t exist or isn’t visible on the website?

---

<div class="post-metadata">

**Author:** ![psm](https://avatars.discourse-cdn.com/v4/letter/p/a9adbd/32.png) [@psm](https://community.openai.com/u/psm)\
**Post date:** [June 21, 2023, 7:04pm UTC](https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987/6 "2023-06-21T19:04:17Z")

</div>

Interesting Idea.  
Is it open source? If not, will it be open source?

---

<div class="post-metadata">

**Author:** ![info21](https://avatars.discourse-cdn.com/v4/letter/i/6de8d8/32.png) [@info21](https://community.openai.com/u/info21)\
**Post date:** [November 4, 2023, 6:19am UTC](https://community.openai.com/t/i-built-an-llm-powered-tool-that-can-comprehend-any-website-structure-and-extract-the-desired-data-in-the-preferred-format/40987/7 "2023-11-04T06:19:08Z")

</div>

Open-sourcing parts of the tool like the CSS selector generator would be very cool.
