# GPTBot makes over 10000 request to my website

**URL:** <https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489>\
**Category:** Bugs\
**Tags:** chatgpt\
**Created:** [April 23, 2025, 12:37am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489 "2025-04-23T00:37:45Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![TheresaQWQ](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/theresaqwq/32/618696_2.png) [@TheresaQWQ](https://community.openai.com/u/TheresaQWQ)\
**Post date:** [April 23, 2025, 12:37am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/1 "2025-04-23T00:37:45Z")

</div>

![The image shows a data analytics dashboard displaying HTTP request logs with a graphical representation and tabulated data details. (Captioned by AI)](https://us1.discourse-cdn.com/openai1/original/4X/a/9/d/a9da73c9dcaaa14a59b79c7b1127cfc984038969.png)

This endpoint is intended solely for file downloads and is backed by an S3 bucket, not a traditional webpage meant for indexing. However, we’ve observed that **GPTBot** has been sending an excessive number of requests to this endpoint, which has led to a **significant increase in our bandwidth usage and incurred substantial traffic costs**. These requests are unnecessary and harmful, as they access non-HTML content and place an unreasonable load on our infrastructure.

---

<div class="post-metadata">

**Author:** ![curt.kennedy](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/curt.kennedy/32/709249_2.png) [@curt.kennedy](https://community.openai.com/u/curt.kennedy)\
**Post date:** [April 23, 2025, 12:46am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/2 "2025-04-23T00:46:28Z")

</div>

Escalated to OpenAI.

(more words for the algorithm)

---

<div class="post-metadata">

**Author:** ![colinr](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/colinr/32/512000_2.png) [@colinr](https://community.openai.com/u/colinr)\
**Post date:** [April 23, 2025, 5:13am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/3 "2025-04-23T05:13:47Z")

</div>

GPTBot user agent information is published at [https://platform.openai.com/docs/bots](https://platform.openai.com/docs/bots), which you can use to restrict crawling in your robots.txt file. Alternatively you could email gptbot at [openai.com](http://openai.com) and let us know which domain you’re referencing.

---

<div class="post-metadata">

**Author:** ![dr.neil](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/dr.neil/32/621179_2.png) [@dr.neil](https://community.openai.com/u/dr.neil)\
**Post date:** [April 26, 2025, 11:20pm UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/4 "2025-04-26T23:20:00Z")

</div>

why is the crawler downloading binary files ?  
Is this expected ?

---

<div class="post-metadata">

**Author:** ![colinr](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/colinr/32/512000_2.png) [@colinr](https://community.openai.com/u/colinr)\
**Post date:** [April 27, 2025, 9:40pm UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/5 "2025-04-27T21:40:45Z")

</div>

Web crawlers generally just follow links and don’t know in advance what type of content is going to be served until they access it. If you’d like to disallow crawlers from visiting all or part of your site you can do so using the [robots.txt](https://developers.google.com/search/docs/crawling-indexing/robots/intro) file. OpenAI’s [crawler user agents are listed here](https://platform.openai.com/docs/bots/).

If you have additional questions or concerns about GPTBot please feel free to email us at gptbot (at) [openai.com](http://openai.com).

---

<div class="post-metadata">

**Author:** ![jake3](https://avatars.discourse-cdn.com/v4/letter/j/a698b9/32.png) [@jake3](https://community.openai.com/u/jake3)\
**Post date:** [January 6, 2026, 1:16am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/6 "2026-01-06T01:16:40Z")

</div>

The user agent list link is broken.

GPTBot is hammering our staging site with a request every 3 seconds. It’s found a calender page and keeps making requests with different months and years in the URL parameters, some 800 years in the future. The staging site is not linked from anywhere and only discoverable from the DNS settings. Altogether the antithesis of intelligence.

I’ve disallowed it in `robots.txt` but the requests keep coming.

This is a denial-of-service attack and should be shut down, using appropriate law enforcement as necessary.

---

<div class="post-metadata">

**Author:** ![colinr](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/colinr/32/512000_2.png) [@colinr](https://community.openai.com/u/colinr)\
**Post date:** [January 6, 2026, 6:09am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/7 "2026-01-06T06:09:47Z")

</div>

Aside from simple misconfigurations, a common mistake is when sites (or their hosting providers) inadvertently fail to actually serve their robots.txt file to web crawlers. Please share more information, such as a domain name that you’re referring to, to help understand your issue. As mentioned above, you can email gptbot (at) [openai.com](http://openai.com/) if you prefer not to share more info here.

---

<div class="post-metadata">

**Author:** ![jake3](https://avatars.discourse-cdn.com/v4/letter/j/a698b9/32.png) [@jake3](https://community.openai.com/u/jake3)\
**Post date:** [January 7, 2026, 12:40am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/8 "2026-01-07T00:40:53Z")

</div>

I’ve sent an email with server logs and domain info. I am quite sure the problem is your end not mine. Please fix it.

---

<div class="post-metadata">

**Author:** ![colinr](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/colinr/32/512000_2.png) [@colinr](https://community.openai.com/u/colinr)\
**Post date:** [January 7, 2026, 12:58am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/9 "2026-01-07T00:58:26Z")

</div>

Thanks Jake, confirmed GPTBot is no longer crawling the site now that it’s serving us content for robots.txt. I’ve replied to your email with more detail.

---

<div class="post-metadata">

**Author:** ![Foxalabs](https://sea2.discourse-cdn.com/openai1/user_avatar/community.openai.com/foxalabs/32/738355_2.png) [@Foxalabs](https://community.openai.com/u/Foxalabs)\
**Post date:** [January 7, 2026, 9:35am UTC](https://community.openai.com/t/gptbot-makes-over-10000-request-to-my-website/1238489/10 "2026-01-07T09:35:37Z")

</div>

I just want to add to this thread to say that 10k requests over a few hours to a day is a very normal level of activity for a well configured web crawler, GoogleBot does this kind of number on many of my sites and it’s a good thing, getting indexed for searching means your content is discoverable.

If this is your first website and you are alarmed at the number of connections in a day, these kinds of numbers are quite typical, a denial of service attack is many thousands in a matter of seconds, but even this is normal for a busy website.

My recommendation is to let the crawlers do their thing if you wish to be found when people ask AI and search engines for answers.
