OpenAI launches webcrawler GPTBot, and instructions on how to block it

August 8, 2023

6 Views

SaveSavedRemoved 0

OpenAI launches webcrawler GPTBot, and instructions on how to block it

OpenAI has launched a web crawler to improve artificial intelligence models like GPT-4.

Called GPTBot, the system combs through the Internet to train and enhance AI’s capabilities. Using GPTBot has the potential to improve existing AI models when it comes to aspects like accuracy and safety, according to a blog post by OpenAI.

“Web pages crawled with the GPTBot user agent may potentially be used to improve future models and are filtered to remove sources that require paywall access, are known to gather personally identifiable information (PII), or have text that violates our policies,” reads the post.

Websites can choose to restrict access to the web crawler, however, and prevent GPTBot from accessing their sites, either partially or by opting out entirely. OpenAI said that website operators can disallow the crawler by blocking its IP address or on a site’s Robots.txt file.

How to prevent GPTBot from using your website’s content

According to OpenAI, you can disallow GPTBot by adding it to your site’s Robots.txt, which is essentially a text file that instructs web crawlers on what they can or cannot access from a website.