If you manage a website, you may have seen two OpenAI crawler names in your server logs: GPTBot and OAI-SearchBot. They are not two names for the same job. GPTBot is associated with content that may be used to improve OpenAI's foundation models. OAI-SearchBot is used to find and retrieve pages for ChatGPT search results.
That difference gives a publisher a useful choice. You can allow one crawler and limit the other. You can also allow both, block both, or set a policy by path. The right choice depends on your rights, audience, and publishing goals. There is no setting that guarantees a ranking, a citation, or a visit.
This guide explains the distinction in plain language. It also shows how to check the policy a crawler sees. For a live test on your own site, use the AI Crawler Access Checker. It reports the public rules and page directives it can observe without claiming that a crawler will use your page.
Start With The Job You Need#
Before editing robots.txt, decide what outcome you want.
If you want your pages to be considered for ChatGPT search answers, the relevant OpenAI crawler is OAI-SearchBot. OpenAI's publisher and developer FAQ says that allowing OAI-SearchBot lets a site appear in ChatGPT search results, subject to the usual limits of search and retrieval. That is an opportunity for discovery, not a promise of placement or traffic.
If you do not want your public pages used for training, GPTBot is the name to review. The same FAQ explains that publishers can disallow GPTBot while still allowing OAI-SearchBot. This separates a training preference from a search preference. It is a better starting point than blocking every automated request because the requests look unfamiliar.
The policy can be different for different folders. A public help center may be useful for search. A private preview, customer export, or draft archive may need a stronger block. Use the narrowest path that matches the right you control.
How The Two OpenAI Crawlers Differ#
The names describe different functions:
| Crawler | Main purpose | Policy question |
|---|---|---|
| GPTBot | May collect web content for model improvement | Do we allow this use of our content? |
| OAI-SearchBot | Retrieves pages for ChatGPT search | Do we want this search discovery path? |
The word “may” matters. A robots rule is a request about access. It is not a contract that describes every later use, and it is not a way to verify what a private system did with a copy. Read the current provider documentation and your own terms before making a rights decision.
OpenAI also documents other user agents and says publishers can use separate rules for them. A policy written for GPTBot should not be assumed to cover every OpenAI request. Check the exact user-agent name in the current OpenAI crawler guidance, then review your logs and hosting controls.
This pattern is not unique to OpenAI. Anthropic documents separate agents for training, search, and user-directed requests in its crawler article. A clear policy names the purpose you are deciding about instead of treating every bot as one category.
What Robots Txt Can Control#
robots.txt is a public instruction file at the root of a host, such as https://example.com/robots.txt. It can describe which user agents may fetch which paths. A common starting point looks like this:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
This example blocks GPTBot and permits OAI-SearchBot. It is only an example, not a recommendation for every site. If you publish content under a subpath, replace / with the path you mean to control. If you use a CDN, framework, or managed host, confirm which file reaches the public web.
Robots rules do not replace access control. They do not protect a private document, erase a page that is already copied, or change a provider's terms. They also do not automatically control an API, an image host, a second domain, or a staging host. Keep private material behind real authentication and network controls.
The file can also be syntactically valid while the result is wrong. A later group can override the rule you intended, a generated file can omit a user-agent block, or a proxy can serve a different file by region. Test the public response instead of trusting the file in your repository.
Choose Access By Purpose#
There are four reasonable policy patterns:
- Search and training allowed. Use this when you want broad public discovery and your content rights permit model improvement.
- Search allowed, training limited. Use separate rules when ChatGPT search is useful but GPTBot access does not match your publishing choice.
- Both limited. Use this for drafts, private material, or a site whose owner is still reviewing the rights and business case.
- Path-specific access. Allow a public documentation folder while limiting customer data, previews, or internal notes.
Do not choose a pattern because a blog promises more “AI visibility.” The practical question is whether a reader benefits from the page being found and whether your team accepts the stated use. Record that decision beside the file so it survives a future redesign.
Test The Rules Before Publishing#
Run a check from outside your development environment. The AI Crawler Access Checker can fetch the public robots.txt, compare documented user-agent groups, inspect page-level robots tags, and show the final response it received. Save the result with the date and the URL tested.
Then check the page itself. A page can allow a crawler in robots.txt and still send a noindex or nofollow directive in HTML or an X-Robots-Tag header. Those directives have different jobs. noindex asks search systems not to include a page. nofollow qualifies links on the page. Neither one is a substitute for deciding whether a crawler may fetch the file.
Test at least one real content URL, one blocked URL, and the root robots.txt. Check the final status, content type, redirects, and the response from your CDN. Repeat after a deployment because a framework route can silently replace a hand-written file.
Where Llms Txt Fits#
An llms.txt file can give language models a curated map of important pages. It is a useful publishing aid, but it is not a replacement for robots.txt, an access-control system, or a guarantee that a model will read or cite the file.
Use the llms.txt Generator and Validator to create a readable list of canonical pages, then review each link by hand. Include the pages that explain your product, method, limits, and sources. Do not use the file to hide a policy decision that belongs in robots.txt.
The same principle applies to a sitemap. A sitemap helps systems discover URLs; it does not grant access or make a thin page valuable. The Sitemap Health Checker can help you test the XML response, canonical URLs, and basic discovery signals together.
Common Crawler Policy Mistakes#
The most common errors are easy to prevent:
- Blocking
GPTBotand assuming every OpenAI crawler is blocked. - Allowing
OAI-SearchBotand treating a search visit as guaranteed. - Editing a local file while a CDN serves an older public copy.
- Putting private data behind a robots rule instead of authentication.
- Adding an
llms.txtfile and assuming it controls crawling. - Using a broad
Disallow: /rule when only one folder needed review. - Forgetting that a second host or asset domain has its own policy.
- Calling a page “indexable” without checking its canonical and page-level directives.
The fix is not more bot names. It is a short, documented policy and a repeatable public test.
Use A Simple Review Checklist#
Before you merge a crawler-policy change, ask:
- Which user agent and purpose are we deciding about?
- Which exact paths should be public, limited, or private?
- Does the public
robots.txtmatch the intended file? - Does a real page return the expected status and content type?
- Are HTML and header directives consistent with the policy?
- Are sitemap and
llms.txtlinks canonical and useful? - Did the CDN, staging host, and asset host receive the same review?
- Where is the decision recorded, and when will we recheck it?
Keep the answers short. A future editor should be able to understand the choice without reverse-engineering a deployment.
Keep The Policy Honest#
Crawler access is one part of a site's publishing system. Clear titles, useful pages, reliable links, and a stable canonical URL still matter. A bot can fetch a page and decide that it is not useful. A search system can retrieve a page and choose another source. That is normal.
Mydentify's directory methodology uses the same evidence-first rule for product listings: separate what a page says from what an independent check confirms. Apply that rule to crawler policies too. Record the file you served, the page you tested, the date, and any provider documentation that informed the choice.
The goal is not to allow every crawler. The goal is to make an intentional choice that respects your rights and helps the right readers find the right pages. Recheck the policy after a redesign, domain move, CDN change, or provider update.
Read The Official Sources#
Start with OpenAI's publisher and developer FAQ for current user-agent definitions and controls. For a second provider's model, search, and user-requested agents, read Anthropic's crawler documentation. Provider names and policies can change, so keep the review date beside your decision.
If you want to inspect your own public behavior next, run the AI Crawler Access Checker, then use the canonical URL identity checker on one representative page. The result should make your policy easier to explain—not promise a ranking or a citation.
See How We Checked It
We name the source, record the review date, and separate reported claims from our own checks. Read the methodology and editorial policy before citing a finding or sending a correction.
Put this guide into practice
Open the ai discovery tools hub, then choose the check that matches your next decision.
Found an outdated fact? Report a correction.
Back to articles