The short answer
The four major AI engines do not share a crawling architecture, and that directly changes what is visible on your site. OpenAI runs four separate bots (including GPTBot for training and OAI-SearchBot for search) and is gradually building its own index while still leaning on Bing ([1]). Anthropic takes the lightest approach: Claude outsources retrieval entirely to the Brave Search API ([2]). Perplexity has built the most independent infrastructure, an index of over 200 billion URLs ([3]), but also faces the most serious suspicion of bypassing robots.txt ([4]). Google has no separate AI crawler at all: Googlebot does everything, with a dedicated token, Google-Extended, to control training use ([5]). And the most consequential technical fact: only Googlebot executes JavaScript. GPTBot, ClaudeBot, and PerplexityBot do not ([6]).
OpenAI: four bots, and a homegrown index still under construction
OpenAI runs the most complex ecosystem, with four separate bots, each independently controllable via robots.txt ([7]):
| Bot | Role |
|---|---|
| GPTBot | Collects training data for generative models |
| OAI-SearchBot | Indexes content for ChatGPT Search results |
| ChatGPT-User | Fetches a specific page in real time when a user requests it |
| OAI-AdsBot | Other product actions |
Notably, since the GPT-5 launch in August 2025, OAI-SearchBot now generates roughly 3.5x more log events than GPTBot ([8]), a sign that search now outweighs training collection in OpenAI's crawl activity. GPTBot alone generated approximately 569 million requests across Vercel's network in December 2024 ([9]).
ChatGPT Search remains a hybrid architecture: queries from Enterprise and Edu workspaces still route through Bing ([10]), but OpenAI is building its own index in parallel, with separate verticals for news, shopping, PDFs, and academic papers, identified by Peec AI through experiment names such as prefer-index-over-serp-v3 ([1]). The retrieval stack reportedly runs on three layers: a discovery index, a reading cache that keeps copies of previously fetched pages, and a small set of pages opened live at query time ([11]).
Anthropic: the lightest approach, the most dependent on Brave
Anthropic runs three bots: ClaudeBot for training, Claude-User for real-time retrieval, and Claude-SearchBot for search index optimization, each independently controllable via robots.txt ([12]).
Claude's defining trait is that it has no search index of its own. When web search is enabled, Claude relies entirely on the Brave Search API, confirmed in March 2025 when Anthropic added "Brave Search" to its subprocessor list ([2]). Claude's web fetch tool then retrieves full page content on demand, with HTML preprocessing that strips scripts and styles ([13]).
It is also the engine with the strongest citation precision, thanks to its Citations API, which chunks source documents into sentences before passing them to the model, enabling exact-sentence-level attribution ([14]). Anthropic reports this feature improves citation recall by up to 15% compared to a hand-built, prompt-based citation implementation ([14]).
ClaudeBot also had its own notoriety moment: in July 2024, iFixit reported roughly one million requests in 24 hours, a crawl rate fast enough to trigger internal alerts ([15]).
Perplexity: the most independent infrastructure, the most contested
Perplexity has invested most aggressively in proprietary search infrastructure: an index of over 200 billion unique URLs, described as an "exabyte-scale index," backed by tens of thousands of CPUs and hundreds of terabytes of RAM ([3]). The company handles approximately 200 million queries per day ([3]) and was already serving roughly 30 million cited answers per day by May 2025 ([16]).
Two bots are declared: PerplexityBot for indexing and Perplexity-User for user-initiated fetches ([17]). The problem is what is not declared. Cloudflare documented "stealth crawling" behavior: user agents modified mid-crawl, and a switch to unannounced IP ranges, which led Cloudflare to de-list Perplexity as a verified bot ([4]). Multiple publishers have taken the matter to court, including the New York Times, Reddit, Dow Jones, and the BBC ([18]).
Google: one bot for everything, one token for AI training
Google takes the opposite approach from the other three: no separate AI crawler at all. Googlebot fetches everything, and Google controls training use through a dedicated robots.txt token, Google-Extended, introduced in September 2023 ([5]). Blocking Google-Extended prevents pages from being used for model training, but does not affect regular search indexing.
A newer fetcher, Google-Agent, handles real-time, user-triggered requests inside Gemini products, and behaves more like a standard web browser than a search crawler ([19]). Gemini relies on a six-stage grounding pipeline that rewrites the query, reranks results, and blends them into the model's context before generating a response ([20]).
The JavaScript gap: the technical fact that decides your visibility
This is the most consequential limitation across the entire ecosystem. GPTBot, ClaudeBot, and PerplexityBot do not execute JavaScript, a finding established by Vercel's analysis of over half a billion AI crawler requests ([6]). ChatGPT-User and ClaudeBot sometimes fetch JavaScript files (11.5% and 23.8% of requests respectively) but never execute them ([6]). Googlebot is the sole exception: it renders JavaScript through a headless Chrome-based pipeline, making it the only AI-adjacent crawler that can see fully rendered single-page application (SPA, an app whose content is built in the browser rather than delivered as complete HTML) content ([21]).
In practice, this means:
- Any content injected client-side, including JSON-LD structured data added through Google Tag Manager, is invisible to GPTBot, ClaudeBot, and PerplexityBot ([22]).
- On an SPA-based site, these crawlers only see the initial HTML shell: navigation skeletons, meta tags, empty placeholder containers ([6]).
- Content loaded via AJAX, infinite scroll, or anything dependent on client-side rendering stays inaccessible ([22]).
- Paywalled content is equally invisible: AI crawlers do not authenticate, they only see whatever the unauthenticated HTML response returns ([6]).
Server-side rendering (SSR, generating the full HTML on the server before sending it to the browser, rather than in the browser) or static HTML are therefore essential for AI engine visibility ([23]).
Who blocks what: blocking rates by crawler
Nearly 4 in 10 top websites now actively block at least one AI crawler, according to Zyte's State of Web Access 2026 analysis of 11,100 top landing pages ([24]):
| Crawler | Explicit block rate |
|---|---|
| GPTBot | 8.4% |
| CCBot | 7.3% |
| ClaudeBot | 6.3% |
| Google-Extended | 5.6% |
| Bytespider | 5.5% |
| PerplexityBot | 4.2% |
| OAI-SearchBot | 2.9% |
(Source: [25])
The ranking is not random. OAI-SearchBot is the least blocked, likely because it sends referral traffic back to publishers, while GPTBot, purely dedicated to training with no traffic in return, is the most blocked ([24]). Among news publishers specifically, blocking rates run far higher: 79% block AI training bots and 71% also block AI retrieval bots ([26]). Cloudflare reports that over 2.5 million websites have chosen to completely disallow AI training through its managed robots.txt feature ([27]).
Volume gaps are just as stark. According to Vercel's network data from late 2024, Googlebot generated 4.5 billion requests per month, against 569 million for GPTBot, 370 million for ClaudeBot, and just 24.4 million for PerplexityBot ([9]). By early 2026, GPTBot and ClaudeBot accounted for 12% and 9.2% of global bot traffic respectively according to Cloudflare Radar data ([28]), and GPTBot's request volume grew 305% between May 2024 and May 2025 ([29]).
A side-by-side comparison of the four engines
| Capability | Perplexity | ChatGPT | Claude | Gemini |
|---|---|---|---|---|
| Own search index | Full | Moderate | None (Brave) | Full (Google) |
| JavaScript rendering | None | None | None | Full |
| Citation accuracy | Moderate | Low | High | Moderate |
| Agentic search / Deep Research | High | Full | Low | Full |
| robots.txt compliance | Minimal | High | High | High |
(Synthesis of the qualitative positioning found across the studies cited above, notably [6] and [4].)
Gemini comes out ahead thanks to Google's index and JavaScript rendering, one of the two advantages no competitor shares. Claude offsets having no index of its own with the strongest citation precision, driven by its sentence-level Citations API. Perplexity matches the others on index scale but pays for it with the weakest robots.txt compliance.
The citation reliability problem, shared by all four
Despite these different architectures, no engine has solved the citation reliability problem. A 2025 study by the Tow Center for Digital Journalism at Columbia University, run across 1,600 test queries on eight generative search engines, found that these engines fail to retrieve correct information more than 60% of the time ([30]). A separate 2025 arXiv study found only 26.5% of bibliographic references were fully correct, versus 39.8% erroneous or fabricated ([31]). And source overlap between AI engines citing the same topic can run as low as 16% ([32]): for the exact same question, two engines can rely on evidence with almost nothing in common.
Over the same period, AI search visit volume jumped from 15.6 to 27.4 billion, a 42.8% rise between Q1 2025 and Q1 2026 ([33]). Audience is growing faster than answer reliability is improving.
Timeline: the dates that matter
- September 2023: Google introduces the Google-Extended robots.txt token to separate training from indexing ([5]).
- June 2024: WIRED and developer Robb Knight independently document Perplexity ignoring robots.txt and using spoofed user agents.
- July 2024: iFixit reports roughly one million ClaudeBot requests in 24 hours ([15]).
- March 2025: Anthropic officially confirms it relies on the Brave Search API for Claude's web search ([2]).
- Summer 2025: Cloudflare documents Perplexity's stealth crawling and de-lists it as a verified bot ([4]).
- August 2025: GPT-5 launches; OpenAI's crawl activity triples in the following weeks ([8]).
- October to December 2025: Reddit, then the New York Times, sue Perplexity for copyright infringement ([18]).
- March 2026: Google rolls out Google-Agent, a fetcher separate from Googlebot for real-time Gemini use cases ([19]).
FAQ
Can ChatGPT, Claude, and Perplexity execute JavaScript? No. GPTBot, ClaudeBot, and PerplexityBot do not render JavaScript; they only see the raw HTML returned by the server. Only Googlebot, through its headless Chrome rendering pipeline, executes JavaScript ([6], [21]).
Why doesn't Claude have its own search index? Anthropic chose not to build its own search infrastructure and instead delegates result retrieval entirely to the Brave Search API, a choice publicly confirmed in March 2025 ([2]).
Does robots.txt actually block every AI crawler? Not uniformly. OpenAI, Anthropic, and Google respect robots.txt directives per bot token. Perplexity, however, has been documented bypassing those directives using spoofed user agents and undeclared IP ranges ([4]).
What is Google-Extended, and how does it differ from Googlebot? Google-Extended is a separate robots.txt token that controls only whether pages can be used to train Google's AI models. Blocking it does not affect regular search indexing, which is still handled by Googlebot ([5]).
Why are these engines' citations often wrong? Studies show retrieval failure rates above 60% on test queries, with links sometimes fabricated and source overlap between engines as low as 16% for the exact same question ([30], [32]).
Does Perplexity respect robots.txt? The evidence documented by Cloudflare and independent developers indicates it does not, at least not consistently: user agents modified mid-crawl, undeclared IP ranges, which led Cloudflare to de-list it from its verified bots list ([4]).
Which crawler generates the most traffic on a given site? It varies by industry, but at the global scale, Googlebot dominates by far (4.5 billion monthly requests per Vercel), followed by GPTBot then ClaudeBot, with PerplexityBot remaining the smallest volume among the four main AI crawlers ([9]).
Your action plan
- Serve server-rendered or static HTML for any critical content. GPTBot, ClaudeBot, and PerplexityBot will never see what only exists in a client-side generated DOM.
- Never inject your structured JSON-LD data through a client-side tag manager. It needs to exist in the raw HTML the server returns.
- Check your robots.txt token by token, not just
User-agent: *. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot, and Claude-SearchBot are each controlled independently. - Separate training bots from search bots in your server logs. Blocking a training bot (GPTBot, ClaudeBot) does not have the same visibility impact as blocking a search bot (OAI-SearchBot, Claude-SearchBot).
- Don't expect paywalled or authenticated content to be seen. No AI crawler authenticates; only freely accessible public content counts.
- Track what actually gets cited, not just what gets crawled. A technically perfect, fully crawlable site can still be absent from answers if the content lacks the evidence density these engines look for.
If you want to know precisely which engines actually see your content and which ones cite you, check out Traaker's method for measuring and improving your visibility in AI search.
Keep reading
These crawling architectures are only the first step. Once an engine can see your content, it still has to choose to cite it, which depends on factors specific to each engine: read our guides on how to get cited by ChatGPT, how to get cited by Claude, and how to get cited by Perplexity.
Methodology and limitations. This article synthesizes public data from third-party sources (Vercel, Zyte, Cloudflare Radar, Perplexity Research, Anthropic, OpenAI) collected at different dates and with different methodologies; the percentages and volumes should be read as orders of magnitude rather than strictly comparable measurements. AI crawler behavior evolves quickly, and some sources do not fully agree on the exact timeline of certain events (notably when Cloudflare first detected Perplexity's stealth crawling). Always verify the current behavior of your own robots.txt and server logs before making a blocking decision.
Sources
- SearchEngineWatch
- TechCrunch
- Perplexity Research
- Cloudflare
- ipregistry.co
- Adobe Experience League
- developers.openai.com
- Search Engine Journal
- Vercel
- OpenAI Help Center
- Search Engine Land
- anthropic.com/crawl
- Anthropic Docs
- Anthropic
- 404 Media
- The AI Engineer
- Perplexity Docs
- The Verge
- MarkTechPost
- WebSearchAPI
- NuxtSEO
- Search Engine Journal
- NuxtSEO
- Zyte
- Zyte, 2026 analysis of 11,100 pages
- BuzzStream
- PCMag
- Oncrawl
- SEOptimer
- CJR
- arXiv
- Frase
- Contently