Computational Marketing

Technical SEO Checklist for AI Crawler Compatibility

AI bots ignore noindex; robots.txt is your only control lever.

Senior Writer · · 10 min read
Cover illustration for “Technical SEO Checklist for AI Crawler Compatibility”
Bots as a New Buyer Class · October 5, 2026 · 10 min read · 2,293 words

A site can hold steady rankings in Google while its presence in ChatGPT, Claude, and Perplexity answers quietly empties out, and the dashboards that track classic search will show nothing wrong. The four crawlers most practitioners now have to govern explicitly, GPTBot from OpenAI, PerplexityBot from Perplexity AI, ClaudeBot from Anthropic, and Google-Extended, Google's AI training and grounding crawler, differ from Googlebot in rendering capability, crawl behavior, and purpose. A WAF rule, CDN configuration, or bot-management policy written without distinguishing training bots from real-time citation bots can drop AI crawler traffic entirely while Google traffic continues as expected, and with AI bots now accounting for roughly 40 to 50 percent of Googlebot-level activity across the web, that failure mode scales with the stakes. The two surfaces, classic search ranking and AI-mediated discovery, have decoupled far enough that strong performance on one tells a team nothing reliable about performance on the other.

AI crawlers versus Googlebot: the differences that drive checklist divergence

The divergence between the two checklists traces back to three concrete differences: rendering capability, crawl depth behavior, and how reliably each crawler type honors standard SEO directives [1][2][3]. On rendering, almost no AI crawler executes JavaScript. The practical consequence is that content or schema markup appearing only after JavaScript executes is invisible to most AI engines, which turns server-side rendering or static site generation into a hard crawlability requirement rather than a performance nicety.

Directive obedience produces a second, subtler gap. AI search and citation crawlers behave differently here: OAI-SearchBot is documented as respecting the noindex signal, even though training crawlers largely ignore it. If an internal link points to a suppressed page, a training bot will follow it and read the content behind the directive meant to hide it. Pages a team deliberately suppressed from Google search results can quietly become the AI-facing surface of a brand, with no one on the team aware that it is happening.

Crawl depth is the third divergence, and it behaves unevenly across the category. AI search bots drop off sharply beyond two to three clicks from the homepage, with average visits falling to roughly one per page once a page sits three or more clicks deep. AI training bots are less constrained by depth, but a training visit alone does not generate user-facing visibility. That distinction matters because "AI crawler" is not one behavior. AI search bots crawl to find new URLs and fetch HTML for answers, and how deep they'll go runs much closer to Googlebot's behavior. AI user bots are the third category: triggered the moment a real person asks a question inside ChatGPT, Claude, or Perplexity, these bots fetch live content to research the answer in real time, and their activity is the closest proxy available to an impression in an AI interface. User bot activity is what actually connects crawling to visibility, which is the reason the rest of this checklist treats crawl access as a means to that end rather than as the end itself.

Diagram: AI Crawler Types: Three Distinct Behaviors. Visualizes: Visualize the three categories of AI crawler and their key behavioral differences to clarify why 'AI crawler' is not one behavior.

Robots.txt: the one access-control file AI bots respect

Canonical tags and meta noindex carry limited weight with AI training bots, so robots.txt is the only reliable lever you have over what AI systems can access. That file is now carrying a job it was never built for: for AI crawlers specifically, it is the primary access-control mechanism, full stop, and treating it as a boilerplate artifact left over from a CMS default is a mistake with real consequences. The checklist item most teams skip is writing explicit rules for GPTBot, ClaudeBot, PerplexityBot, CCBot, and Google-Extended, choosing to allow or block each one by deliberate strategy rather than leaving the decision to omission. Omission defaults to allow, which works fine if that's what you want, but plenty of teams assume noindex handles content protection and never check whether AI bots are reading suppressed pages anyway.

In robots.txt governance, the split that matters most separates discovery from ingestion. AI search and indexing bots, PerplexityBot and OAI-SearchBot among them, should generally be allowed if real-time citation in an AI answer is the goal. Blocking a training bot does not block a citation bot, and allowing a citation bot does not mean training ingestion is also happening. If a team wants AI citation without training ingestion, it needs to write separate, explicit rules for each user-agent, not apply a single blanket allow or block across the board.

A Disallow: / left in robots.txt after a staging deployment is the single most catastrophic finding an audit turns up, a line that blocks every crawler including every AI bot, and it stays invisible to the team because Google traffic may keep flowing from cached index data even after the block goes live. Verifying the production robots.txt file is the first check in any audit, before any other item on this list gets attention. Confirm the file is at the root of the domain, returns a 200 status, and contains no syntax errors that inadvertently block entire user-agents. A file some teams look to as an alternative control layer, llms.txt, does not belong in this conversation: server log data shows the major AI bots do not read it, so it has no place as a crawl-control mechanism regardless of how it gets discussed elsewhere.

Site structure and internal linking as the crawl budget AI bots spend

A correctly configured robots.txt only grants permission to enter a site; it does nothing to help a crawler find its way around once inside, and that is where internal linking takes over. For AI crawlers, internal linking functions as the literal navigation map rather than a ranking signal, and for Perplexity it is reportedly the only discovery mechanism available. The crawl-depth drop-off described earlier creates the same structural risk here, but the stakes run higher, because AI search bots carry a much smaller crawl budget than Googlebot, so they can't tolerate poor site structure the way Googlebot can. Content that Googlebot would eventually find through a patient, deep crawl can remain essentially invisible to an AI search crawler that stops looking after two or three clicks, so critical pages need to sit within three to four clicks of the homepage, and anything deeper should be evaluated for whether it needs to move closer in or pick up links from a shallower hub page. AI bots also do not click buttons, so content tucked behind accordions, tabs, dynamic filters, or anything that loads only on user interaction stays invisible to them; this limitation isn't new for Googlebot either, but a smaller crawl budget makes the cost of interaction-dependent content proportionally higher for AI bots than it ever was for Google's crawler. Faceted navigation, calendar widgets, and session IDs deserve a check as well, since any of them can generate crawl traps or endless URL variations that waste the limited budget an AI bot has to spend. XML sitemaps still matter for the platforms that use them, including ChatGPT and Claude, and the standard hygiene rules still apply: the sitemap should be linked from robots.txt, return a 200 status, stay under size limits, carry a lastmod date that reflects real content changes rather than a blanket redeployment timestamp, and contain only canonical URLs, with no redirected, noindexed, 404, or parameterized duplicates cluttering the list.

JavaScript rendering: what server-side delivery means for AI crawler access

If you want full AI coverage, server-side rendering or static site generation is a crawlability requirement, not a performance optimization you can defer. Any content or schema markup that exists only after client-side JavaScript executes is invisible to GPTBot, PerplexityBot, and ClaudeBot, full stop on the consequence even if the mechanism behind it is simple. Google-Extended remains the exception here too, benefiting from Googlebot's two-wave rendering pipeline the way it benefits from it elsewhere, while the other three major AI crawlers have no equivalent rendering capability and receive nothing more than whatever HTML the server delivers on first response.

In practice, this failure most often traces back to hydration events in frameworks like React, Next.js, Angular, and Vue. Content can render correctly for a human visitor and survive Googlebot's rendering pipeline without issue, yet the raw HTML an AI crawler actually gets can be a near-empty shell with no text it can pull out. Schema markup needs the same scrutiny: it has to be present in the server-delivered HTML, not injected by a client-side script, and placing JSON-LD in the document head is the safest way to guarantee that. INP optimization sits adjacent to this problem: heavy JavaScript execution that blocks the main thread produces poor INP scores for human users, and the same long-running tasks that degrade INP often delay or prevent the same content from rendering in time for a crawler that isn't waiting around.

Structured data and schema markup: the explicit labeling layer AI systems need to extract and cite content

Structured data has moved well past the status of a nice-to-have enhancement. It now functions as the primary language through which an AI system identifies what a page is about, who published it, and whether it carries enough authority to be worth citing. Several schema types carry particular weight in that decision. Article or BlogPosting schema on every content page signals content type explicitly, which helps an AI engine categorize the page correctly and extract from it with confidence. Organization schema on the homepage and About page establishes the entity signals an AI engine uses to identify a company and connect scattered mentions of it across the web. HowTo schema on instructional content gives AI engines a structured format they prioritize when answering process-based questions.

Placement matters as much as presence. JSON-LD in the document head is the correct location for all of this, because schema injected by client-side JavaScript will not be seen by any AI crawler that does not render JavaScript, which is most of them, as the rendering section already established. Rel=canonical, hreflang, sitemap entries, and internal links should all point to the same preferred URL, because conflicting signals create ambiguity that AI systems tend to resolve by deprioritizing the content rather than guessing which version is authoritative. Organization schema should connect to consistent name, address, and phone data across the site and any external directories that list the business, because if the entity signal is fragmented, it's harder for an AI system to confirm that scattered mentions all refer to the same brand.

Content hierarchy and heading structure as the extraction surface for AI-generated answers

AI systems extract answers directly from HTML heading structure, so they treat H2s and H3s as independently extractable answer snippets, not just visual breaks in a page. That makes heading writing a technical SEO decision now, not only an editorial one. Every page needs exactly one H1 that states its topic clearly, because AI engines use it as the primary signal when they extract the topic. H2s need to function as section headers that stand on their own as descriptive phrases, because AI systems often pull H2 content out as an answer snippet, apart from the paragraph text that follows it. H3s should break down their parent H2 sections in a logical order, and you should never skip a heading level or jump from an H1 straight to an H4, because an irregular hierarchy confuses the parsers AI retrieval systems rely on to map a page's structure.

Content depth carries a related but more nuanced signal. Longer, more comprehensive content averages more AI citations on some platforms, ChatGPT among them, but the same correlation nearly disappears on others, Google AI Overviews showing close to no relationship between word count and citation likelihood. Freshness feeds into the same calculation. A concise, direct answer in the opening paragraph of a section, one that answers the heading's implicit question before the prose elaborates further, increases the odds that an AI system extracts it cleanly.

Core Web

A page-speed scoring system still plays a role in this checklist, though a narrower one than its role in classic SEO. One such metric, a responsiveness score that replaced an earlier input-delay measure, measures how responsive a page feels to a human user interacting with it, and it depends heavily on how much JavaScript execution blocks the main thread during page load. The same long-running tasks that produce a poor INP score are frequently the tasks that delay or prevent content from finishing its render in time for a crawler to see it, which ties this metric back to the rendering section's core argument even though its primary audience is a human visitor, not a bot. If a team runs both checklists in parallel, it will find that fixing INP issues often fixes rendering-timing issues for AI crawlers too, because both problems trace back to the same bloated JavaScript execution path.

None of this is verifiable from the outside without measurement, and that is the gap a robots.txt audit alone cannot close. A rule written to allow PerplexityBot or block GPTBot is a statement of intent, not a confirmation that the intent produced the outcome. Letterstory, among other platforms built for this purpose, tracks whether ChatGPT, Claude, Gemini, and Perplexity actually cite a brand in response to real queries, which gives a team a way to check that its robots.txt configuration, rendering setup, and schema markup are together producing the visibility they were designed to produce, rather than assuming compliance because the rule exists on paper. The two checklists in this piece, one for Googlebot and one for the crawlers now reading the same sites on behalf of AI systems, will keep diverging as each ecosystem evolves on its own schedule. Treating them as one checklist is the single fastest way to lose visibility on one surface while believing, based on the other, that everything is fine.

More in Bots as a New Buyer Class