Field note · AI readiness
AI crawler readiness: a practical policy for websites and webshops.
A useful AI policy separates search visibility, user-requested retrieval, and model training. It then connects those choices to crawlable content, authoritative data, structured information, and repeatable verification.
Published 28 July 2026
8 minute read
PitchAI
In brief
The decision is broader than allow or block.
- Assign one business owner for crawler visibility, training permissions, protected paths, and review dates.
- Treat search crawlers, user-requested fetchers, and training crawlers as separate controls.
- Make core company, service, product, policy, price, and availability information clear in crawlable HTML.
- Use canonical URLs, internal links, sitemaps, accurate metadata, and structured data that matches the visible page.
- Verify robots.txt and the actual CDN or firewall response with the relevant user agents; do not assume the file is the whole policy.
01 · The policy gap
Silence in robots.txt is not evidence of a decision.
Emerce reviewed the robots.txt files of 100 Dutch webshops in July 2026. Eleven were hidden behind bot walls; of the 89 that could be analysed, 74—83%—named none of GPTBot, OAI-SearchBot, PerplexityBot, or ClaudeBot. Large chains and smaller shops showed essentially the same pattern.
The article also found repeated host and plugin configurations, suggesting that some of the exceptions came from defaults rather than a deliberate company policy. That distinction matters. A technically valid file can still reflect no considered position on answer-engine visibility, model training, data freshness, or ownership.
The business implication extends beyond e-commerce. Companies need someone to connect marketing intent, governance, web architecture, data ownership, and operational verification. Crawler readiness belongs in an AI implementation programme, not as a one-off SEO edit.
02 · Separate the uses
Search, retrieval, and training are three policy choices.
OpenAI documents OAI-SearchBot and GPTBot as independent controls. OAI-SearchBot is used to surface sites in ChatGPT search; GPTBot is used for potential model training. ChatGPT-User supports user-initiated visits and does not determine search inclusion. A company can therefore support ChatGPT search without treating training access as the same decision.
Anthropic similarly distinguishes Claude-SearchBot, Claude-User, and ClaudeBot. Perplexity distinguishes its search crawler from its user-requested fetcher. The names and behaviour can change over time, so the policy should link to current vendor guidance and carry a review date.
Search discovery
Decide whether public pages should be eligible for indexing, retrieval, and citation in AI-assisted search.
User-requested retrieval
Decide whether an answer engine may visit a public page in response to a user's direct request.
Training and model use
Decide separately whether public material may be collected for model development or other model uses.
03 · Information quality
Crawler access does not make an unclear company understandable.
Google's guidance for AI features is refreshingly plain: established search fundamentals still apply. Important content should exist in textual form, internal links should make it findable, crawling must work through robots and hosting infrastructure, and structured data must match the visible page. Google explicitly says that no special AI text file or AI-specific schema is required.
For a service company, the public source of truth should explain what the company does, who it serves, where it operates, how engagements work, and how to make contact. For a webshop, it also needs accurate product names, descriptions, variants, prices, availability, delivery information, returns, and policy pages. Structured data can clarify those facts, but it must not introduce claims that a visitor cannot verify on the page.
This is where AI crawler readiness becomes data readiness. A crawler can only repeat what the public system makes available. Stale stock, contradictory service descriptions, vague claims, and hidden JavaScript-only copy create an accuracy problem even when access is fully open.
04 · Implementation
Use a short, evidence-based implementation sequence.
-
01
Inventory
List public domains, subdomains, CDN and firewall layers, robots files, sitemaps, canonical rules, protected routes, and accountable owners.
-
02
Decide
Record the policy for search discovery, user-requested retrieval, training, rate limits, sensitive paths, and the next review date.
-
03
Clarify
Create or revise concise pages for services, products, company facts, governance, policies, and common decision questions.
-
04
Structure
Use semantic headings, descriptive links, canonical URLs, useful metadata, current XML and HTML sitemaps, and page-appropriate structured data.
-
05
Configure
Apply the policy in robots.txt and at the CDN or firewall. Where a vendor publishes crawler IP ranges, use the current official source rather than a copied static list.
-
06
Verify
Fetch representative pages with each relevant user agent, confirm status and rendered text, validate sitemaps and structured data, and retain dated evidence.
05 · Verification
Measure eligibility and accuracy—not promised placement.
No agency can guarantee that an answer engine will index, rank, cite, or recommend a page. The controllable outcome is narrower: the intended public URLs are reachable, the crawler policy is explicit, the information is coherent, and technical signals do not contradict the visible page.
Verification should cover HTTP responses, robots matching, meta directives, canonicals, internal links, sitemap membership, rendered text, structured data, redirects, and edge security. Server logs can then show whether legitimate crawlers arrive and whether they encounter errors. Bing's current webmaster guidance adds the same operational point: sound SEO supports Copilot and grounding visibility, while updated sitemaps and IndexNow can help search systems notice changed URLs sooner.
Revisit the evidence after platform changes, website releases, product-feed changes, or policy decisions. A dated review cycle is what turns a configuration into governance.
Primary guidance
Keep the policy tied to current vendor documentation.
AI readiness
Need the policy, website, data, and implementation to agree?
PitchAI helps organisations audit the current state, record the decisions, implement the controls, improve the public source of truth, and verify the result.