PitchAI

AI crawler readiness: a practical policy for websites and webshops.

A useful AI policy separates search visibility, user-requested retrieval, and model training. It then connects those choices to crawlable content, authoritative data, structured information, and repeatable verification.

Published 28 July 2026

8 minute read

PitchAI

The decision is broader than allow or block.

  • Assign one business owner for crawler visibility, training permissions, protected paths, and review dates.
  • Treat search crawlers, user-requested fetchers, and training crawlers as separate controls.
  • Make core company, service, product, policy, price, and availability information clear in crawlable HTML.
  • Use canonical URLs, internal links, sitemaps, accurate metadata, and structured data that matches the visible page.
  • Verify robots.txt and the actual CDN or firewall response with the relevant user agents; do not assume the file is the whole policy.

Silence in robots.txt is not evidence of a decision.

Emerce reviewed the robots.txt files of 100 Dutch webshops in July 2026. Eleven were hidden behind bot walls; of the 89 that could be analysed, 74—83%—named none of GPTBot, OAI-SearchBot, PerplexityBot, or ClaudeBot. Large chains and smaller shops showed essentially the same pattern.

The article also found repeated host and plugin configurations, suggesting that some of the exceptions came from defaults rather than a deliberate company policy. That distinction matters. A technically valid file can still reflect no considered position on answer-engine visibility, model training, data freshness, or ownership.

The business implication extends beyond e-commerce. Companies need someone to connect marketing intent, governance, web architecture, data ownership, and operational verification. Crawler readiness belongs in an AI implementation programme, not as a one-off SEO edit.

Source: Emerce, 28 July 2026

Search, retrieval, and training are three policy choices.

OpenAI documents OAI-SearchBot and GPTBot as independent controls. OAI-SearchBot is used to surface sites in ChatGPT search; GPTBot is used for potential model training. ChatGPT-User supports user-initiated visits and does not determine search inclusion. A company can therefore support ChatGPT search without treating training access as the same decision.

Anthropic similarly distinguishes Claude-SearchBot, Claude-User, and ClaudeBot. Perplexity distinguishes its search crawler from its user-requested fetcher. The names and behaviour can change over time, so the policy should link to current vendor guidance and carry a review date.

Search discovery

Decide whether public pages should be eligible for indexing, retrieval, and citation in AI-assisted search.

User-requested retrieval

Decide whether an answer engine may visit a public page in response to a user's direct request.

Training and model use

Decide separately whether public material may be collected for model development or other model uses.

OpenAI's current crawler documentation

Crawler access does not make an unclear company understandable.

Google's guidance for AI features is refreshingly plain: established search fundamentals still apply. Important content should exist in textual form, internal links should make it findable, crawling must work through robots and hosting infrastructure, and structured data must match the visible page. Google explicitly says that no special AI text file or AI-specific schema is required.

For a service company, the public source of truth should explain what the company does, who it serves, where it operates, how engagements work, and how to make contact. For a webshop, it also needs accurate product names, descriptions, variants, prices, availability, delivery information, returns, and policy pages. Structured data can clarify those facts, but it must not introduce claims that a visitor cannot verify on the page.

This is where AI crawler readiness becomes data readiness. A crawler can only repeat what the public system makes available. Stale stock, contradictory service descriptions, vague claims, and hidden JavaScript-only copy create an accuracy problem even when access is fully open.

Google Search Central guidance for AI features

Use a short, evidence-based implementation sequence.

  1. 01

    Inventory

    List public domains, subdomains, CDN and firewall layers, robots files, sitemaps, canonical rules, protected routes, and accountable owners.

  2. 02

    Decide

    Record the policy for search discovery, user-requested retrieval, training, rate limits, sensitive paths, and the next review date.

  3. 03

    Clarify

    Create or revise concise pages for services, products, company facts, governance, policies, and common decision questions.

  4. 04

    Structure

    Use semantic headings, descriptive links, canonical URLs, useful metadata, current XML and HTML sitemaps, and page-appropriate structured data.

  5. 05

    Configure

    Apply the policy in robots.txt and at the CDN or firewall. Where a vendor publishes crawler IP ranges, use the current official source rather than a copied static list.

  6. 06

    Verify

    Fetch representative pages with each relevant user agent, confirm status and rendered text, validate sitemaps and structured data, and retain dated evidence.

Measure eligibility and accuracy—not promised placement.

No agency can guarantee that an answer engine will index, rank, cite, or recommend a page. The controllable outcome is narrower: the intended public URLs are reachable, the crawler policy is explicit, the information is coherent, and technical signals do not contradict the visible page.

Verification should cover HTTP responses, robots matching, meta directives, canonicals, internal links, sitemap membership, rendered text, structured data, redirects, and edge security. Server logs can then show whether legitimate crawlers arrive and whether they encounter errors. Bing's current webmaster guidance adds the same operational point: sound SEO supports Copilot and grounding visibility, while updated sitemaps and IndexNow can help search systems notice changed URLs sooner.

Revisit the evidence after platform changes, website releases, product-feed changes, or policy decisions. A dated review cycle is what turns a configuration into governance.

Bing webmaster guidance for search and grounding

Need the policy, website, data, and implementation to agree?

PitchAI helps organisations audit the current state, record the decisions, implement the controls, improve the public source of truth, and verify the result.