Bots and Crawler Guidance & Decision Matrix

On July 28, 2026 Technical Standards and Specs

It’s not who the crawler is, it’s what happens to your content after it leaves. Five jobs, five different answers.

The intention of this guidance is to support content owners of all types to gain deeper awareness around crawlers and bots, and then consider establishing a policy and approach to optimise accordingly based on required outcomes. The aim is never to simply block every crawler, rather it’s to be able to tell beneficial, commercial, operational and unknown non-human traffic apart, so that each can be managed on its merits. Volume isn’t always true value as AI sources will often send less traffic than classic search, but that traffic is often higher-intent and converts better (as per the agentic-traffic panel below).

Crawler job What you gain if allowed What you risk if allowed What blocking costs you Publisher trend Brand trend
ADiscovery & searchGooglebot · bingbot · OAI-SearchBot Referral traffic, citations in AI answers, brand visibility The same crawl can also feed AI answer features built on the index De-indexing; loss of organic search traffic Allow Allow
BAI trainingGPTBot · ClaudeBot · CCBot Presence and familiarity inside future models; goodwill with AI operators Uncompensated extraction; forgone licensing value; irreversible once crawled Not learned by future models results in invisible loss & no clear signal Block, or license Selective (product facts in, premium IP out)
CLive AI agentsChatGPT-User · Claude-User · Google-Agent A waiting human gets served; task completion; live-answer citations No pageview or ad impression; paywall probing Invisible at the moment of decision for assistant users Conditional; restrict on ad-funded pages Allow on product & knowledge pages
DOperationalAdsBot-Google · verification · previews Ad delivery, verification, measurement, uptime, link previews Negligible; some competitive visibility via SEO tools Broken ad stack; programmatic revenue collapse; ugly link shares Always allow (verified) Always allow (verified)
EUnknown / unverifiedspoofed UAs · python-requests · proxies Nothing knowable Scraping, fraud, reconnaissance everything Nothing Challenge or block; investigate volume Challenge or block
1 · AllowValue returned justifies unrestricted access.
2 · Allow with conditionsRate limits, path scope, citation or attribution terms.
3 · Require licensingAccess only under a commercial agreement.
4 · BlockNo value, real cost, or unverifiable identity.

Decide per crawler group, then per crawler within high-risk groups. Document each decision with a rationale and a review date. (Framework: IAB Tech Lab Bot & Crawler Management Guidance.)

The asymmetryUnblocking a valuable crawler takes a minute. Recollecting content that has already been crawled is impossible. When unsure, start conservative.
The plumbing warningBlock everything and you break revenue: ads.txt verification alone underpins programmatic yield. Never block verified operational bots.
Before 15 September 2026: if your site sits behind Cloudflare or any managed edge/WAF, review your AI-crawler defaults. New defaults will block Training and Agent bots on ad-displaying pages for new domains and free-tier zones and multi-purpose crawlers are assessed under the most restrictive applicable rule.

Sources: IAB Tech Lab Bot & Crawler Management Guidance (2026) and CoMP v1 (finalised 28 April 2026); Cloudflare Content Independence Day announcements (1 July 2026). Verify all user-agent tokens against the IAB/ABC International Spiders & Bots List and operator documentation before implementation. v1.0 · 22 July 2026. General information, not legal advice.

Managing Crawlers & AI Bots

A starter guide for content owners: who’s at your door, what they do with your content, and what you control.

Three shifts have altered the traditional crawl-for-clicks bargain. Search has become answeringwith results pages increasingly answering directly, making the click optional. Crawling has became multi-purpose, the same page fetch can now feed a search index, AI answers, model training, live agent tasks, verification and analytics. And control moved to the network so whilst robots.txt remains advisory, CDN and WAF enforcement is not. In mid-2026, automated requests surpassed human ones in web-page traffic for the first time, and the majority of identified crawler activity now relates to AI. The practical response is not panic or a blanket block: it is knowing the five jobs bots do on your site, and deciding each on its merits.

How to use this guide. Tokens below are user-agent identifiers you will see in server logs, robots.txt and bot-management tools, verified against operator documentation and Cloudflare’s AI Crawl Control reference as at July 2026. Tokens will change frequently, so regularly re-verify against the IAB/ABC International Spiders & Bots List and each operator’s documentation before implementing rules.

Start here (by role)Three proposed starting positions, with full reasoning in the matrix above and the playbook further down.

Brand / commerce

Allow discovery and live agents. Review training as a separate decision with product facts in and sensitive IP out. Measure AI referrals.

Publisher / premium IP

Allow discovery. Audit training and agent exposure before blocking. Publish a crawler policy with a licensing contact.

Technical owner

Visibility → verification → policy. Verify identity by network, not by name. Enforce at the edge for high-risk traffic.

ADiscovery & search indexersBuild an index so you can be found later. The librarian: reads your work, catalogues it, sends readers back. The category most likely to return measurable referral traffic. Common examples are below.

User-agent token Operator What it does Notes & stance
Googlebot Google Core web-search index; the same crawl also feeds AI answer features Multi-purpose (per Cloudflare’s classification); blocking means de-indexing
Googlebot-Image / -News / -Video · Storebot-Google Google Vertical indexing: images, news, video, shopping Product-specific tokens under the Googlebot family
GoogleOther Google Generic crawler for R&D and other products Purpose-ambiguous by design — worth watching in logs
bingbot Microsoft Bing search index; grounds Copilot answers Multi-purpose (per Cloudflare’s classification)
Applebot Apple Siri, Spotlight and Safari suggestions; Apple Intelligence features Multi-purpose (per Cloudflare’s classification)
Amazonbot Amazon Alexa answers and discovery AI use noted in Amazon’s documentation
OAI-SearchBot OpenAI ChatGPT search index Declared search-only; separate from GPTBot (training)
Claude-SearchBot Anthropic Claude search index Split out from training in 2025
PerplexityBot Perplexity Answer-engine index Powers Perplexity’s cited answers
YouBot You.com AI search index Early Pay-Per-Use compensation partner
DuckDuckBot · MojeekBot · PetalBot DuckDuckGo · Mojeek · Huawei Independent and regional indexes Mojeek runs a fully independent UK index
YandexBot · Baiduspider · Yeti Yandex · Baidu · Naver Regional majors (RU, CN, KR) Relevant for international audiences

BAI training crawlersAbsorb your content into a model. The student who memorises your book and never brings it back. Nothing is returned by default, and once taken it can’t reliably be recalled from copies, corpora or trained models. Common examples are below.

User-agent token Operator What it does Notes & stance
GPTBot OpenAI Model training Widely blocked by site owners
ClaudeBot Anthropic Model training Distinct from Claude-SearchBot and Claude-User
Meta-ExternalAgent Meta Training and model improvement Covers Llama and Meta AI products
CCBot Common Crawl Non-profit corpus feeding many downstream trainers Blocking exits many datasets at once — including open-source models
Bytespider ByteDance Model training Compliance practices vary — use verified identity and edge enforcement for high-risk traffic
cohere-training-data-crawler Cohere Enterprise model training
AI2Bot Allen Institute for AI Research corpora (e.g. Dolma) Academic; a values call for some owners
omgilibot / Webzio-Extended Webz.io Data broker supplying AI buyers Downstream use opaque
Diffbot Diffbot Structured extraction and dataset sales
Google-CloudVertexBot Google Cloud Crawls sites at Vertex AI customers’ request Site-owner-initiated; no effect on Google Search
Control tokens, not crawlers‘Google-Extended’ and ‘Applebot-Extended’ will never visit your site. They are robots.txt directives honoured by the operator’s main crawler, and their scope is model training only so they do not remove your content from answer features built on the search index. This single mechanic is why ‘one crawler, many jobs’ matters: with multi-purpose crawlers, the training opt-out and the answer-feature exposure are separate questions.

CLive agents & user fetchersFetch a page right now because a human asked. The personal shopper: a real person is served, but there’s no pageview and no ad impression. Common examples are below.

User-agent token Operator What it does Notes & stance
ChatGPT-User OpenAI Fetches when a ChatGPT user asks User-initiated; OpenAI’s docs (updated Dec 2025) state robots.txt may not apply — enforce at the edge if needed
Claude-User Anthropic Fetches when a Claude user asks Anthropic states all its bots honour robots.txt, incl. Crawl-delay
Perplexity-User Perplexity User-initiated fetches Docs state user fetches may not honour robots directives
Meta-ExternalFetcher Meta Assistant link fetches
MistralAI-User Mistral Le Chat fetches
DuckAssistBot DuckDuckGo AI answer fetches
Google-Agent Google User-directed browsing on Google infrastructure (Gemini Agent, Chrome Auto Browse) Added early 2026 with its own IP range file; verify current docs, ranges and robots behaviour before writing a rule
Google-NotebookLM Google Fetches user-added sources for notebooks Being renamed after the Gemini Notebook rebrand so the old token will deprecate August 2026; please check current docs
Browser-use agents (various) Various Operate a real browser on a user’s behalf Present as near-human traffic
robots.txt treatment often differs by operator and it’s now three waysFor user-initiated fetchers the rules diverge: Google documents that its user-triggered fetchers (including ‘Google-Agent’) generally ignore robots.txt because a person requested the page; OpenAI’s December 2025 documentation states that robots.txt may not apply to ‘ChatGPT-User’ for the same reason; while Anthropic states all its bots, including ‘Claude-User’ will honour robots.txt. The lesson: for this category, robots.txt is not a reliable control so always check each operator’s current documentation, and enforce at the edge where it matters.

The shape of agentic traffic (mid-2026)Agents are the fastest-moving force on the web and the gap between what they take and what they return has never been wider.

+45%
AI-agent requests, quarter on quarter, 17.7 billion in Q2 2026 (DataDome)
57.5%
of web-page requests are now automated (not human) the first such crossover (Cloudflare, Jun 2026)
~52%
of AI crawler requests are for training; only ~2.6% are real-time, human-triggered fetches (Cloudflare)
80 – 88%
of AI referral traffic comes from ChatGPT even as its crawl volume fell (DataDome)

The defining feature of agent traffic is asymmetry. A person shopping for shoes might visit five sites; an agent completing the same task can visit hundredsand automated traffic is now expanding several times faster than human activity (HUMAN Security, 2026). That volume doesn’t map to value: crawl activity and referral value are moving in opposite directions. ChatGPT crawled around 6% less quarter on quarter yet sent roughly 17% more visitors, while Meta’s two agents generated more than half of all AI traffic on one major network and returned almost none. The lesson mirrors the matrix, so judge a crawler by what comes back, not by how often it knocks.

Composition shifts constantly, which is why the playbook says review quarterly. Operators are splitting their do-everything crawlers into clearer single-purpose ones. Mixed-purpose crawling fell from roughly 49% to 33% of AI requests across the first half of 2026, and the rankings reorder month to month (Meta’s indexing crawler overtook its training crawler in June). Yesterday’s allowlist is not today’s.

A note on MCP and agentic commerceModel Context Protocol traffic has gone from negligible to peaks near 500,000 requests a day (according to DataDome) which is an early signal of agents taking inventory of what they can act on before they act. The commercial stakes are rising with it: Gartner expects AI agents to intermediate more than US$15 trillion in B2B purchasing by 2028, and AI-referred traffic already converts around 42% better than non-AI visits (Adobe, 2026). For brands, being readable by agents is becoming a revenue question, not just a technical one.

Figures are drawn from each provider’s own network and measurement window, so treat them as directional rather than precise, but the direction is clear. Sources: DataDome AI Traffic Report Q2 2026; Cloudflare Radar (mid-2026); HUMAN Security 2026 State of AI Traffic; Gartner; Adobe Digital Insights. Please verify against primary sources regularly.

DOperational & advertising infrastructureKeep the money and machinery working. The meter readers: rarely glamorous, never block the verified ones.

User-agent token Operator What it does Notes & stance
AdsBot-Google / AdsBot-Google-Mobile Google Ad landing-page quality checks Blocking harms Google Ads delivery and quality scoring
Mediapartners-Google Google AdSense contextual crawling Required for AdSense targeting
AdIdxBot Microsoft Microsoft Advertising quality and verification Bing Ads equivalent of AdsBot
GoogleInspectionTool Google Search Console testing (URL Inspection, Rich Results) Commonly seen in logs and wrongly blocked; leave allowed for your own diagnostics
proximic · peer39_crawler Comscore · Peer39 Contextual classification for targeting Blocking reduces contextual demand
Verification fetchers (DV, IAS, HUMAN) Various Viewability, fraud and brand-safety verification Tokens vary, so consult the IAB/ABC Spiders & Bots List if possible; blocking breaks verification and ads.txt-dependent revenue
AhrefsBot · SemrushBot · Screaming Frog · MJ12bot · DotBot SEO vendors Site audit, backlinks, rankings Gray zone: useful to you and to competitors; rate-limit rather than block outright
facebookexternalhit · Twitterbot · LinkedInBot · Slackbot · Discordbot · WhatsApp · TelegramBot · Pinterestbot Platforms Link previews when your URLs are shared Blocking makes your shares ugly
UptimeRobot · Pingdom Monitoring vendors Uptime checks Low volume; often run by you or your host
Feedfetcher-Google · Feedly · podcast fetchers Various Feed and podcast distribution Distribution you probably want
archive.org_bot / ia_archiver · TurnitinBot Internet Archive · Turnitin Web archival; academic integrity A values call as much as a commercial one

EImpostors & unknownsThe gatecrashers. Treat unidentifiable traffic as information, not noise.

A meaningful share of bot traffic matches no known signature: spoofed major-crawler user agents from the wrong networks, default library strings (python-requests, curl, Go-http-client, Scrapy), headless browsers, and residential-proxy fleets. Some operators’ crawler identities also remain poorly documented publicly (xAI/Grok and DeepSeek are current examples) which is the transparency problem in miniature: treat as unknown until verified. High-volume unidentified traffic deserves investigation, not a shrug.

Not every Googlebot is GooglebotUser-agent strings are voluntary text, so any scraper can claim to be anyone. Verify identity with the network, not the name, using each operator’s published method. For Google: reverse-resolve the source IP, confirm the hostname is an approved Google domain (e.g. *.googlebot.com), then forward-resolve that hostname and check it returns the original IP. Most major operators publish IP-range files or a DNS process for exactly this. The three-step check: 1) match the UA against a maintained directory; 2) confirm the source IP against the operator’s published ranges or forward-confirmed reverse DNS (an ASN such as Google’s AS15169 is a supporting signal, not proof on its own); 3) confirm behaviour fits the declared purpose (crawl rate, paths, robots.txt compliance). A claimed major-crawler UA arriving from a residential ISP or generic cloud range is almost certainly an impostor.
Before you block an unknown, check it’s not yours!Some unidentified traffic is your own stack: CMS, site search, translation or personalisation vendors, tag-management, affiliate-compliance platforms, data-feed partners or support bots. Cross-reference high-volume unknowns against procurement, tag inventories and engineering before blocking, or you may break something you’re paying for.

The playbook: five moves, in order

  1. Get visibility first. The problem isn’t crawling, it’s invisibility. Audit at least 30 days of edge or server logs before changing anything; owners who block blind often block crawlers that drive real referrals. Most CDNs and WAFs can surface bot traffic by user-agent, ASN and verified status. Request a regular 30-day report from your vendor(s) if you don’t already receive one.
  2. Evaluate each crawler on value. Traffic return, content use, bandwidth cost, reputation, strategic value. Try to score them accordingly and the decision usually forms naturally.
  3. Apply the four options per group: Allow · Allow with conditions · Require licensing · Block. Within high-risk groups, decide per crawler.
  4. Document and review. Record decision, rationale, conditions and a review date. Check for new high-volume unknowns weekly; review bandwidth monthly; re-run the full assessment quarterly.
  5. Publish your policy. A human-readable crawler policy page, a machine-readable robots.txt, and a licensing contact. A visible policy turns a passive block into a potential commercial relationship.
What to measure per crawlerStep 2 is easier with the right fields. Judge each verified crawler on value returned and cost consumed: requests and bytes served, pages crawled, referrals and conversions, citations, and any licensing revenue against origin cost and risk exposure. You can’t manage what you don’t measure.
Australian context (July 2026)Australia has no US-style ‘fair use’, only specific fair dealing exceptions none of which covers AI training. In October 2025 the Attorney-General ruled out introducing a text-and-data-mining exception, and the Productivity Commission’s final report (December 2025) likewise recommended against one, preferring a wait-and-see approach. Consultation has moved to licensing and compensation models through the Copyright and AI Reference Group (CAIRG). Hence the approach in this guide: content used on terms, not by default. Treat this as general information, not legal advice.

A crawler policy, as a suggested starter worksheetYou can copy this into a doc or spreadsheet. Fill one row per crawler or group, record why, and set a review date.

Crawler / group Decision Condition or rate limit Why (rationale) Review date
Discovery / searchGooglebot · bingbot AllowConditionsLicenseBlock
AI trainingGPTBot · ClaudeBot · CCBot AllowConditionsLicenseBlock
Live AI agentsChatGPT-User · Claude-User AllowConditionsLicenseBlock
OperationalAdsBot · verification · previews AllowConditionsLicenseBlock
Unknown / unverifiedspoofed UAs · proxies AllowConditionsLicenseBlock

Mature policies also vary by page type with public product and help pages usually more open than premium editorial, ad-heavy articles, or account/checkout paths. Add rows per sensitive area as needed. Not implementation advice: test changes on a small set of URLs first, and re-verify tokens against operator documentation.

Suggested Key Considerations

Declarerobots.txt, Content Signals and licence directives state your preferences. Not necessary, but advisory.
VerifyWeb Bot Auth (IETF) cryptographically signs bot requests. Has been implemented by Cloudflare & supported by OpenAI, AWS and others.
EnforceCDN, WAF and bot management apply your policy at the network edge, where it cannot be ignored.
TransactIAB Tech Lab’s CoMP v1 (finalised 28 April 2026) has APIs for licensed AI systems to request content and declare usage.

IAB Tech Lab’s CoMP is very deliberately scoped: it is not a blocking system, not a licensing marketplace, and not the economic model. It enables the communication layer that assumes you already have an access-control strategy. The IAB Tech Lab’s companion Bot & Crawler Management Guidance (2026) supplies that strategy framework, and this guide draws on it throughout.

The emerging third option: licensed accessThe old binary of allow-or-block is becoming allow / license / block. Licensing isn’t a new crawler type, it’s a new relationship that can apply to any of the five categories. Once a commercial agreement exists, CoMP is the standard way for a licensed AI system and a content owner to communicate. This is what turns a passive block into a potential revenue line.

Some Resources

Version 1.0 · 22 July 2026 · Prepared by Jonas Jaanimagi, Technology Lead, IAB Australia. User-agent tokens verified against operator documentation and Cloudflare’s AI Crawl Control reference as at publication; tokens change frequently so please re-verify before implementing rules. Crawler categorisations reflect operator declarations and third-party classifications and may evolve. This document is general information for IAB Australia members and event attendees; it is not legal, commercial or technical advice for any specific implementation.

Recommended

>