It’s not who the crawler is, it’s what happens to your content after it leaves. Five jobs, five different answers.
The intention of this guidance is to support content owners of all types to gain deeper awareness around crawlers and bots, and then consider establishing a policy and approach to optimise accordingly based on required outcomes. The aim is never to simply block every crawler, rather it’s to be able to tell beneficial, commercial, operational and unknown non-human traffic apart, so that each can be managed on its merits. Volume isn’t always true value as AI sources will often send less traffic than classic search, but that traffic is often higher-intent and converts better (as per the agentic-traffic panel below).
| Crawler job | What you gain if allowed | What you risk if allowed | What blocking costs you | Publisher trend | Brand trend |
|---|---|---|---|---|---|
| ADiscovery & searchGooglebot · bingbot · OAI-SearchBot | Referral traffic, citations in AI answers, brand visibility | The same crawl can also feed AI answer features built on the index | De-indexing; loss of organic search traffic | Allow | Allow |
| BAI trainingGPTBot · ClaudeBot · CCBot | Presence and familiarity inside future models; goodwill with AI operators | Uncompensated extraction; forgone licensing value; irreversible once crawled | Not learned by future models results in invisible loss & no clear signal | Block, or license | Selective (product facts in, premium IP out) |
| CLive AI agentsChatGPT-User · Claude-User · Google-Agent | A waiting human gets served; task completion; live-answer citations | No pageview or ad impression; paywall probing | Invisible at the moment of decision for assistant users | Conditional; restrict on ad-funded pages | Allow on product & knowledge pages |
| DOperationalAdsBot-Google · verification · previews | Ad delivery, verification, measurement, uptime, link previews | Negligible; some competitive visibility via SEO tools | Broken ad stack; programmatic revenue collapse; ugly link shares | Always allow (verified) | Always allow (verified) |
| EUnknown / unverifiedspoofed UAs · python-requests · proxies | Nothing knowable | Scraping, fraud, reconnaissance everything | Nothing | Challenge or block; investigate volume | Challenge or block |
Decide per crawler group, then per crawler within high-risk groups. Document each decision with a rationale and a review date. (Framework: IAB Tech Lab Bot & Crawler Management Guidance.)
Sources: IAB Tech Lab Bot & Crawler Management Guidance (2026) and CoMP v1 (finalised 28 April 2026); Cloudflare Content Independence Day announcements (1 July 2026). Verify all user-agent tokens against the IAB/ABC International Spiders & Bots List and operator documentation before implementation. v1.0 · 22 July 2026. General information, not legal advice.
Managing Crawlers & AI Bots
A starter guide for content owners: who’s at your door, what they do with your content, and what you control.
Three shifts have altered the traditional crawl-for-clicks bargain. Search has become answeringwith results pages increasingly answering directly, making the click optional. Crawling has became multi-purpose, the same page fetch can now feed a search index, AI answers, model training, live agent tasks, verification and analytics. And control moved to the network so whilst robots.txt remains advisory, CDN and WAF enforcement is not. In mid-2026, automated requests surpassed human ones in web-page traffic for the first time, and the majority of identified crawler activity now relates to AI. The practical response is not panic or a blanket block: it is knowing the five jobs bots do on your site, and deciding each on its merits.
How to use this guide. Tokens below are user-agent identifiers you will see in server logs, robots.txt and bot-management tools, verified against operator documentation and Cloudflare’s AI Crawl Control reference as at July 2026. Tokens will change frequently, so regularly re-verify against the IAB/ABC International Spiders & Bots List and each operator’s documentation before implementing rules.
Start here (by role)Three proposed starting positions, with full reasoning in the matrix above and the playbook further down.
Allow discovery and live agents. Review training as a separate decision with product facts in and sensitive IP out. Measure AI referrals.
Allow discovery. Audit training and agent exposure before blocking. Publish a crawler policy with a licensing contact.
Visibility → verification → policy. Verify identity by network, not by name. Enforce at the edge for high-risk traffic.
ADiscovery & search indexersBuild an index so you can be found later. The librarian: reads your work, catalogues it, sends readers back. The category most likely to return measurable referral traffic. Common examples are below.
| User-agent token | Operator | What it does | Notes & stance |
|---|---|---|---|
| Googlebot | Core web-search index; the same crawl also feeds AI answer features | Multi-purpose (per Cloudflare’s classification); blocking means de-indexing | |
| Googlebot-Image / -News / -Video · Storebot-Google | Vertical indexing: images, news, video, shopping | Product-specific tokens under the Googlebot family | |
| GoogleOther | Generic crawler for R&D and other products | Purpose-ambiguous by design — worth watching in logs | |
| bingbot | Microsoft | Bing search index; grounds Copilot answers | Multi-purpose (per Cloudflare’s classification) |
| Applebot | Apple | Siri, Spotlight and Safari suggestions; Apple Intelligence features | Multi-purpose (per Cloudflare’s classification) |
| Amazonbot | Amazon | Alexa answers and discovery | AI use noted in Amazon’s documentation |
| OAI-SearchBot | OpenAI | ChatGPT search index | Declared search-only; separate from GPTBot (training) |
| Claude-SearchBot | Anthropic | Claude search index | Split out from training in 2025 |
| PerplexityBot | Perplexity | Answer-engine index | Powers Perplexity’s cited answers |
| YouBot | You.com | AI search index | Early Pay-Per-Use compensation partner |
| DuckDuckBot · MojeekBot · PetalBot | DuckDuckGo · Mojeek · Huawei | Independent and regional indexes | Mojeek runs a fully independent UK index |
| YandexBot · Baiduspider · Yeti | Yandex · Baidu · Naver | Regional majors (RU, CN, KR) | Relevant for international audiences |
BAI training crawlersAbsorb your content into a model. The student who memorises your book and never brings it back. Nothing is returned by default, and once taken it can’t reliably be recalled from copies, corpora or trained models. Common examples are below.
| User-agent token | Operator | What it does | Notes & stance |
|---|---|---|---|
| GPTBot | OpenAI | Model training | Widely blocked by site owners |
| ClaudeBot | Anthropic | Model training | Distinct from Claude-SearchBot and Claude-User |
| Meta-ExternalAgent | Meta | Training and model improvement | Covers Llama and Meta AI products |
| CCBot | Common Crawl | Non-profit corpus feeding many downstream trainers | Blocking exits many datasets at once — including open-source models |
| Bytespider | ByteDance | Model training | Compliance practices vary — use verified identity and edge enforcement for high-risk traffic |
| cohere-training-data-crawler | Cohere | Enterprise model training | — |
| AI2Bot | Allen Institute for AI | Research corpora (e.g. Dolma) | Academic; a values call for some owners |
| omgilibot / Webzio-Extended | Webz.io | Data broker supplying AI buyers | Downstream use opaque |
| Diffbot | Diffbot | Structured extraction and dataset sales | — |
| Google-CloudVertexBot | Google Cloud | Crawls sites at Vertex AI customers’ request | Site-owner-initiated; no effect on Google Search |
CLive agents & user fetchersFetch a page right now because a human asked. The personal shopper: a real person is served, but there’s no pageview and no ad impression. Common examples are below.
| User-agent token | Operator | What it does | Notes & stance |
|---|---|---|---|
| ChatGPT-User | OpenAI | Fetches when a ChatGPT user asks | User-initiated; OpenAI’s docs (updated Dec 2025) state robots.txt may not apply — enforce at the edge if needed |
| Claude-User | Anthropic | Fetches when a Claude user asks | Anthropic states all its bots honour robots.txt, incl. Crawl-delay |
| Perplexity-User | Perplexity | User-initiated fetches | Docs state user fetches may not honour robots directives |
| Meta-ExternalFetcher | Meta | Assistant link fetches | — |
| MistralAI-User | Mistral | Le Chat fetches | — |
| DuckAssistBot | DuckDuckGo | AI answer fetches | — |
| Google-Agent | User-directed browsing on Google infrastructure (Gemini Agent, Chrome Auto Browse) | Added early 2026 with its own IP range file; verify current docs, ranges and robots behaviour before writing a rule | |
| Google-NotebookLM | Fetches user-added sources for notebooks | Being renamed after the Gemini Notebook rebrand so the old token will deprecate August 2026; please check current docs | |
| Browser-use agents (various) | Various | Operate a real browser on a user’s behalf | Present as near-human traffic |
The shape of agentic traffic (mid-2026)Agents are the fastest-moving force on the web and the gap between what they take and what they return has never been wider.
The defining feature of agent traffic is asymmetry. A person shopping for shoes might visit five sites; an agent completing the same task can visit hundredsand automated traffic is now expanding several times faster than human activity (HUMAN Security, 2026). That volume doesn’t map to value: crawl activity and referral value are moving in opposite directions. ChatGPT crawled around 6% less quarter on quarter yet sent roughly 17% more visitors, while Meta’s two agents generated more than half of all AI traffic on one major network and returned almost none. The lesson mirrors the matrix, so judge a crawler by what comes back, not by how often it knocks.
Composition shifts constantly, which is why the playbook says review quarterly. Operators are splitting their do-everything crawlers into clearer single-purpose ones. Mixed-purpose crawling fell from roughly 49% to 33% of AI requests across the first half of 2026, and the rankings reorder month to month (Meta’s indexing crawler overtook its training crawler in June). Yesterday’s allowlist is not today’s.
Figures are drawn from each provider’s own network and measurement window, so treat them as directional rather than precise, but the direction is clear. Sources: DataDome AI Traffic Report Q2 2026; Cloudflare Radar (mid-2026); HUMAN Security 2026 State of AI Traffic; Gartner; Adobe Digital Insights. Please verify against primary sources regularly.
DOperational & advertising infrastructureKeep the money and machinery working. The meter readers: rarely glamorous, never block the verified ones.
| User-agent token | Operator | What it does | Notes & stance |
|---|---|---|---|
| AdsBot-Google / AdsBot-Google-Mobile | Ad landing-page quality checks | Blocking harms Google Ads delivery and quality scoring | |
| Mediapartners-Google | AdSense contextual crawling | Required for AdSense targeting | |
| AdIdxBot | Microsoft | Microsoft Advertising quality and verification | Bing Ads equivalent of AdsBot |
| GoogleInspectionTool | Search Console testing (URL Inspection, Rich Results) | Commonly seen in logs and wrongly blocked; leave allowed for your own diagnostics | |
| proximic · peer39_crawler | Comscore · Peer39 | Contextual classification for targeting | Blocking reduces contextual demand |
| Verification fetchers (DV, IAS, HUMAN) | Various | Viewability, fraud and brand-safety verification | Tokens vary, so consult the IAB/ABC Spiders & Bots List if possible; blocking breaks verification and ads.txt-dependent revenue |
| AhrefsBot · SemrushBot · Screaming Frog · MJ12bot · DotBot | SEO vendors | Site audit, backlinks, rankings | Gray zone: useful to you and to competitors; rate-limit rather than block outright |
| facebookexternalhit · Twitterbot · LinkedInBot · Slackbot · Discordbot · WhatsApp · TelegramBot · Pinterestbot | Platforms | Link previews when your URLs are shared | Blocking makes your shares ugly |
| UptimeRobot · Pingdom | Monitoring vendors | Uptime checks | Low volume; often run by you or your host |
| Feedfetcher-Google · Feedly · podcast fetchers | Various | Feed and podcast distribution | Distribution you probably want |
| archive.org_bot / ia_archiver · TurnitinBot | Internet Archive · Turnitin | Web archival; academic integrity | A values call as much as a commercial one |
EImpostors & unknownsThe gatecrashers. Treat unidentifiable traffic as information, not noise.
A meaningful share of bot traffic matches no known signature: spoofed major-crawler user agents from the wrong networks, default library strings (python-requests, curl, Go-http-client, Scrapy), headless browsers, and residential-proxy fleets. Some operators’ crawler identities also remain poorly documented publicly (xAI/Grok and DeepSeek are current examples) which is the transparency problem in miniature: treat as unknown until verified. High-volume unidentified traffic deserves investigation, not a shrug.
*.googlebot.com), then forward-resolve that hostname and check it returns the original IP. Most major operators publish IP-range files or a DNS process for exactly this. The three-step check: 1) match the UA against a maintained directory; 2) confirm the source IP against the operator’s published ranges or forward-confirmed reverse DNS (an ASN such as Google’s AS15169 is a supporting signal, not proof on its own); 3) confirm behaviour fits the declared purpose (crawl rate, paths, robots.txt compliance). A claimed major-crawler UA arriving from a residential ISP or generic cloud range is almost certainly an impostor.The playbook: five moves, in order
- Get visibility first. The problem isn’t crawling, it’s invisibility. Audit at least 30 days of edge or server logs before changing anything; owners who block blind often block crawlers that drive real referrals. Most CDNs and WAFs can surface bot traffic by user-agent, ASN and verified status. Request a regular 30-day report from your vendor(s) if you don’t already receive one.
- Evaluate each crawler on value. Traffic return, content use, bandwidth cost, reputation, strategic value. Try to score them accordingly and the decision usually forms naturally.
- Apply the four options per group: Allow · Allow with conditions · Require licensing · Block. Within high-risk groups, decide per crawler.
- Document and review. Record decision, rationale, conditions and a review date. Check for new high-volume unknowns weekly; review bandwidth monthly; re-run the full assessment quarterly.
- Publish your policy. A human-readable crawler policy page, a machine-readable robots.txt, and a licensing contact. A visible policy turns a passive block into a potential commercial relationship.
A crawler policy, as a suggested starter worksheetYou can copy this into a doc or spreadsheet. Fill one row per crawler or group, record why, and set a review date.
| Crawler / group | Decision | Condition or rate limit | Why (rationale) | Review date |
|---|---|---|---|---|
| Discovery / searchGooglebot · bingbot | AllowConditionsLicenseBlock | … | … | … |
| AI trainingGPTBot · ClaudeBot · CCBot | AllowConditionsLicenseBlock | … | … | … |
| Live AI agentsChatGPT-User · Claude-User | AllowConditionsLicenseBlock | … | … | … |
| OperationalAdsBot · verification · previews | AllowConditionsLicenseBlock | … | … | … |
| Unknown / unverifiedspoofed UAs · proxies | AllowConditionsLicenseBlock | … | … | … |
Mature policies also vary by page type with public product and help pages usually more open than premium editorial, ad-heavy articles, or account/checkout paths. Add rows per sensitive area as needed. Not implementation advice: test changes on a small set of URLs first, and re-verify tokens against operator documentation.
Suggested Key Considerations
IAB Tech Lab’s CoMP is very deliberately scoped: it is not a blocking system, not a licensing marketplace, and not the economic model. It enables the communication layer that assumes you already have an access-control strategy. The IAB Tech Lab’s companion Bot & Crawler Management Guidance (2026) supplies that strategy framework, and this guide draws on it throughout.
Some Resources
- IAB/ABC International Spiders & Bots List: the maintained directory for advertising-relevant bots, and the authoritative reference for verification-vendor tokens.
- IAB Australia AI Hub: local AI education, guidance and resources.
- IAB Tech Lab CoMP v1 and Bot & Crawler Management Guidance.
- Operator crawler documentation: Google, Microsoft, OpenAI, Anthropic, Apple, Meta and Perplexity each publish token lists, IP ranges and verification methods.
- Cloudflare AI Crawl Control and Radar bots directory: live classification and traffic data for known crawlers.
Version 1.0 · 22 July 2026 · Prepared by Jonas Jaanimagi, Technology Lead, IAB Australia. User-agent tokens verified against operator documentation and Cloudflare’s AI Crawl Control reference as at publication; tokens change frequently so please re-verify before implementing rules. Crawler categorisations reflect operator declarations and third-party classifications and may evolve. This document is general information for IAB Australia members and event attendees; it is not legal, commercial or technical advice for any specific implementation.