18 August 2026 · Methodology v1.0
State of AI Crawler Access 2026
We live-scanned 27 of the world's best-known brands across SaaS, media, retail, travel, and finance. The New York Times and BBC block every one of the 5 major AI crawlers. Media sites score the worst of any category.
Scan your site free Methodology
Average score by category
Block rate by AI bot
| Bot | Owner | Block rate | Blocks it (examples) |
|---|---|---|---|
| GPTBot | OpenAI / ChatGPT | 18.5% | TechCrunch, Figma, Forbes |
| PerplexityBot | Perplexity | 11.1% | Forbes, The New York Times, BBC |
| ClaudeBot | Anthropic / Claude | 18.5% | TechCrunch, Figma, Forbes |
| Google-Extended | Google Gemini / AI Overviews | 14.8% | TechCrunch, Figma, The New York Times |
| Applebot-Extended | Apple Intelligence | 14.8% | TechCrunch, Forbes, The New York Times |
Even the biggest brands have a schema gap
Only 0.0% of the 27 brands have FAQPage schema and only 3.7% have BreadcrumbList — meaning even trillion-dollar companies are leaving basic AI-citation signals on the table.
Full leaderboard
| # | Brand | Category | Score | |
|---|---|---|---|---|
| 1 | HubSpot hubspot.com |
SaaS / Dev Tools | 90 | Scan |
| 2 | Atlassian atlassian.com |
SaaS / Dev Tools | 90 | Scan |
| 3 | Cloudflare cloudflare.com |
SaaS / Dev Tools | 83 | Scan |
| 4 | Zendesk zendesk.com |
SaaS / Dev Tools | 80 | Scan |
| 5 | Stripe stripe.com |
SaaS / Dev Tools | 77 | Scan |
| 6 | Walmart walmart.com |
Retail / E-commerce | 77 | Scan |
| 7 | PayPal paypal.com |
Finance | 77 | Scan |
| 8 | Zoom zoom.us |
SaaS / Dev Tools | 75 | Scan |
| 9 | Target target.com |
Retail / E-commerce | 70 | Scan |
| 10 | Notion notion.so |
SaaS / Dev Tools | 67 | Scan |
| 11 | Slack slack.com |
SaaS / Dev Tools | 67 | Scan |
| 12 | Shopify shopify.com |
Retail / E-commerce | 67 | Scan |
| 13 | Airbnb airbnb.com |
Travel | 67 | Scan |
| 14 | GitHub github.com |
SaaS / Dev Tools | 64 | Scan |
| 15 | Airtable airtable.com |
SaaS / Dev Tools | 63 | Scan |
| 16 | Wired wired.com |
Media / News | 63 | Scan |
| 17 | CNN cnn.com |
Media / News | 61 | Scan |
| 18 | Visa visa.com |
Finance | 60 | Scan |
| 19 | American Express americanexpress.com |
Finance | 60 | Scan |
| 20 | Vercel vercel.com |
SaaS / Dev Tools | 58 | Scan |
| 21 | The Guardian theguardian.com |
Media / News | 55 | Scan |
| 22 | TechCrunch techcrunch.com |
Media / News | 51 | Scan |
| 23 | Booking.com booking.com |
Travel | 51 | Scan |
| 24 | Figma figma.com |
SaaS / Dev Tools | 42 | Scan |
| 25 | Forbes forbes.com |
Media / News | 39 | Scan |
| 26 | The New York Times nytimes.com |
Media / News | 35 | Scan |
| 27 | BBC bbc.com |
Media / News | 32 | Scan |
Bonus: do AI companies practice what they preach?
We scanned 14 leading AI companies' own corporate sites (including OpenAI's competitors). Average score: 69.1/100 — not perfect. Only 57.1% (8/14) publish their own llms.txt. Meta AI and Groq score below 60.
| Company | Score | llms.txt | |
|---|---|---|---|
| ElevenLabs elevenlabs.io |
97 | ✅ | Scan |
| Cohere cohere.com |
78 | ✅ | Scan |
| Suno suno.com |
77 | ✅ | Scan |
| Google DeepMind deepmind.google |
76 | ✅ | Scan |
| Together AI together.ai |
76 | ✅ | Scan |
| Runway runwayml.com |
67 | ✅ | Scan |
| Replicate replicate.com |
67 | ✅ | Scan |
| Stability AI stability.ai |
65 | ❌ | Scan |
| Anthropic anthropic.com |
63 | ❌ | Scan |
| Mistral AI mistral.ai |
63 | ✅ | Scan |
| xAI x.ai |
63 | ❌ | Scan |
| Hugging Face huggingface.co |
63 | ❌ | Scan |
| Meta AI ai.meta.com |
58 | ❌ | Scan |
| Groq groq.com |
54 | ❌ | Scan |
How this report was built
27 well-known global brands (SaaS, media, retail, travel, finance) were scanned live with AICompatible's scanner. Scores combine robots.txt AI-bot access, schema.org structured-data coverage, and llms.txt / sitemap signals. No fabricated numbers — every row is a live scan you can reproduce yourself.
FAQ
Where does this data come from?
27 global brands were scanned live with AICompatible's public scanner. Nothing is fabricated — every score is a reproducible, timestamped live scan.
Why do media sites score so low?
Most news publishers block AI crawlers like GPTBot and ClaudeBot via robots.txt, directly reducing their odds of being cited by AI answer engines.
What is llms.txt and why does it matter?
llms.txt is a structured file that tells AI models what a site is and how to use it. Only 48.1% of the brands scanned publish one.
How do I fix a low score?
Don't block AI bots in robots.txt, add Organization/FAQPage schema, and publish llms.txt. Run the free scanner for a full fix list.
Related: What is GEO? · Turkey Media AI Index · Türkçe versiyon
Note: Scores are automated live scans. Brand names are for identification only.