Hacker News — AI on Front Page · · 23 min read

Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot

Mirrored from Hacker News — AI on Front Page for archival readability. Support the source by reading on the original site.

284 pts · 215 comments on Hacker News

The Agentic Web Index

The internet is rapidly evolving from an environment built primarily for humans, into one increasingly used by machines. See how AI agents, crawlers, scrapers, and other bots are reshaping the way information is discovered, accessed, and used across the web.

Overview

Key ecosystem metrics across 5,000+ websites using Agent Analytics and AI Chat Referral Tracking.

Bot vs. Human Traffic

35%
↓ 1%
Compared to the previous 90 days
The amount of visits from bots vs. humans

Agentrification

29%
↑ 11%
Compared to the previous 90 days
The percentage of bot traffic that's AI-related

AI Chat Referral Volume

0.1%
↓ 9%
Compared to the previous 90 days
The percentage of human website visits that come from AI chat
See Referral Trends ↓

Robots.txt Effectiveness

98.5%
How often bots follow robots.txt rules
See Per Agent ↓

Traffic by Agent Type

AI Agent
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
AI Assistant
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
AI Coding Agent
AI Coding Agent
Fetches documentation and other resources to help build software
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
AI Search Crawler
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Archiver
Archiver
Captures and stores historical website snapshots for long-term digital preservation
Automated Agent
Automated Agent
Automates browser interactions programmatically without direct human supervision
Developer Helper
Developer Helper
Assists with testing, debugging, and ensuring website functionality
Fetcher
Fetcher
Retrieves web page metadata to power app features like link previews or feeds
Intelligence Gatherer
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
Scraper
Scraper
Extracts large amounts of web data, often without explicit website permission
Search Engine Crawler
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
Security Scanner
Security Scanner
Scans websites for security vulnerabilities, threats, and configuration weaknesses
SEO Crawler
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Uncategorized
Uncategorized
Not yet assigned a type
Undocumented AI Agent
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case

Hover over each agent type for more information about what they do

Top Agent Types

Agent types with the most activity

Top Visiting Agents

bingbot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
Googlebot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
AhrefsBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Known Agent
DEV
Developer Helper
Assists with testing, debugging, and ensuring website functionality
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
PetalBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
SemrushBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
facebookexternalhit
FTCH
Fetcher
Retrieves web page metadata to power app features like link previews or feeds
meta-externalagent
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Amzn-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
MJ12bot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
DotBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Applebot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results

Agents with the most activity

Top Operators

Operators with the most activity

AI Scraping Activity

These bots scrape website content to train AI models. Some belong to AI companies, while others belong to third-party services that resell the data. Automatic Robots.txt can block unwanted scraping. Included agent types include AI Data Providers and AI Data Scrapers.

AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs

AI scraping activity by agent type over time

Top Visited Website Categories

Law and Government
Online Communities
Home and Garden
Computers and Electronics
Shopping
Jobs and Education
Pets and Animals
Business and Industrial
Games
Autos and Vehicles
Hobbies and Leisure
Internet and Telecom
Beauty and Fitness

Website categories with most activity

Top Agents

ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
meta-externalagent
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
GPTBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Bytespider
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
ShapBot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
GoogleOther
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
CCBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Timpibot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
YouBot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
VelenPublicWebCrawler
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Diffbot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
DeepSeekBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
AIWebIndex
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
TerraCotta
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service

Agents doing the most AI scraping

Top Operators

Operators doing the most AI scraping

AI Fetching Activity

These bots fetch website content in real time to power AI assistants, coding agents, and other retrieval-augmented generation (RAG) tasks. Pages inform responses on the spot, such as when an assistant summarizes an article or a coding agent references documentation. Included agent types include AI Assistants and AI Coding Agents.

AI Assistant
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
AI Coding Agent
AI Coding Agent
Fetches documentation and other resources to help build software

AI fetching activity by agent type over time

Top Visited Website Categories

Reference
Science
Finance
Law and Government
Computers and Electronics
Beauty and Fitness
Travel and Transportation
Business and Industrial
Internet and Telecom
Health
Autos and Vehicles
Real Estate
Sports

Website categories with most activity

Top Agents

ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
DuckAssistBot
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
Perplexity-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
Claude-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
MistralAI-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
Claude-Code
CODE
AI Coding Agent
Fetches documentation and other resources to help build software
Shap-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
Google-NotebookLM
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
Cursor
CODE
AI Coding Agent
Fetches documentation and other resources to help build software
Gemini-Deep-Research
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
GoogleAgent-URLContext
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
meta-externalfetcher
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
Code
CODE
AI Coding Agent
Fetches documentation and other resources to help build software
kagi-fetcher
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
QualifiedBot
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer

Agents doing the most AI fetching

Top Operators

Operators doing the most AI fetching

AI Search Indexing Activity

These bots crawl website content so it can be surfaced in AI search engines and AI-generated answers. Those answers often include citations or links back to the source pages. Included agent types include AI Search Crawlers.

AI Search Crawler
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results

AI search indexing activity by agent type over time

Top Visited Website Categories

Travel and Transportation
Reference
Health
Games
Food and Drink
Books and Literature
Jobs and Education
Computers and Electronics
Arts and Entertainment
People and Society
Law and Government
Shopping
Hobbies and Leisure

Website categories with most activity

Top Agents

PetalBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Amzn-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Applebot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
OAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
meta-webindexer
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Claude-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
PerplexityBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
LinkupBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Google-CloudVertexBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
xAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
AzureAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
AddSearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Anomura
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
ExaSearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
MistralAI-Index
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results

Agents doing the most AI search indexing

Top Operators

Operators doing the most AI search indexing

AI Browsing Activity

These bots use browsers to autonomously navigate websites, click through pages, and make decisions to complete tasks for people. Agentic UX best practices and Google PageSpeed Insights help evaluate how well websites support them. Included agent types include AI Agents.

Average Session Duration

1 minute
↓ 2%
Compared to the previous 90 days
The average duration of a session

Average Pages per Session

8
↑ 7%
Compared to the previous 90 days
The average number of pages visited per session
AI Agent
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

AI browsing activity by agent type over time

Top Visited Website Categories

Computers and Electronics
Law and Government
Autos and Vehicles
Internet and Telecom
Business and Industrial
Shopping
Travel and Transportation
Arts and Entertainment
Beauty and Fitness
Sports
Books and Literature
Finance
Health

Website categories with most activity

Top Agents

Google-Agent
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
Manus-User
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
ChatGPT Agent
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
NovaAct
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
GoogleAgent-Mariner
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
AmazonBuyForMe
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
TwinAgent
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user

Agents doing the most AI browsing

Top Operators

Operators doing the most AI browsing

Robots.txt & Compliance

See which robots.txt rules are set across the web and how well agents follow them. An agent's Robots.txt Effectiveness measures the effectiveness of a disallow rule for it by estimating how much the agent reduces its traffic after it's blocked.

Effectiveness (Overall)

98.5%
How often all agents across all agent types follow robots.txt rules

Effectiveness (AI Scrapers & Data Providers)

97.4%
How often AI Data Scrapers and AI Data Providers follow robots.txt rules

Top Rule-Following Agents

LinkupBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
serpstatbot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
proximic
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
um-IC
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
ClarityBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Barkrowler
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Dragonfly
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
AffsignalCrawler
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
meta-externalads
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
SEOkicks
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
UptimeRobot
DEV
Developer Helper
Assists with testing, debugging, and ensuring website functionality
Leikibot
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
TTD-Content
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
Sucuri Uptime Monitor
UNC
Uncategorized
Not yet assigned a type
heritrix
ARCH
Archiver
Captures and stores historical website snapshots for long-term digital preservation

Agents with the best Robots.txt Effectiveness percentages

Top Rule-Breaking Agents

Baiduspider
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
SirdataBot
UNC
Uncategorized
Not yet assigned a type
YandexBot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
ShapBot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
Sogou web spider
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
peer39_crawler
UNC
Uncategorized
Not yet assigned a type
Mediapartners-Google
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
Scrapy
SCRP
Scraper
Extracts large amounts of web data, often without explicit website permission
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
HeadlessChrome
AUTO
Automated Agent
Automates browser interactions programmatically without direct human supervision
archive.org_bot
ARCH
Archiver
Captures and stores historical website snapshots for long-term digital preservation
linkfluence
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
DotBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
OAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer

Agents with the worst Robots.txt Effectiveness percentages

Robots.txt Rules by Top Blocked Agent

The percentage of the top 1,000 websites blocking each agent in robots.txt over time

Top Blocked Agents

GPTBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
CCBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Bytespider
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
PerplexityBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
anthropic-ai
UND
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case
meta-externalagent
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Diffbot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
Claude-Web
UND
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case
cohere-ai
UND
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case
omgili
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs

Agents blocked by the most top websites

Spoofing & Security

See which agents are most frequently impersonated, and how spoofing activity changes over time. A visit is considered spoofed when it claims a recognized agent identity but fails that agent's supported authentication method, such as verified IP or Web Bot Auth.

Active Threat: AI Bot Spoofing Campaign
We are observing a widespread campaign impersonating AI bots to scan websites for vulnerabilities. The attacker appears to be targeting credential and configuration paths used by AI coding tools. Contact us for more information, or inspect your own traffic.

Spoofed Traffic by Agent Identity

The percentage of impersonated website traffic for each agent identity over time

Top Spoofed Agent Identities

Googlebot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
OAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
GPTBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
PerplexityBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
Applebot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
bingbot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
Perplexity-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
MistralAI-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
GoogleOther
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
AhrefsBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Claude-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
Claude-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Amzn-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results

The most impersonated agent identities

Recent Top Targeted Paths

/.config/anthropic/credentials/default.json
/.claude/settings.json
/.claude.json
/.hermes/.env
/.openclaw/.env
/.codex/config.toml
/.continue/config.json
/.aider.conf.yml
/service-account.json
/serviceaccountkey.json
/service_account.json
/firebase-adminsdk.json
/firebase-service-account.json
/.aws/credentials
/.aws/config
/.s3cfg
/.boto
/.npmrc
/.env.example
/.env.local
/.env.production
/.env.backup
/.env.old
/backend/.env
/api/.env
/admin/.env
/dockerfile
/docker-compose.yaml
/.docker/config.json
/terraform.tfstate
/credentials.json
/secrets.json
/secrets.yml
/key.json
/rclone.conf

Examples of recent top targeted request paths

AI Chat Referrals

See which AI platforms like ChatGPT, Perplexity, and Gemini cite websites and send them human referral traffic. Citations are estimated. Google's guide explains how websites can optimize their content to be more visible in AI chat responses (GEO).

Traffic by AI Chat Platform

ChatGPT
Claude
Copilot
DeepSeek
Gemini
Meta AI
Mistral
Perplexity

Referral activity by AI chat platform over time

Top Mentioned (Cited) Website Categories

Reference
Science
Finance
Law and Government
Computers and Electronics
Beauty and Fitness
Travel and Transportation
Business and Industrial
Internet and Telecom
Health
Autos and Vehicles
Real Estate
Sports
Games
Jobs and Education

Website categories most frequently cited in AI chat responses

Top Clicked (Referred) Website Categories

Travel and Transportation
Autos and Vehicles
Beauty and Fitness
Finance
Computers and Electronics
Jobs and Education
Business and Industrial
Home and Garden
Sports
Health
Real Estate
Internet and Telecom
Science
Food and Drink
Shopping

Website categories receiving the most referrals from AI chat

Methodology

Data Scope

The Index is updated daily with completed days of traffic, security, and referral data from more than 5,000 websites using Agent Analytics and AI Chat Referral Tracking. The current partial day is excluded. Percentage-change tags compare the current period with the preceding period of the same duration. Agent names, operators, and classifications come from the Agent Directory, which is updated as new agents are discovered or existing agents change. Website categories follow the taxonomy used by Google AdSense. Participating websites are not a random sample of the entire web, and the qualifying set can change as websites connect, disconnect, or cross activity thresholds. Results characterize the observed network and broader directional trends; they should not be interpreted as a precise census of global web traffic.

Qualification & Aggregation

Only websites meeting minimum activity and data-quality requirements are included. Internal, test, incomplete, or anomalous data is excluded. Bot traffic percentages use total server traffic as their denominator. AI chat referral percentages use estimated human traffic, calculated by excluding identified bot visits from total server traffic. Rates are calculated for each qualifying website first, then averaged across websites and completed days. This gives each website equal weight regardless of traffic volume and prevents a small number of high-traffic websites from dominating the results. Daily charts are not smoothed, allowing normal seasonality to remain visible.

Measuring Robots.txt Effectiveness

An agent's Robots.txt Effectiveness estimates the reduction in its request rate associated with a full disallow rule. For each completed day, Known Agents establishes an agent-specific baseline from qualifying websites where that agent is allowed, adjusts the baseline for the overall traffic of each website where the agent is disallowed, and compares the expected activity with the activity actually observed. Only website-day observations with sufficient site traffic, agent activity, cross-site coverage, and expected volume qualify. Scores also require repeated observations across multiple websites and days. When an agent publishes a supported authentication method, only verified traffic is attributed to it.

Each qualifying website-day contributes equally. Scores range from 0%, meaning no measurable reduction, to 100%, meaning no qualifying requests were observed where the agent was disallowed. The headline Robots.txt Effectiveness metric gives each qualifying agent equal weight. Because this is an observational estimate rather than a controlled experiment, it measures an association with robots.txt rules but does not claim that robots.txt caused every observed difference. Top Blocked Bots is calculated separately using daily robots.txt scans of Similarweb's top 1,000 websites.

Identifying Spoofed Bots

Spoofing statistics measure traffic from visits that claim the identity of a known agent but fail a supported authentication method, such as published IP verification or HTTP message signatures. Each agent's daily rate is calculated against total server traffic for every qualifying website, then averaged across websites. A failed check indicates that the visit was likely impersonating the named agent; it does not identify the software or operator that actually made the request. Agents without a supported authentication method are not included in these measurements.

AI Chat Citations & Referrals

AI chat referral statistics count directly observed human visits carrying a recognized AI platform in the referring URL or campaign source. Visits without usable referral information cannot be attributed to an AI platform. Citation statistics are estimates based on requests from agents known to retrieve content for AI platforms. Those requests indicate that content may have informed a response, but they do not confirm that a source appeared as a citation to a user. Because AI platforms do not provide a complete public record of their sources, citation results should be interpreted as directional patterns rather than exact citation counts.

Frequently Asked Questions

Can journalists and media organizations use this data?

Absolutely. You may cite The Agentic Web Index with attribution and a link to this page. For interviews, fact-checking, background context, or a more specific breakdown for a story, contact us and include your deadline.


Do you work with researchers?

Absolutely. We welcome thoughtful research into how agents and bots are changing the web. Tell us about your research question, timeframe, and intended use. Depending on the scope and data constraints, we may be able to provide additional context, compare approaches, or explore a joint analysis.


Can I request a specific analysis?

Yes. If you need a breakdown by agent, operator, activity type, website category, or time period that is not shown here, contact us. When the underlying data supports it, we can examine the question and provide a focused analysis.


How do I see these trends on my own website?

Agent Analytics shows which agents and bots visit your website, what they access, and how their activity changes over time. AI Chat Referral Tracking measures the human traffic arriving from AI chat platforms. Automatic Robots.txt helps manage which bots can access your content.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hacker News — AI on Front Page