Key ecosystem metrics across 5,000+ websites using Agent Analytics and AI Chat Referral Tracking.
↓ 3%
Compared to the previous 90 days
The amount of visits from known agents vs. humans
↑ 7%
Compared to the previous 90 days
The percentage of bot traffic that's AI-related
↑ 7%
Compared to the previous 90 days
AI Agent
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
AI Assistant
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
AI Coding Agent
AI Coding Agent
Fetches documentation and other resources to help build software
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
AI Search Crawler
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
Archiver
Archiver
Captures and stores historical website snapshots for long-term digital preservation
Automated Agent
Automated Agent
Automates browser interactions programmatically without direct human supervision
Developer Helper
Developer Helper
Assists with testing, debugging, and ensuring website functionality
Fetcher
Fetcher
Retrieves web page metadata to power app features like link previews or feeds
Intelligence Gatherer
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
Scraper
Scraper
Extracts large amounts of web data, often without explicit website permission
Search Engine Crawler
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
Security Scanner
Security Scanner
Scans websites for security vulnerabilities, threats, and configuration weaknesses
SEO Crawler
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
Uncategorized
Uncategorized
Not yet assigned a type
Undocumented AI Agent
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case
Hover over each agent type for more information about what they do
Agent types with the most activity
bingbot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
8.7%
AhrefsBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
6.9%
Googlebot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
6.5%
Known Agent
DEV
Developer Helper
Assists with testing, debugging, and ensuring website functionality
6.1%
ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
3.4%
PetalBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
3.3%
SemrushBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
2.9%
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
2.7%
facebookexternalhit
FTCH
Fetcher
Retrieves web page metadata to power app features like link previews or feeds
2.7%
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
2.4%
meta-externalagent
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
2.4%
Baiduspider
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
2.3%
Amzn-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
2.3%
DotBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
2.1%
MJ12bot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
2.1%
Agents with the most activity
Operators with the most activity
These bots scrape website content to train AI models. Some belong to AI companies, while others belong to third-party services that resell the data. Automatic Robots.txt can block unwanted scraping. Included agent types include AI Data Providers and AI Data Scrapers.
AI Data Provider
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
AI Data Scraper
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
AI scraping activity by agent type over time
Computers and Electronics
6.0%
Business and Industrial
5.6%
Internet and Telecom
5.3%
Website categories with most activity
ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
26.8%
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
18.7%
meta-externalagent
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
18.6%
GPTBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
7.8%
Bytespider
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
7.0%
ShapBot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
6.6%
GoogleOther
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
2.8%
YouBot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
2.4%
CCBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
2.3%
Timpibot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
1.6%
Reflectionbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
1.3%
Diffbot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
1.0%
DeepSeekBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
1.0%
VelenPublicWebCrawler
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
0.8%
FacebookBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
0.5%
Agents doing the most AI scraping
Operators doing the most AI scraping
These bots fetch website content in real time to power AI assistants, coding agents, and other retrieval-augmented generation (RAG) tasks. Pages inform responses on the spot, such as when an assistant summarizes an article or a coding agent references documentation. Included agent types include AI Assistants and AI Coding Agents.
AI Assistant
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
AI Coding Agent
AI Coding Agent
Fetches documentation and other resources to help build software
AI fetching activity by agent type over time
Computers and Electronics
2.3%
Travel and Transportation
1.9%
Internet and Telecom
1.8%
Business and Industrial
1.7%
Website categories with most activity
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
88.7%
DuckAssistBot
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
4.3%
Claude-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
2.8%
Claude-Code
CODE
AI Coding Agent
Fetches documentation and other resources to help build software
1.7%
Shap-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
1.0%
Cursor
CODE
AI Coding Agent
Fetches documentation and other resources to help build software
0.5%
Google-NotebookLM
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.4%
Perplexity-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.2%
opencode
CODE
AI Coding Agent
Fetches documentation and other resources to help build software
0.2%
MistralAI-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
GoogleAgent-URLContext
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
Gemini-Deep-Research
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
Code
CODE
AI Coding Agent
Fetches documentation and other resources to help build software
0.0%
meta-externalfetcher
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
amazon-QBusiness
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
Agents doing the most AI fetching
Operators doing the most AI fetching
These bots crawl website content so it can be surfaced in AI search engines and AI-generated answers. Those answers often include citations or links back to the source pages. Included agent types include AI Search Crawlers.
AI Search Crawler
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
AI search indexing activity by agent type over time
Travel and Transportation
7.0%
Books and Literature
5.1%
Arts and Entertainment
4.8%
Computers and Electronics
4.6%
Website categories with most activity
PetalBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
26.7%
Amzn-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
18.6%
Applebot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
14.1%
Claude-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
12.2%
meta-webindexer
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
11.6%
OAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
10.3%
LinkupBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
2.9%
PerplexityBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
2.2%
Google-CloudVertexBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.3%
AzureAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.3%
xAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.3%
ExaSearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.3%
AIWebIndex
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.2%
AddSearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.0%
Anomura
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.0%
Agents doing the most AI search indexing
Operators doing the most AI search indexing
These bots use browsers to autonomously navigate websites, click through pages, and make decisions to complete tasks for people. Agentic UX best practices and Google PageSpeed Insights help evaluate how well websites support them. Included agent types include AI Agents.
↑ 30%
Compared to the previous 90 days
The average duration of a session
↑ 70%
Compared to the previous 90 days
The average number of pages visited per session
AI Agent
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
AI browsing activity by agent type over time
Computers and Electronics
0.1%
Business and Industrial
0.0%
Arts and Entertainment
0.0%
Internet and Telecom
0.0%
Books and Literature
0.0%
Travel and Transportation
0.0%
Website categories with most activity
Manus-User
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
37.4%
Google-Agent
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
34.2%
ChatGPT Agent
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
28.5%
NovaAct
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
GoogleAgent-Mariner
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
AmazonBuyForMe
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
TwinAgent
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
Agents doing the most AI browsing
Operators doing the most AI browsing
See which robots.txt rules are set across the web and how well agents follow them. An agent's Robots.txt Effectiveness measures the effectiveness of a disallow rule for it by estimating how much the agent reduces its traffic after it's blocked.
How often
all agents across all agent types follow robots.txt rules
LinkupBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
100.0%
YouBot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
100.0%
serpstatbot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
100.0%
bingbot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
100.0%
Diffbot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
99.9%
Dataprovider.com
SCRP
Scraper
Extracts large amounts of web data, often without explicit website permission
99.9%
AhrefsSiteAudit
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
99.8%
SiteAuditBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
99.7%
Barkrowler
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
99.6%
Googlebot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
99.6%
SERankingBacklinksBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
99.6%
meta-externalads
INT
Intelligence Gatherer
Analyzes web content for brand safety, competitive insights, and ad targeting
99.6%
PerplexityBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
99.5%
Amzn-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
99.4%
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
99.3%
Agents with the best Robots.txt Effectiveness percentages
Baiduspider
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
61.4%
ShapBot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
73.7%
YandexBot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
80.7%
facebookexternalhit
FTCH
Fetcher
Retrieves web page metadata to power app features like link previews or feeds
83.1%
Reflectionbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
91.3%
OAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
92.4%
DeepSeekBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
93.4%
Sogou web spider
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
93.7%
SemrushBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
93.9%
DotBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
94.0%
Applebot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
94.4%
PetalBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
95.0%
Bytespider
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
95.5%
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
95.8%
DataForSeoBot
SEO
SEO Crawler
Analyzes website structure and content to identify SEO improvement opportunities
95.8%
Agents with the worst Robots.txt Effectiveness percentages
●
CCBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
24.8%
●
Bytespider
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
22.7%
●
GPTBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
22.0%
●
ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
21.7%
●
meta-externalagent
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
19.7%
●
omgili
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
19.4%
●
Diffbot
PVDR
AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service
19.1%
●
anthropic-ai
UND
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case
18.9%
●
cohere-ai
UND
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case
18.3%
●
PerplexityBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
17.7%
●
Claude-Web
UND
Undocumented AI Agent
Crawls websites without disclosing its purpose, collecting data for an unknown AI use case
17.4%
●
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
17.1%
●
FacebookBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
16.6%
Agents blocked by the most top websites
See which agents are most frequently impersonated, and how spoofing activity changes over time. A visit is considered spoofed when it claims a recognized agent identity but fails that agent's supported authentication method, such as verified IP or Web Bot Auth.
Active Threat: AI Bot Spoofing Campaign
We are observing a widespread campaign impersonating AI bots to scan websites for vulnerabilities. The attacker appears to be targeting credential and configuration paths used by AI coding tools.
Contact us for more information, or
inspect your own traffic.
The percentage of impersonated website traffic for each agent identity over time
●
Googlebot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
0.5%
●
ChatGPT-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.2%
●
OAI-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.1%
●
PerplexityBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.1%
●
GPTBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
0.1%
●
ClaudeBot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
0.1%
●
Applebot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.1%
●
Amzn-SearchBot
SRCH
AI Search Crawler
Indexes website content to possibly include as citations in AI-powered search results
0.1%
●
Perplexity-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.1%
●
MistralAI-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
●
Claude-User
ASST
AI Assistant
Fetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
●
bingbot
SRCH
Search Engine Crawler
Systematically scans and indexes web pages to include in search results
0.0%
●
GoogleOther
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
0.0%
●
Amazonbot
SCRP
AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs
0.0%
●
Google-Agent
AGNT
AI Agent
Uses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
The most impersonated agent identities
/.config/anthropic/credentials/default.json
/.claude/settings.json
/.claude.json
/.hermes/.env
/.openclaw/.env
/.codex/config.toml
/.continue/config.json
/.aider.conf.yml
/service-account.json
/serviceaccountkey.json
/service_account.json
/firebase-adminsdk.json
/firebase-service-account.json
/.aws/credentials
/.aws/config
/.s3cfg
/.boto
/.npmrc
/.env.example
/.env.local
/.env.production
/.env.backup
/.env.old
/backend/.env
/api/.env
/admin/.env
/dockerfile
/docker-compose.yaml
/.docker/config.json
/terraform.tfstate
/credentials.json
/secrets.json
/secrets.yml
/key.json
/rclone.conf
Examples of recent top targeted request paths
See which AI platforms like ChatGPT, Perplexity, and Gemini cite websites and send them human referral traffic. Citations are estimated. Google's guide explains how websites can optimize their content to be more visible in AI chat responses (GEO).
Computers and Electronics
2.3%
Travel and Transportation
1.9%
Internet and Telecom
1.8%
Business and Industrial
1.7%
Website categories most frequently cited in AI chat responses
Travel and Transportation
0.3%
Computers and Electronics
0.2%
Business and Industrial
0.1%
Internet and Telecom
0.0%
Website categories receiving the most referrals from AI chat
The Index updates daily using completed days of traffic, security, and referral data from more than 5,000 websites using Agent Analytics and AI Chat Referral Tracking. Percentage changes compare equal-length periods. Agent data comes from the Agent Directory, and website categories follow Google AdSense. Participating websites are not a random sample, and the qualifying set changes over time, so results show observed directional trends rather than a census of global web traffic.
Only website-days meeting minimum activity and data-quality requirements are included; internal, test, incomplete, and anomalous data is excluded. Bot percentages use all visits, while AI chat referral percentages use estimated human visits after identified bots are removed. Rates are calculated per website-day and then averaged, giving each website equal weight regardless of traffic. Daily charts are not smoothed.
Robots.txt Effectiveness estimates the traffic reduction associated with blocking an agent sitewide using Disallow: /. For each agent and day, the baseline allowed rate includes every qualifying website that allows it, even those with zero agent visits:
For each blocked website, expected visits are:
Given observed disallowed visits O, its score is:
A 0% score means visits met or exceeded expectations, 50% means half as many visits as expected, and 100% means none were observed.
Scores require meaningful activity across multiple allowed and blocked websites over at least 7 of the latest 14 completed days. Each qualifying website-day has equal weight, and the headline score gives each qualifying agent equal weight. Only verified visits count for agents supporting authentication. This observational score measures outcomes, not intent or causation.
Top Blocked Bots is calculated separately from robots.txt scans of Similarweb's top 1,000 websites.
Spoofing statistics count visits that claim a known agent's identity but fail a supported authentication method, such as IP verification or HTTP message signatures. Each agent's daily rate is calculated against total visits per website and then averaged across websites. A failure suggests impersonation but does not identify the actual software or operator. Agents without supported authentication are excluded.
AI chat referrals count observed human visits with a recognized AI platform in the referrer or campaign source; visits without that data cannot be attributed. Citation estimates use requests from agents that retrieve content for AI platforms. Those requests may inform a response but do not prove a user saw a citation, so citation results are directional rather than exact counts.
Can journalists and media organizations use this data?
Absolutely. You may cite The Agentic Web Index with attribution and a link to this page. For interviews, fact-checking, background context, or a more specific breakdown for a story, contact us and include your deadline.
Do you work with researchers?
Absolutely. We welcome thoughtful research into how agents and bots are changing the web. Tell us about your research question, timeframe, and intended use. Depending on the scope and data constraints, we may be able to provide additional context, compare approaches, or explore a joint analysis.
Can I request a specific analysis?
Yes. If you need a breakdown by agent, operator, activity type, website category, or time period that is not shown here, contact us. When the underlying data supports it, we can examine the question and provide a focused analysis.
How do I see these trends on my own website?