AI Bots & Web Crawlers Directory
439 bots
AdagioBot
AdagioBot is associated with Adagio's advertising platform and appears to inspect publisher websites for demand optimization.
AddSearchBot
As AddSearch adds content from your site to the search, the AddSearch bot gets counted as traffic by most analytics software.
AddThis
AddThis crawls pages to gather and refresh content used by its website marketing tools.
AdIdxBot
An AdIdxBot visit usually starts with a Bing Ads workflow rather than an organic search crawl.
Adsense
The AdSense crawler visits participating sites in order to provide them with relevant ads.
adsnaver
adsnaver is tied to paid search advertising, not Naver's general web index.
Adyen Webhook
Adyen sends HTTP POST webhooks for events that a payment integration cannot safely infer from a browser return or an immediate API response.
Agency Analytics Crawler
A web crawler by Agency Analytics that allows their clients to check their own sites for SEO.
AGI Agent
The AGI Agent is a productivity assistant that takes actions and makes purchases on behalf of users.
AhrefsBot
Powers the database for both Ahrefs, a marketing intelligence platform, and Yep, an independent, privacy-focused search engine.
AhrefsSiteAudit
Powers Ahrefs’ Site Audit tool. Ahrefs users can use Site Audit to analyze websites and find both technical SEO and on-page SEO issues.
AI Search
Cloudflare AI Search is a managed service that lets you connect your data and easily build AI-powered search.
AI Search External
Cloudflare AI Search is a managed service that lets you connect your data and easily build AI-powered search.
AI2Bot
AI2Bot is operated by the Allen Institute for Artificial Intelligence (Ai2) to crawl the web for content to train open-source AI models.
aiHitBot
aiHitBot collects and maintains historical information about companies.
AirsoftdbBot
Airsoftdb is a search engine for airsoft prod.
Algolia
The Algolia Crawler extracts content from your site and makes it searchable.
All Africa Crawler
All Africa Crawler is part of AllAfrica Global Media's news aggregation and distribution operation.
Alli AI Bot
Alli AI bot crawls customer websites to generate SEO recommendations.
Alphalens Bot
Indexes companies and their offerings to power more effective business discovery.
Amazon AdBot
Amazon AdBot scans pages that request ads from Amazon's advertising systems.
Amazon Bedrock AgentCore Browser
AgentCore Browser provides a secure, cloud-based browser that enables AI agents to interact with websites.
Amazon Bedrock Bot
Amazon Bedrock Bot runs when an Amazon Bedrock customer configures a website as a data source for a knowledge base.
Amazon Kendra
Amazon Kendra is a managed information retrieval and intelligent search service that uses natural language processing and advanced deep learning model.
Amazon Product Discovery
Amazon Product Discovery collects public product details from Amazon selling-partner, brand, and retailer websites.
Amazon Q
Amazon Q Business is a generative artificial intelligence (generative AI)-powered assistant that you can tailor to your business needs.
Amazon Route 53 Health Check Service
Amazon Route 53 Health Check Service
Amazon Seller Initiated Listing
Amazon's web crawler that helps sellers succeed by giving them the option to provide a URL to a website and create high-quality product pages in Amazon's store.
Amazonbot
Amazonbot is Amazon's general web crawler for improving Amazon products and services.
Amzn-SearchBot
Amzn-SearchBot crawls for improving Amazon search experiences (Alexa, Rufus).
Anchor Browser
The Web Browser for AI Agents.
Andibot
Andibot gathers web content for Andi, a conversational AI search assistant that answers questions with summaries and sources.
Apify Website Content Crawler
Crawl websites and extract content to feed AI apps. Convert web data to Markdown or HTML, download files, and more.
APIs-Google
Crawling preferences addressed to the APIs-Google user agent affect the delivery of push notification messages by Google APIs.
Apple App Site Association
Apple App Site Association, usually shortened to AASA, is part of the trust setup behind Universal Links.
Apple Podcasts
Apple Podcasts crawler that only accesses URLs associated with registered content on Apple Podcasts. Does not follow robots.txt.
Applebot
Applebot powers search features in Apple's ecosystem (Spotlight, Siri, Safari) and may be used to train Apple's foundation models for generative AI features.
Applebot-Extended
Applebot-Extended is a robots. txt control token, not a crawler that makes its own page requests. Apple calls it a secondary user agent.
Arena Bot
Link preview bot for Arena.im's live engagement platform. Fetches page previews when URLs are shared in Arena's live chat, live blog, and community tools.
Arquivo Web Crawler
Web crawler archives the Portuguese web.
Artemis Web Crawler
Artemis is a calm web reader with which you can follow websites and blogs.
Artsdata Crawler
Web crawler that collects publicly available LOD for arts and culture in Canada.
Atlassian Jira Webhooks
Delivers webhook notifications from Jira Cloud when issues, projects, or other resources change.
Atlassian Rovo
Crawls and indexes web content for Atlassian Rovo's AI-powered search, chat, and agents.
atlassian-bot
atlassian-bot is a crawler for custom 3P websites that indexes data for rovo search.
Attracta
The Attracta bot is analyzes user website content as part of Attracta's SEO services.
Authory
The Authory bot visits websites to back up articles on behalf of journalists and other writers who use the service.
Awario Bot
Awario's web crawler used to discover and collect new and updated web data for their social media monitoring and brand mention tracking platform.
Awario RSS Bot
One of Awario's primary web crawlers specialized in collecting RSS feed data.
Awario Smart Bot
One of Awario's primary web crawlers that discovers and collects new and updated web data.
Baidu ADS Server Proxy
Baidu ADS Server Proxy is identified in this directory as Baidu's scrubbing proxy.
BaiduSpider
Baiduspider is Baidu’s web crawler that indexes websites for inclusion in its Chinese-market search results.
Barkrowler
Barkrowler is Babbar's web crawler that fuels and updates their graph representation of the web, providing SEO tools for the marketing community.
BestChange Bot
The BestChange bot downloads exchange rate information from 600 websites every 5 seconds.
Better Stack
Better Stack is a platform for monitoring and alerting on your applications.
Bibliothèque nationale de France Crawler
Bibliothèque nationale de France's mission is to collect, catalog, preserve, enrich and communicate the national documentary heritage.
Big Sur AI
Big Sur AI Crawler, crawlers users websites to enable AI-infused experiences.
Bing Preview
BingPreview generates page snapshots for Bing. Note that BingPreview has desktop and mobile variants.
Bingbot
Bingbot is Microsoft's web crawler used for indexing websites for Bing Search.
BLEXBot
SEO PowerSuite Link Explorer (webmeup.com) is the world's freshest backlink index, and the primary source of backlink-related data for the SEO PowerSuite tools.
Bluesky Link Preview Service
Bluesky social pulls links in advance to render webpage previews.
BoardGamePrices Bot
Price comparison site for board games. Need to crawl store pages for participating stores. All stores give permission to be crawled.
BorderxBot
E-commerce product crawler operated by Borderxlab.
Botify
Botify SiteCrawler is part of the Botify Analytics suite.
Brandwatch
The Magpie Crawler indexes content for its soical media monitoring solution.
BraveBot
BraveBot crawls and indexes web pages for the Brave Search index, which also grounds answers from Brave's Leo AI assistant.
Brightbot
Brightbot is Bright Data's crawler layer that monitors the health of websites and enforces ethical web data collection.
Browserbase
Runs headless browser automation on behalf of Browserbase customers for web scraping, form submission, and testing.
Buffer Link Preview Bot
Helps Buffer users create better social media posts by generating rich previews when they share links
Bytespider
Bytespider is ByteDance's web crawler used to gather training data for their AI large language models.
CaliberBot
Conductor uses Caliperbot to fetch websites belonging to clients and prospects.
Capital One Bot
Capital One Bot crawls dealer websites for getting the usage information for Capital One lead navigator button.
CCBot
CCBot is operated by the Common Crawl Foundation to crawl web content for AI training and research.
CensysInspectBot
CensysInspectBot belongs to Censys's internet-wide scanning activity.
Channel3Bot
Crawls product detail pages to index content for AI-powered product discovery, routing shoppers to original websites.
ChatGPT agent
Agent that can use its own browser to perform tasks for user.
ChatGPT-Operator
Handles user-initiated requests from ChatGPT operator accessing external content; not used for automated crawling or AI training.
ChatGPT-User
Handles user-initiated requests in ChatGPT, accessing external content to provide real-time information; not used for automated crawling or AI training.
Chathive crawler
Chathive crawler is the Inteso Group BV crawler that Chathive customers use to bring their own website content into an AI assistant.
Checkly
Checkly is a platform for monitoring and alerting on your applications.
Chrome Lighthouse
PageSpeed Insights (PSI) reports on the user experience of a page on both mobile and desktop devices, and provides suggestions on how that page may be improved.
Chrome Privacy Preserving Prefetch Proxy
Chrome's Privacy Preserving Prefetch Proxy service that fetches /.well-known/traffic-advice to enable privacy-preserving prefetch hints.
CitibotSiteCrawler
CitibotSiteCrawler collects public data from government websites to power Citibot’s AI civic engagement tools.
ClarityBot
ClarityBot is seoClarity's web crawler that performs technical SEO audits, analyzes content, and monitors website performance.
Claude Web
Claude Web is a legacy Anthropic crawler identified by the token claude-web. It fetched recent web content for the Claude assistant.
Claude-SearchBot
Claude-SearchBot is Anthropic's crawler for improving search results shown to Claude users.
Claude-User
Claude-User supports Claude AI users. When individuals ask questions to Claude, it may access websites using a Claude-User agent.
ClaudeBot
ClaudeBot is Anthropic's crawler for gathering public web content that could contribute to the development of its generative AI models.
ClearscopeBot
Clearscope is an AI-driven SEO content optimization platform developed by Mushi Labs.
Cledara SaaS Management Agent
Cledara’s agent automates customer-approved SaaS admin tasks, including invoice collection and user management.
Cloudflare Browser Run
Cloudflare Browser Run is the signed agent identity for page loads made through Cloudflare Browser Rendering.
Cloudflare Crawler
Cloudflare Crawler is the bot identity used by the /crawl endpoint in Cloudflare Browser Rendering.
CMU-Cylab
An academic research bot, conducting research on patterns in counterfeit document sales over the internet.
Cốc Cốc
Coccocbot scrapes websites that are request from the Vietnamese search engine Coc Coc.
cognitiveSEO Crawler
CognitiveSEO is an SEO toolset that crawls the web and analyzes links.
Cohere AI
Cohere AI collects publicly available web text that helps train and refine Cohere's large language models for enterprise generative AI.
ContentKingBot
ContentKing (now Conductor Website Monitoring) is a website monitoring tool that continuously audits websites to help improve their performance and visibility.
Cookiebot
Cookiebot automates compliance with cookie laws and helps you manage your cookie consent preferences.
CookieScript
A cookie scanning bot that examines websites for cookie usage to help maintain GDPR and other privacy regulation compliance.
CoostiBot
CoostiBot crawls merchant websites for Coosti.dk, a Danish price comparison platform. Respects robots.txt.
Cotoyogi
Cotoyogi collects Japanese-language data resources for the Center for Research and Development on Data Lake within Japan's Research Organization of Informati.
Coveobot
Coveobot is a crawler operated by Coveo that indexes content for enterprise search, recommendations, and generative experience platforms.
Crawlson
Crawlson is a search engine crawler for the crawlson.com search engine.
creobot
This bot download assets from the origin server.
CriteoBot
CriteoBot is a crawler operated by Criteo that analyzes web content to serve relevant contextual ads.
Customer.io webhooks
Customer.io's webhook service for event-driven marketing automation and customer data platform.
CuteStat
CuteStat indexes page content for two stated outputs: its search results and its insights.
Cxense
The Cxensebot performs SEO monitoring and analysis of customer webpages.
Cybaa Agent
Cybaa Agent appears when a Cybaa customer asks the service to check a particular public domain or URL.
Dash0 Synthetic Monitoring
Dash0's Synthetic Monitoring provides proactive, automated insights into the availability and performance of your websites and APIs.
Datadog Synthetic Monitoring Robot
Datadog's automated monitoring service that performs synthetic tests to verify website availability and performance.
DataForSEO Site Auditor
DataForSEO Site Auditor scans websites for critical on-site SEO errors.
DataForSeoBot
DataForSeoBot is a backlink checker bot operated by DataForSEO that crawls websites to build and maintain their backlink database.
Dataprovider.com
Dataprovider.com indexes the web and structures the data.
Daum
A request ending in Daum/4. 1 is recorded as automated traffic for Daum's Korean search engine.
DeepSeek Bot
DeepSeek Bot crawls web content used to train and improve DeepSeek's generative AI models.
Detectify
Detectify is a web security scanner that performs automated security tests on web applications and attack surface monitoring.
Devin
Devin is a collaborative AI teammate built to help ambitious engineering teams achieve more.
Diffbot
Diffbot crawls and structures web pages into a knowledge graph that is sold for AI training, retrieval, and data enrichment.
DigitalOceanUptimeBot
DigitalOcean Uptime is a monitoring service that checks the health of any URL or IP address.
Direqt Anomura
Anomura is Direqt’s search crawler, it discovers and indexes pages their customers websites.
Discord Bot
Discord's link preview bot that crawls URLs shared in Discord chats to generate rich previews.
DotBot
DotBot is a web crawler operated by Moz (formerly SEOmoz) that collects data for their Link Explorer tool and Links API.
DuckAssistBot
DuckAssistBot is DuckDuckGo's crawler for AI-assisted answers in DuckDuckGo Search.
DuckDuckBot
DuckDuckBot is DuckDuckGo's crawler for improving its search results.
EasyScan
Automated scanning service that reviews online content on behalf of end users to identify potential legal issues.
Echobox Bot
Echobox provides newsletter and social-media automation for digital publishers.
Element451Bot
Element451Bot is the Knowledge Hub crawler for Element451, indexing pages so its AI assistant can answer questions for students.
Embedly
Embedly link preview service (operated by Medium) that fetches metadata and embeds for URLs.
eMoney Advisor
Collects raw financial data that can later be used for financial planning and analysis.
Epivoz Crawler
News aggregator needs to crawl news/blog articles to generate short summaries for page preview of attributed links.
EzoicBot
Ezoic is a technology platform for digital publishers. You can learn more about what Ezoic does here.
Facebook Webhooks
Facebook's webhook service that delivers real-time event notifications for Meta platform events and changes.
FacebookBot
FacebookBot is a Meta crawler that collects public web content.
FacebookExternalHit
Fetches content for shared links on Meta platforms to generate rich previews.
Factset_spyderbot
Factset uses a Python Selenium Crawler for web scraping to deliver reliable, current financial data.
FalBot
fal.ai's webhook service that delivers asynchronous notifications for AI model processing and generation tasks.
Fastmail Bot
Fastmail fetch and image proxy bot.
Fedicabot
Fedicabot fetches content for people using Fedica's social publishing tools. Fedica says the bot does not crawl whole websites.
FedReporter Bot for FFIEC
Bot to download data from the ffiec, Active/Closed/Branches File, Holding Company Data, 002 Data.
FirmlyAI Bot
FirmlyAI Bot navigates e-commerce websites for agentic commerce workflows.
FishBot
FishBot crawls webpages to deliver Open Source AI for All.
FlipboardProxy
Fetches and prepares website content for presentation in the Flipboard application.
FlyingPress
FlyingPress bot optimizes pages by generating critical CSS, detecting above-fold images, and delaying 3rd-party scripts.
FN Legislative and Regulatory Bot
FN Bot makes targeted rate-limited requests to a variety of publically available legislative and regulatory sources.
Freespoke
Freespoke is a search engine that believes in free speech and shows you all viewpoints.
Funnelback
Funnelback crawls the websites and data repositories configured for a Squiz enterprise search collection.
GeedoProductSearchBot
GeedoProductSearch is a web crawler operated by Geedo SIA that indexes product information from e-commerce websites.
GeedoShopProductFinder
GeedoShopProductFinder is the automated crawl.
Gemini Deep Research
Gemini Deep Research is a user-started research workflow in Google's Gemini apps.
GhostExplore
An aggregator service for websites powered by ghost.org.
Gigabot
Gigablast is the only non-Big Tech search engine in the U.S. that uses its own search index and algorithms.
GitHub Camo
GitHub Camo sits between a GitHub reader and an externally hosted image.
GitHub Hookshot
GitHub's webhooks for events like push, pull request, etc.
Google AdMob Reward Verification
Sends server-side verification callbacks to confirm users completed rewarded ad views.
Google Ads Creatives Assistant
Fetches website content for Google Ads creative generation and enhancement tools.
Google AdsBot
Google AdsBot is Google's web crawler for quality control of Google Ads.
Google Association Service
Verifies associations between apps and websites for Digital Asset Links.
Google Business Link Verification
Verifies that business links in Google Business Profile are accessible and return valid HTTP status codes.
Google Docs
Fetches images and page content when users insert links into Google Docs.
Google Feedfetcher
Feedfetcher is used for crawling RSS or Atom feeds for Google News and PubSubHubbub.
Google Image Proxy
Google's image caching proxy service used by Gmail and other Google services to cache and serve images.
Google Images
The Google Images bot is the search engine crawler for Google Images Search.
Google NotebookLM
Google NotebookLM fetches a web page when a person adds that URL as a notebook source.
Google PageRenderer
Upon user request, Google Page Renderer fetches and renders web pages.
Google Publisher Center
Google Publisher Center fetches and processes feeds that publishers explicitly supplied for use in Google News landing pages.
Google Read Aloud
Upon user request, Google Read Aloud fetches and reads out web pages using text-to-speech (TTS).
Google Scholar
Google Scholar crawls scholarly literature from publisher sites, institutional repositories, and university websites for its academic search engine.
Google Site Verifier
Google Site Verifier fetches Search Console verification tokens.
Google StoreBot
Google StoreBot gathers commerce information for Google Shopping.
Google Videos
The Google Videos bot is the search engine crawler for Google Video Search.
Google-AdWords-Express
Google-AdWords-Express is associated with a Google Ads product for small businesses.
Google-Adwords-Instant
Fetches advertiser landing pages when triggered by user actions in the Google Ads platform.
Google-Agent
Google-Agent navigates the web and performs actions upon user request, used by agents hosted on Google infrastructure such as Project Mariner.
Google-CloudVertexBot
Google-CloudVertexBot is the crawler Google documents for site-owner-requested crawls used to build Vertex AI Agents.
Google-Display-Ads-Bot
Verifies site eligibility during the AdSense approval process.
Google-Extended
Google-Extended is a robots. txt product token, not a separate crawler that identifies itself in HTTP requests.
Google-Flights-Search-Spackle
Google-Flights-Search-Spackle is a user-triggered Google Flights fetcher.
Google-InspectionTool
Google-InspectionTool retrieves a page when someone uses a Google Search testing feature such as the Rich Results Test or live URL inspection in Search Console.
Google-Safety
A Google-Safety request is an abuse check. Google's own example is malware detection for a link posted publicly on a Google property.
google-xrawler
Google's user-triggered fetcher for merchant product feeds. Fetches XML product feeds from e-commerce sites for syncing with Google Merchant Center.
Googlebot
Googlebot is the main crawler for Google Search.
GoogleOther
Crawling preferences addressed to the GoogleOther user agent don't affect any specific product.
GoogleStackdriverMonitoringBot
GoogleStackdriverMonitoringBot is operated by Google Cloud to perform uptime checks and monitor availability of services.
GPT-Actions
GPT-Actions identifies requests made when a GPT connects to an external HTTP API.
GPTBot
Crawls web content to improve OpenAI's generative AI models and ChatGPT; respects 'robots.txt' directives to exclude sites from training data.
Grok DeepSearch
Grok DeepSearch performs multi-step research across the web to answer complex Grok queries with cited sources.
Grok Search
Grok Search fetches web pages in real time to power Grok's search and answer features inside X and the Grok apps.
GrokBot
GrokBot is xAI's crawler used to gather web content for training the Grok family of models. xAI publishes limited documentation for it.
GTmetrix
GTmetrix provides metrics and insights for your site's loading speed and performance.
HarkBot
Hark is building the most advanced personal intelligence in the world.
HelloWork
HelloworkJobPostingBot turns vacancies from public career pages and other sources into listings for HelloWork's French recruitment platform.
Henry Shopping Agent
Executes checkout via browser automation using a user's card and signed mandate.
HetrixTools Uptime Monitoring Bot
HetrixTools website monitoring begins with an HTTP or HTTPS URL entered by a customer.
HEY Email Privacy Proxy
HEY email stops spy pixels and prevents user IP tracking by proxying all HTML email images, fonts, and external assets.
HIFIBot
HIFIBot is operated by HIFI, a financial services company for musicians and professional creators.
Hookdeck
A reliable Event Gateway for event-driven applications
HubSpot Page Fetcher
When posting to LinkedIn from Hubspot, images need to be pulled through to LinkedIn when published. The crawler performs this function.
Huckabuy Bot
Huckabot is Huckabuy’s main crawler which is utilized by almost all of Huckabuy’s products.
Hydrozen
Hydrozen is a tool for monitoring availability of your websites, Cronjobs, APIs, Domains, SSL etc.
Hype Machine
Since 2005, Hype Machine monitors music publications/blogs for posts about new artists and builds playlists using this metadata for listeners.
i-search-crawler
i-search-crawler is a web crawler operated by.
IASBot
IAS (Integral Ad Science) crawler, formerly known as AdmantX, is used for analyzing web content to ensure brand safety and suitability for advertisers.
iAskBot
iAskBot crawls and indexes web content to power iAsk.ai, an AI question-answering search engine.
IbouBot
IbouBot is the crawler of the Ibou Search Engine.
ICC Crawler
ICC-Crawler automatically crawls the Internet and collects web pages.
Idealo
Germany's largest price comparison service.
Iframely
An Iframely request usually starts with a person sharing a URL in an app that uses Iframely.
ImagesiftBot
ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support Hive's suite of web intelligence products.
IndeedJobBot
Indeed's job crawling bot that crawls job and job related information.
Inngest
Inngest is a platform for building event-driven applications.
Instapaper
Instapaper is an app that lets people save articles to read later.
Internet Archive - Archive-It
Internet Archive’s Archive-It service preserves publicly accessible web pages for the historical record.
Internet Archive Bot
Internet Archive Bot crawls public web pages for the Internet Archive's Wayback Machine. The operator also calls it archive.
InternetMeasurementBot
InternetMeasurementBot is operated by driftnet.io to discover and measure services that network owners and operators have publicly exposed.
JobicyBot
JobicyBot monitors and verifies job listings.
Jobs with GPT
Crawls job-related pages to power jobswithgpt.com, a platform for discovering AI-enhanced career opportunities.
jobswithgptcom-bot
Simple crawler focussing on only job postings for job search site.
Kagi Bot
Kagi Bot is the web crawler for the Kagi search engine. It crawls the web to build its own search index, which supports its ad-free search product.
KakaoTalk Scrap
The KakaoTalk scrap server collects and processes webpage data to create optimized previews for URLs.
kb.dk_bot
Royal Danish Library collects the Danish Internet according to the Danish Legal Deposit Act for research purposes.
KeldanNewsCrawler
Icelandic news article crawler for a news search on keldan.is.
Kernel Browsers
Runs browser automation on behalf of Kernel customers for web agents, automations, and web scraping.
Kernel Search
Kernel's search crawler that indexes web pages for AI-powered search and retrieval.
KimiBot
KimiBot crawls content potentially used to train Kimi's foundation models.
KlaviyoAIBot
Klaviyo’s web crawler for its Kai Customer Agent feature.
Lane Agent
The Lane Agent is a signed agent that acts on behalf of an end user to discover and purchase from a merchant.
Level9SearchBot
The public record for Level9SearchBot is unusually sparse.
Library Of Congress Web Archiving
Subject experts at the Library of Congress select websites for the Library's web archive.
LINE OGP Scraper
A scraper to get OGP by LINE Corporation.
LinerBot
LinerBot gathers web content for Liner, an AI research and answer assistant that cites the sources behind its responses.
Linespider
Linespider collects web data for search results inside LINE services.
LinkCheck
LinkCheck is a Siteimprove crawler used to inspect site content and links.
LinkCheckerBot
LinkCheckerBot is a backlink monitoring crawl.
LinkedInBot
LinkedInBot is a bot that renders links shared on LinkedIn.
LinksIndexerBot
LinksIndexerBot is an SEO bot that crawls websites to index backlinks and aggregate website summaries.
LinkupBot
LinkupBot discovers and retrieves publicly accessible webpages to build and maintain the Linkup web search index.
LogicMonitor SiteMonitor
LogicMonitor SiteMonitor monitors your website's uptime, performance, and availability from multiple global regions.
LogRocketBot
LogRocketBot is the asset-fetching client associated with LogRocket session replay.
Loomly Bot
LoomlyBot is part of Loomly's post-planning interface.
Lumar
Lumar runs cloud-based crawls for sites selected by its customers and turns the fetched page data into private website analytics.
MagiBot
MagiBot is a Peak Labs crawler associated with research into information extraction and retrieval from natural-language material.
MagnetmeBot
For Magnet. me customers, MagnetmeBot is part of the job-sync process.
MailRUBot
The mail.ru bot is a mail fetcher on behalf of the Mail.ru email service.
Make.com
Make.com workflow automation platform connector that fetches data from customer endpoints.
Manus Bot
Manus is the action engine that goes beyond answers to execute tasks, automate workflows, and extend your human reach.
Marfeel Audits Crawler
Marfeel's audit crawlers that periodically re-crawl traffic-receiving URLs to detect structured data, meta tags, and HTML issues.
Marfeel Flowcards Crawler
Marfeel's crawler that fetches content for Flowcards that load directly from specific URLs.
Marfeel Preview Crawler
Marfeel's previewer crawler used to render preview experiences for both mobile and desktop views.
Marfeel Social Crawler
Marfeel's crawler used for social experiences (Facebook, X/Twitter, Telegram, Reddit, LinkedIn).
Marginalia Search
Marginalia Search looks for parts of the web that large commercial indexes often bury, especially personal sites, older pages, and independent blogs.
marketgoo
marketgoo provides white label SEO tools.
Mars Finder
Mars Finder crawls for site search, not a web-wide consumer engine.
MediaMonitoringBot
News publishers see MediaMonitoringBot because its subscribers want alerts about newly published material.
Mediatoolkitbot
A mention in a Determ report can start with a visit from Mediatoolkitbot.
MelonMesa Bot
This bot is used to aggregate data about a popular online multiplayer game from consenting hosts who have opted-in to this collection.
meta-externalads
Crawls the web to improve advertising and business-related products and services.
meta-externalagent
The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly.
meta-externalfetcher
The Meta-ExternalFetcher crawler performs user-initiated fetches of individual links to support specific product functions.
Meta-ExternalTest
Meta bot used to test external integrations.
meta-webindexer
Crawls web content to provide search results for Meta AI users.
MicrosoftPreview
MicrosoftPreview fetches a URL when a Microsoft product needs a page snapshot.
MirrorWebCrawler
MirrorWebCrawler makes archived copies of websites for MirrorWeb Ltd.
MistralAI-Index
MistralAI-Index crawls and indexes web content for Mistral's search feature in Le Chat. Content it indexes is not used to train Mistral's generative models.
MistralAI-User
MistralAI-User fetches a page when someone asks Le Chat a question that calls for current web information.
MJ12bot
MJ12bot is a web crawler operated by Majestic-12 Ltd, a UK-based company that builds a search engine focused on backlink analysis and web structure mapping.
Mojeek
Details and information for webmasters regarding Mojeekbot, the web crawler for the Mojeek search engine.
MomenticBot
Momentic traffic comes from an end-to-end test written for a particular application.
MotoMinerBot
MotoMinerBot moves vehicle detail pages from dealership websites into MotoMiner's search index.
Moz rogerbot
Rogerbot is Moz's site audit crawler for Moz Pro Campaigns.
MRGbot
Search engine aimed at generating a corpus of data to be able to aggregate data in various ways.
MSN
MSNBot was Microsoft's crawler for the former MSN Search engine. It visited web pages so they could be considered for MSN's search index.
naver-blueno
A Naver user can trigger naver-blueno by inserting a link into a Naver service such as a blog or café.
naverbot
Naver's web crawler (also known as Yeti) is used by Naver, South Korea's largest search engine, to crawl and index web content.
Navu
Navu crawls the websites requested by their customers and prospects to train AIs for them.
Neevabot
Neevabot is the web crawler for the search engine neeva.com.
netEstate Imprint Crawler
The NetEstate Imprint crawler crawls websites for public contact information.
New York Times Newsgathering
New York Times Newsgathering is a shared identity for scripts written inside the Times newsroom.
NewRelic Minions
New Relic Synthetic monitoring infrastructure that performs API checks and virtual browser instances to monitor websites and applications from global locations
NewsBank
NewsBank aggregates licensed publisher content for schools, libraries, and government research, learning, and archiving.
NewsNow
The NewsNow bot is the web crawler for the news aggregator service NewsNow.
NitroBot
NitroBot is service traffic from NitroPack's cloud-based site-speed optimizer.
Nostra
Nostra accelerates site speed for managed web platforms.
Notabot
Crawler to integrate Helpfeel external search engine.
Novellum AI Crawl
Novellum.ai is building out tools for building agents. This MCP tool will be used by agents to crawl sites.
OAI-AdsBot
Validates the safety of web pages submitted as ads on ChatGPT; data collected is not used to train generative AI foundation models.
OAI-SearchBot
Indexes websites for inclusion in ChatGPT's search results; does not crawl content for AI model training.
Observer
Crawler looking for broken links to help improve world wide web as a whole.
OhDearBot
OhDearBot is a monitoring bot operated by Oh Dear that performs uptime checks, broken link detection, and mixed content scanning.
Omgilibot
Omgilibot crawls public web content for Webz.io, which packages and licenses web data feeds that are commonly used to train AI models.
OMIM Litrack
OMIM Litrack is a literature tracking crawler.
OnCrawl
Enterprise SEO platform powered by the industry-leading SEO Crawler and Log Analyzer.
Online Webceo Bot
A bot associated with WebCEO, a company that provides SEO tools and services.
OnticaBot
Feed-fetching crawler that discovers and fetches public articles for curated content feeds.
OpenGraph.io Bot
OpenGraph. io Bot appears when an OpenGraph. io customer asks the service to inspect a URL, often so a consumer product can display a shared-link preview.
OpenGraphXYZBot
Bot for opengraph.xyz service that generates and previews Open Graph meta tags and dynamic social media images
Orlo Link Preview
Orlo Link Preview works inside Orlo's social media publishing workflow.
Oseox
Oseox is a French SEO platform that provides.
Ozon Web Grabber
A component that serves to load previews for external and internal links.
Pagefreezer Website Archiving
Pagefreezer Website Archiving is a compliance.
PanguBot
PanguBot crawls web content used to train Huawei's Pangu family of large language models.
ParticleNewsBot
Particle is an AI powered aggregator that collects news from many sources.
Payhawk Invoice Fetching Agent
Automated browser bot that fetches invoices for users from supplier websites and attaches them to their expense records.
PayPal
PayPal delivers real-time event notifications for payments, subscriptions, and account updates.
payroll-bot
payroll-bot is an AI crawler operated by ADP, Inc. to collect publicly available legal and payroll documentation.
Perplexity-User
Handles user-initiated requests in Perplexity, accessing external content to provide real-time information; not used for automated crawling or AI training.
PerplexityBot
Indexes websites for inclusion in Perplexity's search results; does not crawl content for AI model training.
PetalBot
PetalBot is a web crawler operated by Huawei's Petal Search engine.
PhindBot
PhindBot crawls technical and developer-focused web content to power Phind, an AI answer engine aimed at programmers.
Pingdom Bot
Pingdom Bot is used by Pingdom's monitoring services to perform various checks on websites, including uptime and performance monitoring.
Pinterest Bot
Pinterest Bot, documented by Pinterest as Pinterestbot, gathers public web content for Pinterest.
Placer.ai-Crawler
Collects publicly available metadata from official websites in the open web.
Platebreaker
Recipe nutrition search engine. Indexes schema.org/Recipe JSON-LD. Respects robots.txt. Verified via Web Bot Auth.
Polar Webhooks
Polar's webhook service delivers real-time event notifications for payment processing, including purchases, subscriptions, cancellations, and refunds.
Potions
The Potions bot fetches product feeds and crawls data from its customers' websites, used for e-commerce related services.
prerender
It's HTML pre rendering service for SPA(Single Page Application) Website SEO.
PressEngine Bot
The PressEngine Bot verifies coverage created by video games press as genuine and their own creation.
Pricey
Pricey collects and compares product prices, showing trends and helping users find the best time to buy deals online.
Promptwatch Bot
Promptwatch Bot is Promptwatch's verification crawler that validates crawler-log integrations and runs on-demand SEO and performance audits.
ProximicBot
Proximic is Comscore's web crawler that performs contextual content analysis to help advertisers determine the best matching campaigns for a page's content.
PulsePoint Crawler
A web crawler used by PulsePoint, a digital advertising technology company, for content indexing and ads.txt verification.
QA.tech
The QA.tech web agent browses the website and identifies potential test cases, and executes tests against a web application
QStash
QStash is a platform for building event-driven applications.
QualifiedBot
Bot crawls customer websites to provide information to customer hosted chatbots.
Quantcastbot
Quantcast Bot is a web crawler used for advertisement quality assurance and to understand page content for Interest-Based Audiences.
Quartr Crawler
Quartr uses a crawler to obtain and deliver investor relations material.
Qwantbot
Crawls and indexes web content for Qwant search engine.
Rakuten Image extraction bot
Rakuten's image extraction bot fetches product images from sites that partner with Rakuten.
Razorpay-Webhook
Razorpay-Webhook delivers asynchronous payment notifications to a merchant's configured URL.
RDTvlokipBot
Official crawler for RDTvlokip Search, an independent French search engine. Respects robots.txt.
Redirect pizza destination monitor
redirect.pizza's destination monitor ensures that the redirect destination URLs are reachable.
Retool
Retool is the platform user agent attached to outbound HTTP requests from Retool.
Revvim
RevvimGort is listed as Revvim's crawler for examining customer websites and finding SEO opportunities.
RyeBot
Powers automated checkout on behalf of shoppers with explicit consent.
Sanity Webhooks
Sanity's webhook service that delivers real-time event notifications for content changes and other events.
Sansec Security Monitor
Sansec Security Monitor is a web crawler that monitors online stores for malicious code, data breaches, and digital skimming attacks.
SBIntuitionsBot
SBIntuitionsBot is a crawler operated by SB Intuitions Corp. that collects web data for AI development and information analysis.
ScreamingFrogBot
Screaming Frog SEO Spider is a website crawler used by SEO professionals for site audits and technical SEO analysis.
Screpy
SEO Checker for Screpy Bot that SEO.
SE Ranking Backlinks
SE Ranking's backlink analysis crawler that discovers and analyzes backlink profiles for SEO research and competitive analysis.
SearchAtlas Bot
Bot used to evaluate customer's websites and provide SEO optimization strategy.
SeekportBot
SeekportBot builds the independent web index used by the Seekport search engine.
Selectika AI
Selectika AI enrichment compute vision for Fashion.
SemanticScholarBot
SemanticScholarBot looks for academic PDFs on selected domains for Semantic Scholar.
Sentry Uptime Monitoring Bot
Sentry's Uptime Monitoring Bot performs health checks on configured URLs to monitor the availability and reliability of web services.
SEO Audit Check Bot
SEO audit check bot is likely an automated tool within the WebCEO platform that performs comprehensive SEO audits on websites.
seo4ajax
seo4ajax belongs to a prerendering service for single-page applications.
Seobility
Seobility is a browser-based online SEO software that helps you improve your website’s search engine rankings.
SerpstatBot
SerpstatBot is the Serpstat bot collects data for Serpstat's Backlink Analysis tool.
ServerHunterSpider
Hosting providers see ServerHunterSpider because Server Hunter imports plan information for its comparison site.
SeznamBot
SeznamBot is the web crawler operated by Seznam.cz, the leading Czech search engine.
ShapBot
Crawls and indexes web content to power Parallel's search and content extraction APIs for AI applications.
Shopify Webhooks
Shopify webhooks are useful for keeping your app in sync with Shopify data, or as a trigger to perform an additional action after that event has occurred.
Shortwave Image Fetcher
An email client that proxies all images found in HTML emails from to protect end customer's IP address and connection private.
SISTRIX Optimizer Uptime
A homepage request arriving once every minute may come from the uptime feature in a SISTRIX Optimizer project.
Site24x7
Site24x7 Bot is used by Site24x7's monitoring services to perform various checks on websites, including uptime and performance monitoring.
Sitebulb
Sitebulb is a desktop and cloud-based website crawler used by SEO professionals for technical SEO audits.
SiteGuru
SiteGuru is an SEO auditing tool that crawls.
Siteimprove Crawl
Siteimprove content suite (i.e. Quality Assurance, Accessibility, Policy, and SEO). Crawls run on ports are 80 for HTTP and 443 for HTTPS.
SiteSearch360
SiteSearch360 crawls websites whose owners use the Site Search 360 hosted search box.
Skroutz ImageBot
Skroutz ImageBot to fetch the individual product images.
Skype
Skype's URI Preview services fetches a page preview when someone posts a URL in a Skype message.
Slack-ImgProxy
Slack-ImgProxy is a bot operated by Slack that fetches and caches images posted in Slack channels.
Slackbot
Slackbot is Slack's default, general-purpose bot that handles various API requests and integrations.
SlackLinkExpandingBot
Slackbot Link Expanding is a bot operated by Slack that fetches metadata from shared links to create rich previews.
SMTnet PM Bot
SMTnet PM Bot visits partner company websites so their content can appear in SMTnet's on-site search.
SnapchatAdsBot
SnapchatAdsBot is a crawler operated by Snapchat that verifies and analyzes websites for their advertising platform.
SnapURLPreviewBot
SnapURLPreviewBot is a crawler operated by Snap Inc. that analyzes and generates previews of URLs shared on Snapchat and other Snap platforms.
Sogou Web Spider
Sogou Web Spider visits pages for the search index at sogou. com.
sourcedash
sourcedash is the crawler identity for Source Dashboard, a UK property search and management platform.
Spyglasses
Spyglasses accesses site content to assist with AEO capabilities. Learn more at https://www.spyglasses.io.
Stably
Stably is a QA testing bot that users run to E2E test their websites for functionality testing and protecting user flows against regressions.
Statabot
Statabot searches for stata.toc files and indexes their contents.
StatistikAustria
Bot to collect product prices for the official consumer price index of Austria.
StatsDroneBot
The StatsDrone affiliate marketing statistics scraping and aggregating tool.
StatusCake Page Speed
StatusCake Page Speed monitors your page load and render speeds.
StatusCake SSL Monitoring
StatusCake SSL monitors your website certificates for common issues
StatusCake Uptime
StatusCake monitors the uptime of your website.
Steam Chat
The Steam Chat bot fetches previews of URLs shared within the Steam client's chat feature.
Stripe Webhooks
Stripe's webhook service that delivers real-time event notifications for payment processing and account updates.
Stripebot
Crawls Stripe merchant websites to collect data for service delivery and financial regulatory compliance.
Strivve Automation
Strivve Web Bot Authentication Agent.
svix
svix is a webhook service for sending events to webhooks.
TangibleeBot
TangibleeBot reads retailer product pages for Tangiblee's visualization and virtual try-on services.
Telegram Bot
TelegramBot crawls websites to render a link preview when people send a message containing a URL in the Telegram messaging service.
TermlyBot
Crawls websites to detect and categorize cookies set by first and third parties.
Terracotta
The Terracotta bot scrapes websites for use in generating indices for serving searches using Ceramic's search product.
TikTokSpider
TikTokSpider is a TikTok and ByteDance web crawler used to index and analyze content outside the TikTok platform.
Timpibot
Timpibot crawls the web to build Timpi's decentralized search and data index, which is used to supply training and grounding data for AI applications.
TNOThesisCrawler
Research crawler by TNO collecting publicly available academic theses from university websites for the DIAMONDS platform.
Toutiao
Toutiao is ByteDance's automated news aggregation bot collecting content across web platforms.
Trellis-Services
Critical CSS Generator to Optimize Websites.
Trendiction Bot
Trendiction's web crawler that discovers and collects public web data for their social media monitoring and media intelligence platform.
TrustedSite
Crawl a customer website to perform basic validation.
TTD-Content
TTD-Content is a crawler operated by The Trade Desk that verifies content and quality of ad placements for their demand-side platform.
Tumblr
A Tumblr request usually starts with a post author pasting a URL.
TurnitinBot
TurnitinBot collects online material for Turnitin's plagiarism prevention service.
Twilio Knowledge
Twilio Knowledge fetches website content when a Twilio customer adds a web source to a knowledge base.
Twilio Proxy
Twilio's proxy service that handles communications between end-users and applications through Twilio's programmable voice and messaging platform.
TwinAgent
Automate complex operations end-to-end.
Twitterbot
Fetches content for shared links on X/Twitter to generate rich previews.
upday
upday is a news aggregator app, and its bot crawls news sources. It collects and indexes articles to be recommended to users on its platform.
Updown.io
Performs uptime and performance checks on websites.
Uptime Robot
Uptime Robot is a platform for monitoring and alerting on your applications.
UsercentricsBot
UsercentricsBot is operated by Usercentrics GmbH to scan websites for data processing services and third-party technologies.
v0bot
v0bot is listed as an AI crawler for v0 services, with v0bot as its stable user-agent token. That is the full purpose stated in the supplied bot record.
Velen Public Web Crawler
Velen Public Web Crawler collects public web content for Webz.io's data feeds, which are licensed for AI training, market intelligence, and monitoring.
Vemetric Favicon Bot
Fetches favicons from websites in the highest quality available.
Vercel build container
System-initiated requests made from Vercel's build container during a build
Vercel Favicon Bot
Vercel Favicon Bot is categorized as a preview client, and its name indicates a narrow job: obtaining a site's favicon for use in a Vercel interface or preview.
Vercel Screenshot Bot
Vercel Screenshot Bot is listed as a preview client whose name points to page-image capture.
vercelflags
vercelflags is a monitoring operated by its operator.
verceltracing
verceltracing is a monitoring operated by its operator.
videootv Bot
videootv Bot works on sites that publish through Digital Green's native advertising module.
Visually.io Shopify Editor
Shopify theme editor alternative for live, real-time store editing via a secure iframe and controlled proxy.
W3 Validator Services
W3C provides various free validation services that help check the conformance of Web sites against open standards.
WARDBot
WARDBot tracks URL status codes, helping users monitor the availability of web pages they have added to the monitoring list.
WebSpiderMount
Job wrapping data processor handling jobs distribution from employer websites to multiple endpoints, like job boards, advertisement platforms, job alerts etc.
WindowsForum-AI
Technology news aggregation bot for WindowsForum.com.
WMF Citoid
Citoid handles the automatic citation lookup in Wikimedia's VisualEditor.
WMF Zotero Translation Server
WMF Zotero Translation Server sits behind Citoid rather than facing Wikimedia editors directly.
WordCountBot
WordCountBot analyzes website word count based on public pages. All words belonging to public pages and included in HTML source code.
XY Archive Compliance Bot
XY Archive Compliance Bot visits websites selected by customers with recordkeeping requirements. The operator describes a two-part job.
Yahoo Ad Monitoring
A landing page listed in a Yahoo advertisement may be opened by Yahoo Ad Monitoring.
Yahoo Japan SEO Crawler
Yahoo Japan search engine crawler for SEO analysis.
Yahoo Link Preview
Yahoo Link Preview's bot fetches data from URLs shared on Yahoo platforms.
Yahoo! JAPAN
This Yahoo! JAPAN entry uses J-DLC as its stable token.
Yahoo! Slurp
Yahoo! Slurp is the web crawler (robot) used by Yahoo! Search to discover and index web pages for its search engine.
YahooCacheSystem
YahooCacheSystem caches website contents as part of the Yahoo! Search Service.
YahooMailProxy
Yahoo Mail Proxy is a content fetch proxy that retrieves the page content of URLs that are embedded within emails sent to Yahoo Mail users.
YandexAdditional
YandexAdditional is Yandex's dedicated user agent for content used by YandexGPT and other generative AI features.
Yandexbot
YandexBot is a web crawler operated by Yandex, a major Russian search engine.
Yeti
Yeti is the web crawler for Naver, a South Korean search engine. It indexes websites to provide search results and power other services on the Naver platform.
YGS Group Falconer Scraper
YGS Group Falconer Scraper supports Falconer, a coverage discovery product presented by YGS Content Licensing.
YisouSpider
The directory cannot independently verify YisouSpider's identity.
YouBot
YouBot crawls and indexes web pages to power the You.com AI search engine and its cited answers.
Zoombot
ZoomBot is SEOZoom's web crawler that builds.
ZoomInfo
Zoominfobot is an indexing robot for a web search engine, similar to Google. Created by Zoom Information Inc.(www.zoominfo.
ZumBot
ZumBot is a web crawler that indexes webpages for Zum Open Internet Search.
Be the brand AI recommends
Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.
