Promptwatch Logo

AI Bots & Web Crawlers Directory

Know exactly which AI crawlers, search engine bots, and automated agents reach your site: who runs them, what they do with your content, and how to allow or block each one.
AI search relevance:

439 bots

AdagioBot

AdagioBot is associated with Adagio's advertising platform and appears to inspect publisher websites for demand optimization.

AdvertisingNot relevant

AddSearchBot

As AddSearch adds content from your site to the search, the AddSearch bot gets counted as traffic by most analytics software.

SEOIndirect

AddThis

AddThis crawls pages to gather and refresh content used by its website marketing tools.

SEOIndirect

AdIdxBot

An AdIdxBot visit usually starts with a Bing Ads workflow rather than an organic search crawl.

Search Engine CrawlerRelevant

Adsense

The AdSense crawler visits participating sites in order to provide them with relevant ads.

Search Engine CrawlerRelevant

adsnaver

adsnaver is tied to paid search advertising, not Naver's general web index.

Search Engine CrawlerRelevant

Adyen Webhook

Adyen sends HTTP POST webhooks for events that a payment integration cannot safely infer from a browser return or an immediate API response.

WebhookNot relevant

Agency Analytics Crawler

A web crawler by Agency Analytics that allows their clients to check their own sites for SEO.

SEOIndirect

AGI Agent

The AGI Agent is a productivity assistant that takes actions and makes purchases on behalf of users.

AI AssistantRelevant

AhrefsBot

Powers the database for both Ahrefs, a marketing intelligence platform, and Yep, an independent, privacy-focused search engine.

SEOIndirect

AhrefsSiteAudit

Powers Ahrefs’ Site Audit tool. Ahrefs users can use Site Audit to analyze websites and find both technical SEO and on-page SEO issues.

SEOIndirect

AI Search

Cloudflare AI Search is a managed service that lets you connect your data and easily build AI-powered search.

AI SearchRelevant

AI Search External

Cloudflare AI Search is a managed service that lets you connect your data and easily build AI-powered search.

AI SearchRelevant

AI2Bot

AI2Bot is operated by the Allen Institute for Artificial Intelligence (Ai2) to crawl the web for content to train open-source AI models.

AI CrawlerRelevantUnverifiable

aiHitBot

aiHitBot collects and maintains historical information about companies.

AI CrawlerRelevantUnverifiable

AirsoftdbBot

Airsoftdb is a search engine for airsoft prod.

AggregatorIndirect

Algolia

The Algolia Crawler extracts content from your site and makes it searchable.

Search Engine CrawlerRelevant

All Africa Crawler

All Africa Crawler is part of AllAfrica Global Media's news aggregation and distribution operation.

Search Engine CrawlerRelevant

Alli AI Bot

Alli AI bot crawls customer websites to generate SEO recommendations.

SEOIndirect

Alphalens Bot

Indexes companies and their offerings to power more effective business discovery.

AI SearchRelevant

Amazon AdBot

Amazon AdBot scans pages that request ads from Amazon's advertising systems.

Search Engine CrawlerRelevant

Amazon Bedrock AgentCore Browser

AgentCore Browser provides a secure, cloud-based browser that enables AI agents to interact with websites.

AI AssistantRelevant

Amazon Bedrock Bot

Amazon Bedrock Bot runs when an Amazon Bedrock customer configures a website as a data source for a knowledge base.

AI CrawlerRelevant

Amazon Kendra

Amazon Kendra is a managed information retrieval and intelligent search service that uses natural language processing and advanced deep learning model.

AI AssistantRelevant

Amazon Product Discovery

Amazon Product Discovery collects public product details from Amazon selling-partner, brand, and retailer websites.

Search Engine CrawlerRelevantUnverifiable

Amazon Q

Amazon Q Business is a generative artificial intelligence (generative AI)-powered assistant that you can tailor to your business needs.

AI AssistantRelevant

Amazon Route 53 Health Check Service

Amazon Route 53 Health Check Service

MonitoringNot relevant

Amazon Seller Initiated Listing

Amazon's web crawler that helps sellers succeed by giving them the option to provide a URL to a website and create high-quality product pages in Amazon's store.

E-commerceRelevantUnverifiable

Amazonbot

Amazonbot is Amazon's general web crawler for improving Amazon products and services.

AI CrawlerRelevant

Amzn-SearchBot

Amzn-SearchBot crawls for improving Amazon search experiences (Alexa, Rufus).

AI SearchRelevant

Anchor Browser

The Web Browser for AI Agents.

AI CrawlerRelevant

Andibot

Andibot gathers web content for Andi, a conversational AI search assistant that answers questions with summaries and sources.

AI AssistantRelevantUnverifiable

Apify Website Content Crawler

Crawl websites and extract content to feed AI apps. Convert web data to Markdown or HTML, download files, and more.

AI AssistantRelevant

APIs-Google

Crawling preferences addressed to the APIs-Google user agent affect the delivery of push notification messages by Google APIs.

Search Engine CrawlerRelevant

Apple App Site Association

Apple App Site Association, usually shortened to AASA, is part of the trust setup behind Universal Links.

Social MediaIndirect

Apple Podcasts

Apple Podcasts crawler that only accesses URLs associated with registered content on Apple Podcasts. Does not follow robots.txt.

Feed FetcherNot relevant

Applebot

Applebot powers search features in Apple's ecosystem (Spotlight, Siri, Safari) and may be used to train Apple's foundation models for generative AI features.

AI CrawlerRelevant

Applebot-Extended

Applebot-Extended is a robots. txt control token, not a crawler that makes its own page requests. Apple calls it a secondary user agent.

AI TrainingRelevant

Arena Bot

Link preview bot for Arena.im's live engagement platform. Fetches page previews when URLs are shared in Arena's live chat, live blog, and community tools.

PreviewIndirect

Arquivo Web Crawler

Web crawler archives the Portuguese web.

ArchiverIndirect

Artemis Web Crawler

Artemis is a calm web reader with which you can follow websites and blogs.

Feed FetcherNot relevant

Artsdata Crawler

Web crawler that collects publicly available LOD for arts and culture in Canada.

AggregatorIndirect

Atlassian Jira Webhooks

Delivers webhook notifications from Jira Cloud when issues, projects, or other resources change.

WebhookNot relevant

Atlassian Rovo

Crawls and indexes web content for Atlassian Rovo's AI-powered search, chat, and agents.

AI CrawlerRelevant

atlassian-bot

atlassian-bot is a crawler for custom 3P websites that indexes data for rovo search.

AI CrawlerRelevant

Attracta

The Attracta bot is analyzes user website content as part of Attracta's SEO services.

SEOIndirect

Authory

The Authory bot visits websites to back up articles on behalf of journalists and other writers who use the service.

ArchiverIndirect

Awario Bot

Awario's web crawler used to discover and collect new and updated web data for their social media monitoring and brand mention tracking platform.

MonitoringNot relevantUnverifiable

Awario RSS Bot

One of Awario's primary web crawlers specialized in collecting RSS feed data.

Feed FetcherNot relevantUnverifiable

Awario Smart Bot

One of Awario's primary web crawlers that discovers and collects new and updated web data.

AnalyticsNot relevantUnverifiable

Baidu ADS Server Proxy

Baidu ADS Server Proxy is identified in this directory as Baidu's scrubbing proxy.

Search Engine CrawlerRelevant

BaiduSpider

Baiduspider is Baidu’s web crawler that indexes websites for inclusion in its Chinese-market search results.

Search Engine CrawlerRelevant

Barkrowler

Barkrowler is Babbar's web crawler that fuels and updates their graph representation of the web, providing SEO tools for the marketing community.

SEOIndirect

BestChange Bot

The BestChange bot downloads exchange rate information from 600 websites every 5 seconds.

AggregatorIndirect

Better Stack

Better Stack is a platform for monitoring and alerting on your applications.

MonitoringNot relevant

Bibliothèque nationale de France Crawler

Bibliothèque nationale de France's mission is to collect, catalog, preserve, enrich and communicate the national documentary heritage.

Academic ResearchIndirect

Big Sur AI

Big Sur AI Crawler, crawlers users websites to enable AI-infused experiences.

AI CrawlerRelevant

Bing Preview

BingPreview generates page snapshots for Bing. Note that BingPreview has desktop and mobile variants.

PreviewIndirect

Bingbot

Bingbot is Microsoft's web crawler used for indexing websites for Bing Search.

Search Engine CrawlerRelevant

BLEXBot

SEO PowerSuite Link Explorer (webmeup.com) is the world's freshest backlink index, and the primary source of backlink-related data for the SEO PowerSuite tools.

SEOIndirect

Bluesky Link Preview Service

Bluesky social pulls links in advance to render webpage previews.

PreviewIndirect

BoardGamePrices Bot

Price comparison site for board games. Need to crawl store pages for participating stores. All stores give permission to be crawled.

AggregatorIndirect

BorderxBot

E-commerce product crawler operated by Borderxlab.

AI CrawlerRelevant

Botify

Botify SiteCrawler is part of the Botify Analytics suite.

SEOIndirect

Brandwatch

The Magpie Crawler indexes content for its soical media monitoring solution.

AI CrawlerRelevant

BraveBot

BraveBot crawls and indexes web pages for the Brave Search index, which also grounds answers from Brave's Leo AI assistant.

AI AssistantRelevant

Brightbot

Brightbot is Bright Data's crawler layer that monitors the health of websites and enforces ethical web data collection.

AnalyticsNot relevant

Browserbase

Runs headless browser automation on behalf of Browserbase customers for web scraping, form submission, and testing.

AI CrawlerRelevant

Buffer Link Preview Bot

Helps Buffer users create better social media posts by generating rich previews when they share links

PreviewIndirect

Bytespider

Bytespider is ByteDance's web crawler used to gather training data for their AI large language models.

AI CrawlerRelevantUnverifiable

CaliberBot

Conductor uses Caliperbot to fetch websites belonging to clients and prospects.

SEOIndirect

Capital One Bot

Capital One Bot crawls dealer websites for getting the usage information for Capital One lead navigator button.

AggregatorIndirect

CCBot

CCBot is operated by the Common Crawl Foundation to crawl web content for AI training and research.

AI CrawlerRelevant

CensysInspectBot

CensysInspectBot belongs to Censys's internet-wide scanning activity.

AnalyticsNot relevantUnverifiable

Channel3Bot

Crawls product detail pages to index content for AI-powered product discovery, routing shoppers to original websites.

AI CrawlerRelevant

ChatGPT agent

Agent that can use its own browser to perform tasks for user.

AI AssistantRelevant

ChatGPT-Operator

Handles user-initiated requests from ChatGPT operator accessing external content; not used for automated crawling or AI training.

AI AssistantRelevant

ChatGPT-User

Handles user-initiated requests in ChatGPT, accessing external content to provide real-time information; not used for automated crawling or AI training.

AI AssistantRelevantTracked by Promptwatch

Chathive crawler

Chathive crawler is the Inteso Group BV crawler that Chathive customers use to bring their own website content into an AI assistant.

AI AssistantRelevant

Checkly

Checkly is a platform for monitoring and alerting on your applications.

MonitoringNot relevant

Chrome Lighthouse

PageSpeed Insights (PSI) reports on the user experience of a page on both mobile and desktop devices, and provides suggestions on how that page may be improved.

AnalyticsNot relevant

Chrome Privacy Preserving Prefetch Proxy

Chrome's Privacy Preserving Prefetch Proxy service that fetches /.well-known/traffic-advice to enable privacy-preserving prefetch hints.

PreviewIndirect

CitibotSiteCrawler

CitibotSiteCrawler collects public data from government websites to power Citibot’s AI civic engagement tools.

AI CrawlerRelevant

ClarityBot

ClarityBot is seoClarity's web crawler that performs technical SEO audits, analyzes content, and monitors website performance.

SEOIndirectUnverifiable

Claude Web

Claude Web is a legacy Anthropic crawler identified by the token claude-web. It fetched recent web content for the Claude assistant.

AI CrawlerRelevantTracked by Promptwatch

Claude-SearchBot

Claude-SearchBot is Anthropic's crawler for improving search results shown to Claude users.

AI AssistantRelevantTracked by Promptwatch

Claude-User

Claude-User supports Claude AI users. When individuals ask questions to Claude, it may access websites using a Claude-User agent.

AI AssistantRelevantTracked by Promptwatch

ClaudeBot

ClaudeBot is Anthropic's crawler for gathering public web content that could contribute to the development of its generative AI models.

AI CrawlerRelevantTracked by Promptwatch

ClearscopeBot

Clearscope is an AI-driven SEO content optimization platform developed by Mushi Labs.

SEOIndirect

Cledara SaaS Management Agent

Cledara’s agent automates customer-approved SaaS admin tasks, including invoice collection and user management.

AI AssistantRelevant

Cloudflare Browser Run

Cloudflare Browser Run is the signed agent identity for page loads made through Cloudflare Browser Rendering.

AI AssistantRelevant

Cloudflare Crawler

Cloudflare Crawler is the bot identity used by the /crawl endpoint in Cloudflare Browser Rendering.

AI CrawlerRelevant

CMU-Cylab

An academic research bot, conducting research on patterns in counterfeit document sales over the internet.

Academic ResearchIndirect

Cốc Cốc

Coccocbot scrapes websites that are request from the Vietnamese search engine Coc Coc.

Search Engine CrawlerRelevant

cognitiveSEO Crawler

CognitiveSEO is an SEO toolset that crawls the web and analyzes links.

SEOIndirect

Cohere AI

Cohere AI collects publicly available web text that helps train and refine Cohere's large language models for enterprise generative AI.

AI CrawlerRelevantTracked by Promptwatch

ContentKingBot

ContentKing (now Conductor Website Monitoring) is a website monitoring tool that continuously audits websites to help improve their performance and visibility.

AnalyticsNot relevantUnverifiable

Cookiebot

Cookiebot automates compliance with cookie laws and helps you manage your cookie consent preferences.

MonitoringNot relevant

CookieScript

A cookie scanning bot that examines websites for cookie usage to help maintain GDPR and other privacy regulation compliance.

MonitoringNot relevantUnverifiable

CoostiBot

CoostiBot crawls merchant websites for Coosti.dk, a Danish price comparison platform. Respects robots.txt.

AggregatorIndirect

Cotoyogi

Cotoyogi collects Japanese-language data resources for the Center for Research and Development on Data Lake within Japan's Research Organization of Informati.

AI CrawlerRelevantUnverifiable

Coveobot

Coveobot is a crawler operated by Coveo that indexes content for enterprise search, recommendations, and generative experience platforms.

AI AssistantRelevantUnverifiable

Crawlson

Crawlson is a search engine crawler for the crawlson.com search engine.

Search Engine CrawlerRelevant

creobot

This bot download assets from the origin server.

PreviewIndirect

CriteoBot

CriteoBot is a crawler operated by Criteo that analyzes web content to serve relevant contextual ads.

AdvertisingNot relevant

Customer.io webhooks

Customer.io's webhook service for event-driven marketing automation and customer data platform.

WebhookNot relevant

CuteStat

CuteStat indexes page content for two stated outputs: its search results and its insights.

Search Engine CrawlerRelevant

Cxense

The Cxensebot performs SEO monitoring and analysis of customer webpages.

SEOIndirect

Cybaa Agent

Cybaa Agent appears when a Cybaa customer asks the service to check a particular public domain or URL.

MonitoringNot relevant

Dash0 Synthetic Monitoring

Dash0's Synthetic Monitoring provides proactive, automated insights into the availability and performance of your websites and APIs.

MonitoringNot relevant

Datadog Synthetic Monitoring Robot

Datadog's automated monitoring service that performs synthetic tests to verify website availability and performance.

MonitoringNot relevant

DataForSEO Site Auditor

DataForSEO Site Auditor scans websites for critical on-site SEO errors.

SEOIndirect

DataForSeoBot

DataForSeoBot is a backlink checker bot operated by DataForSEO that crawls websites to build and maintain their backlink database.

SEOIndirect

Dataprovider.com

Dataprovider.com indexes the web and structures the data.

Search Engine CrawlerRelevant

Daum

A request ending in Daum/4. 1 is recorded as automated traffic for Daum's Korean search engine.

Search Engine CrawlerRelevant

DeepSeek Bot

DeepSeek Bot crawls web content used to train and improve DeepSeek's generative AI models.

AI CrawlerRelevantTracked by Promptwatch

Detectify

Detectify is a web security scanner that performs automated security tests on web applications and attack surface monitoring.

MonitoringNot relevant

Devin

Devin is a collaborative AI teammate built to help ambitious engineering teams achieve more.

AI AssistantRelevant

Diffbot

Diffbot crawls and structures web pages into a knowledge graph that is sold for AI training, retrieval, and data enrichment.

AI CrawlerRelevant

DigitalOceanUptimeBot

DigitalOcean Uptime is a monitoring service that checks the health of any URL or IP address.

MonitoringNot relevantUnverifiable

Direqt Anomura

Anomura is Direqt’s search crawler, it discovers and indexes pages their customers websites.

AI SearchRelevant

Discord Bot

Discord's link preview bot that crawls URLs shared in Discord chats to generate rich previews.

PreviewIndirectUnverifiable

DotBot

DotBot is a web crawler operated by Moz (formerly SEOmoz) that collects data for their Link Explorer tool and Links API.

SEOIndirectUnverifiable

DuckAssistBot

DuckAssistBot is DuckDuckGo's crawler for AI-assisted answers in DuckDuckGo Search.

AI AssistantRelevant

DuckDuckBot

DuckDuckBot is DuckDuckGo's crawler for improving its search results.

Search Engine CrawlerRelevant

EasyScan

Automated scanning service that reviews online content on behalf of end users to identify potential legal issues.

AI AssistantRelevant

Echobox Bot

Echobox provides newsletter and social-media automation for digital publishers.

AI CrawlerRelevant

Element451Bot

Element451Bot is the Knowledge Hub crawler for Element451, indexing pages so its AI assistant can answer questions for students.

AI SearchRelevant

Embedly

Embedly link preview service (operated by Medium) that fetches metadata and embeds for URLs.

PreviewIndirect

eMoney Advisor

Collects raw financial data that can later be used for financial planning and analysis.

AggregatorIndirect

Epivoz Crawler

News aggregator needs to crawl news/blog articles to generate short summaries for page preview of attributed links.

AggregatorIndirect

EzoicBot

Ezoic is a technology platform for digital publishers. You can learn more about what Ezoic does here.

SEOIndirect

Facebook Webhooks

Facebook's webhook service that delivers real-time event notifications for Meta platform events and changes.

WebhookNot relevant

FacebookBot

FacebookBot is a Meta crawler that collects public web content.

AI CrawlerRelevant

FacebookExternalHit

Fetches content for shared links on Meta platforms to generate rich previews.

PreviewIndirect

Factset_spyderbot

Factset uses a Python Selenium Crawler for web scraping to deliver reliable, current financial data.

AggregatorIndirect

FalBot

fal.ai's webhook service that delivers asynchronous notifications for AI model processing and generation tasks.

WebhookNot relevant

Fastmail Bot

Fastmail fetch and image proxy bot.

PreviewIndirect

Fedicabot

Fedicabot fetches content for people using Fedica's social publishing tools. Fedica says the bot does not crawl whole websites.

Social MediaIndirect

FedReporter Bot for FFIEC

Bot to download data from the ffiec, Active/Closed/Branches File, Holding Company Data, 002 Data.

AggregatorIndirect

FirmlyAI Bot

FirmlyAI Bot navigates e-commerce websites for agentic commerce workflows.

AI AssistantRelevant

FishBot

FishBot crawls webpages to deliver Open Source AI for All.

AI CrawlerRelevant

FlipboardProxy

Fetches and prepares website content for presentation in the Flipboard application.

Feed FetcherNot relevantUnverifiable

FlyingPress

FlyingPress bot optimizes pages by generating critical CSS, detecting above-fold images, and delaying 3rd-party scripts.

SEOIndirect

FN Legislative and Regulatory Bot

FN Bot makes targeted rate-limited requests to a variety of publically available legislative and regulatory sources.

AggregatorIndirect

Freespoke

Freespoke is a search engine that believes in free speech and shows you all viewpoints.

Search Engine CrawlerRelevant

Funnelback

Funnelback crawls the websites and data repositories configured for a Squiz enterprise search collection.

Search Engine CrawlerRelevant

GeedoProductSearchBot

GeedoProductSearch is a web crawler operated by Geedo SIA that indexes product information from e-commerce websites.

E-commerceRelevant

GeedoShopProductFinder

GeedoShopProductFinder is the automated crawl.

AggregatorIndirect

Gemini Deep Research

Gemini Deep Research is a user-started research workflow in Google's Gemini apps.

AI AssistantRelevant

GhostExplore

An aggregator service for websites powered by ghost.org.

AggregatorIndirect

Gigabot

Gigablast is the only non-Big Tech search engine in the U.S. that uses its own search index and algorithms.

Search Engine CrawlerRelevant

GitHub Camo

GitHub Camo sits between a GitHub reader and an externally hosted image.

PreviewIndirect

GitHub Hookshot

GitHub's webhooks for events like push, pull request, etc.

WebhookNot relevant

Google AdMob Reward Verification

Sends server-side verification callbacks to confirm users completed rewarded ad views.

AdvertisingNot relevant

Google Ads Creatives Assistant

Fetches website content for Google Ads creative generation and enhancement tools.

AI AssistantRelevant

Google AdsBot

Google AdsBot is Google's web crawler for quality control of Google Ads.

Search Engine CrawlerRelevant

Google Association Service

Verifies associations between apps and websites for Digital Asset Links.

VerificationNot relevant

Google Business Link Verification

Verifies that business links in Google Business Profile are accessible and return valid HTTP status codes.

VerificationNot relevant

Google Docs

Fetches images and page content when users insert links into Google Docs.

PreviewIndirect

Google Feedfetcher

Feedfetcher is used for crawling RSS or Atom feeds for Google News and PubSubHubbub.

Feed FetcherNot relevant

Google Image Proxy

Google's image caching proxy service used by Gmail and other Google services to cache and serve images.

PreviewIndirect

Google Images

The Google Images bot is the search engine crawler for Google Images Search.

Search Engine CrawlerRelevant

Google NotebookLM

Google NotebookLM fetches a web page when a person adds that URL as a notebook source.

AI AssistantRelevant

Google PageRenderer

Upon user request, Google Page Renderer fetches and renders web pages.

PreviewIndirect

Google Publisher Center

Google Publisher Center fetches and processes feeds that publishers explicitly supplied for use in Google News landing pages.

Feed FetcherNot relevant

Google Read Aloud

Upon user request, Google Read Aloud fetches and reads out web pages using text-to-speech (TTS).

User InitiatedRelevant

Google Scholar

Google Scholar crawls scholarly literature from publisher sites, institutional repositories, and university websites for its academic search engine.

Search Engine CrawlerRelevant

Google Site Verifier

Google Site Verifier fetches Search Console verification tokens.

VerificationNot relevant

Google StoreBot

Google StoreBot gathers commerce information for Google Shopping.

Search Engine CrawlerRelevant

Google Videos

The Google Videos bot is the search engine crawler for Google Video Search.

Search Engine CrawlerRelevant

Google-AdWords-Express

Google-AdWords-Express is associated with a Google Ads product for small businesses.

SEOIndirect

Google-Adwords-Instant

Fetches advertiser landing pages when triggered by user actions in the Google Ads platform.

AdvertisingNot relevant

Google-Agent

Google-Agent navigates the web and performs actions upon user request, used by agents hosted on Google infrastructure such as Project Mariner.

AgentRelevantTracked by Promptwatch

Google-CloudVertexBot

Google-CloudVertexBot is the crawler Google documents for site-owner-requested crawls used to build Vertex AI Agents.

AI AssistantRelevant

Google-Display-Ads-Bot

Verifies site eligibility during the AdSense approval process.

Search Engine CrawlerRelevant

Google-Extended

Google-Extended is a robots. txt product token, not a separate crawler that identifies itself in HTTP requests.

AI CrawlerRelevantTracked by Promptwatch

Google-Flights-Search-Spackle

Google-Flights-Search-Spackle is a user-triggered Google Flights fetcher.

Search Engine CrawlerRelevant

Google-InspectionTool

Google-InspectionTool retrieves a page when someone uses a Google Search testing feature such as the Rich Results Test or live URL inspection in Search Console.

MonitoringNot relevant

Google-Safety

A Google-Safety request is an abuse check. Google's own example is malware detection for a link posted publicly on a Google property.

MonitoringNot relevant

google-xrawler

Google's user-triggered fetcher for merchant product feeds. Fetches XML product feeds from e-commerce sites for syncing with Google Merchant Center.

Search Engine CrawlerRelevant

Googlebot

Googlebot is the main crawler for Google Search.

Search Engine CrawlerRelevant

GoogleOther

Crawling preferences addressed to the GoogleOther user agent don't affect any specific product.

Search Engine CrawlerRelevant

GoogleStackdriverMonitoringBot

GoogleStackdriverMonitoringBot is operated by Google Cloud to perform uptime checks and monitor availability of services.

MonitoringNot relevantUnverifiable

GPT-Actions

GPT-Actions identifies requests made when a GPT connects to an external HTTP API.

AI AssistantRelevant

GPTBot

Crawls web content to improve OpenAI's generative AI models and ChatGPT; respects 'robots.txt' directives to exclude sites from training data.

AI CrawlerRelevantTracked by Promptwatch

Grok DeepSearch

Grok DeepSearch performs multi-step research across the web to answer complex Grok queries with cited sources.

AI AssistantRelevantUnverifiableTracked by Promptwatch

Grok Search

Grok Search fetches web pages in real time to power Grok's search and answer features inside X and the Grok apps.

AI AssistantRelevantUnverifiableTracked by Promptwatch

GrokBot

GrokBot is xAI's crawler used to gather web content for training the Grok family of models. xAI publishes limited documentation for it.

AI CrawlerRelevantUnverifiableTracked by Promptwatch

GTmetrix

GTmetrix provides metrics and insights for your site's loading speed and performance.

AnalyticsNot relevant

HarkBot

Hark is building the most advanced personal intelligence in the world.

AI AssistantRelevant

HelloWork

HelloworkJobPostingBot turns vacancies from public career pages and other sources into listings for HelloWork's French recruitment platform.

AggregatorIndirect

Henry Shopping Agent

Executes checkout via browser automation using a user's card and signed mandate.

AI AssistantRelevant

HetrixTools Uptime Monitoring Bot

HetrixTools website monitoring begins with an HTTP or HTTPS URL entered by a customer.

MonitoringNot relevant

HEY Email Privacy Proxy

HEY email stops spy pixels and prevents user IP tracking by proxying all HTML email images, fonts, and external assets.

PreviewIndirect

HIFIBot

HIFIBot is operated by HIFI, a financial services company for musicians and professional creators.

AI AssistantRelevant

Hookdeck

A reliable Event Gateway for event-driven applications

WebhookNot relevant

HubSpot Page Fetcher

When posting to LinkedIn from Hubspot, images need to be pulled through to LinkedIn when published. The crawler performs this function.

PreviewIndirect

Huckabuy Bot

Huckabot is Huckabuy’s main crawler which is utilized by almost all of Huckabuy’s products.

SEOIndirect

Hydrozen

Hydrozen is a tool for monitoring availability of your websites, Cronjobs, APIs, Domains, SSL etc.

MonitoringNot relevant

Hype Machine

Since 2005, Hype Machine monitors music publications/blogs for posts about new artists and builds playlists using this metadata for listeners.

Search Engine CrawlerRelevant

i-search-crawler

i-search-crawler is a web crawler operated by.

Search Engine CrawlerRelevant

IASBot

IAS (Integral Ad Science) crawler, formerly known as AdmantX, is used for analyzing web content to ensure brand safety and suitability for advertisers.

AdvertisingNot relevantUnverifiable

iAskBot

iAskBot crawls and indexes web content to power iAsk.ai, an AI question-answering search engine.

AI AssistantRelevantUnverifiable

IbouBot

IbouBot is the crawler of the Ibou Search Engine.

Search Engine CrawlerRelevant

ICC Crawler

ICC-Crawler automatically crawls the Internet and collects web pages.

AI CrawlerRelevant

Idealo

Germany's largest price comparison service.

AggregatorIndirect

Iframely

An Iframely request usually starts with a person sharing a URL in an app that uses Iframely.

PreviewIndirect

ImagesiftBot

ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support Hive's suite of web intelligence products.

AI CrawlerRelevant

IndeedJobBot

Indeed's job crawling bot that crawls job and job related information.

AggregatorIndirect

Inngest

Inngest is a platform for building event-driven applications.

WebhookNot relevant

Instapaper

Instapaper is an app that lets people save articles to read later.

AI AssistantRelevant

Internet Archive - Archive-It

Internet Archive’s Archive-It service preserves publicly accessible web pages for the historical record.

ArchiverIndirect

Internet Archive Bot

Internet Archive Bot crawls public web pages for the Internet Archive's Wayback Machine. The operator also calls it archive.

ArchiverIndirect

InternetMeasurementBot

InternetMeasurementBot is operated by driftnet.io to discover and measure services that network owners and operators have publicly exposed.

MonitoringNot relevantUnverifiable

JobicyBot

JobicyBot monitors and verifies job listings.

Search Engine CrawlerRelevant

Jobs with GPT

Crawls job-related pages to power jobswithgpt.com, a platform for discovering AI-enhanced career opportunities.

Search Engine CrawlerRelevant

jobswithgptcom-bot

Simple crawler focussing on only job postings for job search site.

Search Engine CrawlerRelevant

Kagi Bot

Kagi Bot is the web crawler for the Kagi search engine. It crawls the web to build its own search index, which supports its ad-free search product.

Search Engine CrawlerRelevant

KakaoTalk Scrap

The KakaoTalk scrap server collects and processes webpage data to create optimized previews for URLs.

PreviewIndirect

kb.dk_bot

Royal Danish Library collects the Danish Internet according to the Danish Legal Deposit Act for research purposes.

Academic ResearchIndirect

KeldanNewsCrawler

Icelandic news article crawler for a news search on keldan.is.

Search Engine CrawlerRelevant

Kernel Browsers

Runs browser automation on behalf of Kernel customers for web agents, automations, and web scraping.

AI CrawlerRelevant

Kernel Search

Kernel's search crawler that indexes web pages for AI-powered search and retrieval.

AI SearchRelevant

KimiBot

KimiBot crawls content potentially used to train Kimi's foundation models.

AI CrawlerRelevant

KlaviyoAIBot

Klaviyo’s web crawler for its Kai Customer Agent feature.

AI SearchRelevant

Lane Agent

The Lane Agent is a signed agent that acts on behalf of an end user to discover and purchase from a merchant.

AI AssistantRelevant

Level9SearchBot

The public record for Level9SearchBot is unusually sparse.

Search Engine CrawlerRelevant

Library Of Congress Web Archiving

Subject experts at the Library of Congress select websites for the Library's web archive.

Academic ResearchIndirect

LINE OGP Scraper

A scraper to get OGP by LINE Corporation.

Social MediaIndirect

LinerBot

LinerBot gathers web content for Liner, an AI research and answer assistant that cites the sources behind its responses.

AI AssistantRelevantUnverifiable

Linespider

Linespider collects web data for search results inside LINE services.

Search Engine CrawlerRelevant

LinkCheck

LinkCheck is a Siteimprove crawler used to inspect site content and links.

SEOIndirect

LinkCheckerBot

LinkCheckerBot is a backlink monitoring crawl.

SEOIndirect

LinkedInBot

LinkedInBot is a bot that renders links shared on LinkedIn.

PreviewIndirect

LinksIndexerBot

LinksIndexerBot is an SEO bot that crawls websites to index backlinks and aggregate website summaries.

PreviewIndirect

LinkupBot

LinkupBot discovers and retrieves publicly accessible webpages to build and maintain the Linkup web search index.

Search Engine CrawlerRelevant

LogicMonitor SiteMonitor

LogicMonitor SiteMonitor monitors your website's uptime, performance, and availability from multiple global regions.

MonitoringNot relevant

LogRocketBot

LogRocketBot is the asset-fetching client associated with LogRocket session replay.

AnalyticsNot relevantUnverifiable

Loomly Bot

LoomlyBot is part of Loomly's post-planning interface.

PreviewIndirect

Lumar

Lumar runs cloud-based crawls for sites selected by its customers and turns the fetched page data into private website analytics.

SEOIndirect

MagiBot

MagiBot is a Peak Labs crawler associated with research into information extraction and retrieval from natural-language material.

Search Engine CrawlerRelevant

MagnetmeBot

For Magnet. me customers, MagnetmeBot is part of the job-sync process.

AggregatorIndirect

MailRUBot

The mail.ru bot is a mail fetcher on behalf of the Mail.ru email service.

PreviewIndirect

Make.com

Make.com workflow automation platform connector that fetches data from customer endpoints.

AI CrawlerRelevant

Manus Bot

Manus is the action engine that goes beyond answers to execute tasks, automate workflows, and extend your human reach.

AI AssistantRelevant

Marfeel Audits Crawler

Marfeel's audit crawlers that periodically re-crawl traffic-receiving URLs to detect structured data, meta tags, and HTML issues.

SEOIndirect

Marfeel Flowcards Crawler

Marfeel's crawler that fetches content for Flowcards that load directly from specific URLs.

PreviewIndirect

Marfeel Preview Crawler

Marfeel's previewer crawler used to render preview experiences for both mobile and desktop views.

PreviewIndirect

Marfeel Social Crawler

Marfeel's crawler used for social experiences (Facebook, X/Twitter, Telegram, Reddit, LinkedIn).

PreviewIndirect

Marginalia Search

Marginalia Search looks for parts of the web that large commercial indexes often bury, especially personal sites, older pages, and independent blogs.

Search Engine CrawlerRelevant

marketgoo

marketgoo provides white label SEO tools.

SEOIndirect

Mars Finder

Mars Finder crawls for site search, not a web-wide consumer engine.

Search Engine CrawlerRelevant

MediaMonitoringBot

News publishers see MediaMonitoringBot because its subscribers want alerts about newly published material.

AggregatorIndirect

Mediatoolkitbot

A mention in a Determ report can start with a visit from Mediatoolkitbot.

AggregatorIndirect

MelonMesa Bot

This bot is used to aggregate data about a popular online multiplayer game from consenting hosts who have opted-in to this collection.

AggregatorIndirect

meta-externalads

Crawls the web to improve advertising and business-related products and services.

AdvertisingNot relevant

meta-externalagent

The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly.

AI CrawlerRelevant

meta-externalfetcher

The Meta-ExternalFetcher crawler performs user-initiated fetches of individual links to support specific product functions.

User InitiatedRelevant

Meta-ExternalTest

Meta bot used to test external integrations.

AI AssistantRelevant

meta-webindexer

Crawls web content to provide search results for Meta AI users.

AI CrawlerRelevant

MicrosoftPreview

MicrosoftPreview fetches a URL when a Microsoft product needs a page snapshot.

PreviewIndirect

MirrorWebCrawler

MirrorWebCrawler makes archived copies of websites for MirrorWeb Ltd.

AggregatorIndirect

MistralAI-Index

MistralAI-Index crawls and indexes web content for Mistral's search feature in Le Chat. Content it indexes is not used to train Mistral's generative models.

AI AssistantRelevantTracked by Promptwatch

MistralAI-User

MistralAI-User fetches a page when someone asks Le Chat a question that calls for current web information.

AI AssistantRelevantTracked by Promptwatch

MJ12bot

MJ12bot is a web crawler operated by Majestic-12 Ltd, a UK-based company that builds a search engine focused on backlink analysis and web structure mapping.

Search Engine CrawlerRelevantUnverifiable

Mojeek

Details and information for webmasters regarding Mojeekbot, the web crawler for the Mojeek search engine.

Search Engine CrawlerRelevant

MomenticBot

Momentic traffic comes from an end-to-end test written for a particular application.

MonitoringRelevant

MotoMinerBot

MotoMinerBot moves vehicle detail pages from dealership websites into MotoMiner's search index.

AggregatorIndirect

Moz rogerbot

Rogerbot is Moz's site audit crawler for Moz Pro Campaigns.

SEOIndirect

MRGbot

Search engine aimed at generating a corpus of data to be able to aggregate data in various ways.

Search Engine CrawlerRelevant

MSN

MSNBot was Microsoft's crawler for the former MSN Search engine. It visited web pages so they could be considered for MSN's search index.

Search Engine CrawlerRelevant

naver-blueno

A Naver user can trigger naver-blueno by inserting a link into a Naver service such as a blog or café.

PreviewIndirect

naverbot

Naver's web crawler (also known as Yeti) is used by Naver, South Korea's largest search engine, to crawl and index web content.

Search Engine CrawlerRelevant

Navu

Navu crawls the websites requested by their customers and prospects to train AIs for them.

AI CrawlerRelevant

Neevabot

Neevabot is the web crawler for the search engine neeva.com.

Search Engine CrawlerRelevant

netEstate Imprint Crawler

The NetEstate Imprint crawler crawls websites for public contact information.

AI CrawlerRelevant

New York Times Newsgathering

New York Times Newsgathering is a shared identity for scripts written inside the Times newsroom.

AggregatorIndirect

NewRelic Minions

New Relic Synthetic monitoring infrastructure that performs API checks and virtual browser instances to monitor websites and applications from global locations

MonitoringNot relevant

NewsBank

NewsBank aggregates licensed publisher content for schools, libraries, and government research, learning, and archiving.

ArchiverIndirect

NewsNow

The NewsNow bot is the web crawler for the news aggregator service NewsNow.

Search Engine CrawlerRelevant

NitroBot

NitroBot is service traffic from NitroPack's cloud-based site-speed optimizer.

SEOIndirect

Nostra

Nostra accelerates site speed for managed web platforms.

SEOIndirect

Notabot

Crawler to integrate Helpfeel external search engine.

Search Engine CrawlerRelevant

Novellum AI Crawl

Novellum.ai is building out tools for building agents. This MCP tool will be used by agents to crawl sites.

AI CrawlerRelevant

OAI-AdsBot

Validates the safety of web pages submitted as ads on ChatGPT; data collected is not used to train generative AI foundation models.

AdvertisingRelevantTracked by Promptwatch

OAI-SearchBot

Indexes websites for inclusion in ChatGPT's search results; does not crawl content for AI model training.

AI AssistantRelevantTracked by Promptwatch

Observer

Crawler looking for broken links to help improve world wide web as a whole.

SEOIndirect

OhDearBot

OhDearBot is a monitoring bot operated by Oh Dear that performs uptime checks, broken link detection, and mixed content scanning.

MonitoringNot relevant

Omgilibot

Omgilibot crawls public web content for Webz.io, which packages and licenses web data feeds that are commonly used to train AI models.

AI CrawlerRelevant

OMIM Litrack

OMIM Litrack is a literature tracking crawler.

Academic ResearchIndirect

OnCrawl

Enterprise SEO platform powered by the industry-leading SEO Crawler and Log Analyzer.

SEOIndirect

Online Webceo Bot

A bot associated with WebCEO, a company that provides SEO tools and services.

SEOIndirect

OnticaBot

Feed-fetching crawler that discovers and fetches public articles for curated content feeds.

Social MediaIndirect

OpenGraph.io Bot

OpenGraph. io Bot appears when an OpenGraph. io customer asks the service to inspect a URL, often so a consumer product can display a shared-link preview.

PreviewIndirect

OpenGraphXYZBot

Bot for opengraph.xyz service that generates and previews Open Graph meta tags and dynamic social media images

PreviewIndirectUnverifiable

Orlo Link Preview

Orlo Link Preview works inside Orlo's social media publishing workflow.

PreviewIndirect

Oseox

Oseox is a French SEO platform that provides.

SEOIndirect

Ozon Web Grabber

A component that serves to load previews for external and internal links.

PreviewIndirect

Pagefreezer Website Archiving

Pagefreezer Website Archiving is a compliance.

ArchiverIndirect

PanguBot

PanguBot crawls web content used to train Huawei's Pangu family of large language models.

AI CrawlerRelevantUnverifiable

ParticleNewsBot

Particle is an AI powered aggregator that collects news from many sources.

AggregatorIndirect

Payhawk Invoice Fetching Agent

Automated browser bot that fetches invoices for users from supplier websites and attaches them to their expense records.

AI AssistantRelevant

PayPal

PayPal delivers real-time event notifications for payments, subscriptions, and account updates.

WebhookNot relevant

payroll-bot

payroll-bot is an AI crawler operated by ADP, Inc. to collect publicly available legal and payroll documentation.

AI CrawlerRelevant

Perplexity-User

Handles user-initiated requests in Perplexity, accessing external content to provide real-time information; not used for automated crawling or AI training.

AI AssistantRelevantTracked by Promptwatch

PerplexityBot

Indexes websites for inclusion in Perplexity's search results; does not crawl content for AI model training.

AI AssistantRelevantTracked by Promptwatch

PetalBot

PetalBot is a web crawler operated by Huawei's Petal Search engine.

AI AssistantRelevant

PhindBot

PhindBot crawls technical and developer-focused web content to power Phind, an AI answer engine aimed at programmers.

AI AssistantRelevantUnverifiable

Pingdom Bot

Pingdom Bot is used by Pingdom's monitoring services to perform various checks on websites, including uptime and performance monitoring.

MonitoringNot relevant

Pinterest Bot

Pinterest Bot, documented by Pinterest as Pinterestbot, gathers public web content for Pinterest.

Search Engine CrawlerRelevant

Placer.ai-Crawler

Collects publicly available metadata from official websites in the open web.

AggregatorIndirect

Platebreaker

Recipe nutrition search engine. Indexes schema.org/Recipe JSON-LD. Respects robots.txt. Verified via Web Bot Auth.

Search Engine CrawlerRelevant

Polar Webhooks

Polar's webhook service delivers real-time event notifications for payment processing, including purchases, subscriptions, cancellations, and refunds.

WebhookNot relevant

Potions

The Potions bot fetches product feeds and crawls data from its customers' websites, used for e-commerce related services.

Search Engine CrawlerRelevant

prerender

It's HTML pre rendering service for SPA(Single Page Application) Website SEO.

SEOIndirect

PressEngine Bot

The PressEngine Bot verifies coverage created by video games press as genuine and their own creation.

PreviewIndirect

Pricey

Pricey collects and compares product prices, showing trends and helping users find the best time to buy deals online.

AggregatorRelevant

Promptwatch Bot

Promptwatch Bot is Promptwatch's verification crawler that validates crawler-log integrations and runs on-demand SEO and performance audits.

VerificationIndirect

ProximicBot

Proximic is Comscore's web crawler that performs contextual content analysis to help advertisers determine the best matching campaigns for a page's content.

AdvertisingNot relevantUnverifiable

PulsePoint Crawler

A web crawler used by PulsePoint, a digital advertising technology company, for content indexing and ads.txt verification.

AdvertisingNot relevant

QA.tech

The QA.tech web agent browses the website and identifies potential test cases, and executes tests against a web application

MonitoringRelevant

QStash

QStash is a platform for building event-driven applications.

WebhookNot relevant

QualifiedBot

Bot crawls customer websites to provide information to customer hosted chatbots.

AI CrawlerRelevant

Quantcastbot

Quantcast Bot is a web crawler used for advertisement quality assurance and to understand page content for Interest-Based Audiences.

AdvertisingNot relevant

Quartr Crawler

Quartr uses a crawler to obtain and deliver investor relations material.

AggregatorIndirect

Qwantbot

Crawls and indexes web content for Qwant search engine.

Search Engine CrawlerRelevant

Rakuten Image extraction bot

Rakuten's image extraction bot fetches product images from sites that partner with Rakuten.

AggregatorIndirect

Razorpay-Webhook

Razorpay-Webhook delivers asynchronous payment notifications to a merchant's configured URL.

WebhookNot relevant

RDTvlokipBot

Official crawler for RDTvlokip Search, an independent French search engine. Respects robots.txt.

Search Engine CrawlerRelevant

Redirect pizza destination monitor

redirect.pizza's destination monitor ensures that the redirect destination URLs are reachable.

MonitoringNot relevant

Retool

Retool is the platform user agent attached to outbound HTTP requests from Retool.

AI AssistantRelevant

Revvim

RevvimGort is listed as Revvim's crawler for examining customer websites and finding SEO opportunities.

SEOIndirect

RyeBot

Powers automated checkout on behalf of shoppers with explicit consent.

AI AssistantRelevant

Sanity Webhooks

Sanity's webhook service that delivers real-time event notifications for content changes and other events.

WebhookNot relevant

Sansec Security Monitor

Sansec Security Monitor is a web crawler that monitors online stores for malicious code, data breaches, and digital skimming attacks.

MonitoringNot relevant

SBIntuitionsBot

SBIntuitionsBot is a crawler operated by SB Intuitions Corp. that collects web data for AI development and information analysis.

AI CrawlerRelevantUnverifiable

ScreamingFrogBot

Screaming Frog SEO Spider is a website crawler used by SEO professionals for site audits and technical SEO analysis.

SEOIndirectUnverifiable

Screpy

SEO Checker for Screpy Bot that SEO.

SEOIndirect

SE Ranking Backlinks

SE Ranking's backlink analysis crawler that discovers and analyzes backlink profiles for SEO research and competitive analysis.

SEOIndirect

SearchAtlas Bot

Bot used to evaluate customer's websites and provide SEO optimization strategy.

SEOIndirect

SeekportBot

SeekportBot builds the independent web index used by the Seekport search engine.

Search Engine CrawlerRelevant

Selectika AI

Selectika AI enrichment compute vision for Fashion.

AI CrawlerRelevant

SemanticScholarBot

SemanticScholarBot looks for academic PDFs on selected domains for Semantic Scholar.

AI CrawlerRelevantUnverifiable

Sentry Uptime Monitoring Bot

Sentry's Uptime Monitoring Bot performs health checks on configured URLs to monitor the availability and reliability of web services.

MonitoringNot relevant

SEO Audit Check Bot

SEO audit check bot is likely an automated tool within the WebCEO platform that performs comprehensive SEO audits on websites.

SEOIndirect

seo4ajax

seo4ajax belongs to a prerendering service for single-page applications.

SEOIndirect

Seobility

Seobility is a browser-based online SEO software that helps you improve your website’s search engine rankings.

Search Engine CrawlerRelevant

SerpstatBot

SerpstatBot is the Serpstat bot collects data for Serpstat's Backlink Analysis tool.

SEOIndirect

ServerHunterSpider

Hosting providers see ServerHunterSpider because Server Hunter imports plan information for its comparison site.

AggregatorIndirect

SeznamBot

SeznamBot is the web crawler operated by Seznam.cz, the leading Czech search engine.

Search Engine CrawlerRelevant

ShapBot

Crawls and indexes web content to power Parallel's search and content extraction APIs for AI applications.

AI CrawlerRelevant

Shopify Webhooks

Shopify webhooks are useful for keeping your app in sync with Shopify data, or as a trigger to perform an additional action after that event has occurred.

E-commerceRelevantUnverifiable

Shortwave Image Fetcher

An email client that proxies all images found in HTML emails from to protect end customer's IP address and connection private.

PreviewIndirect

SISTRIX Optimizer Uptime

A homepage request arriving once every minute may come from the uptime feature in a SISTRIX Optimizer project.

MonitoringNot relevantUnverifiable

Site24x7

Site24x7 Bot is used by Site24x7's monitoring services to perform various checks on websites, including uptime and performance monitoring.

MonitoringNot relevant

Sitebulb

Sitebulb is a desktop and cloud-based website crawler used by SEO professionals for technical SEO audits.

SEOIndirectUnverifiable

SiteGuru

SiteGuru is an SEO auditing tool that crawls.

SEOIndirect

Siteimprove Crawl

Siteimprove content suite (i.e. Quality Assurance, Accessibility, Policy, and SEO). Crawls run on ports are 80 for HTTP and 443 for HTTPS.

SEOIndirect

SiteSearch360

SiteSearch360 crawls websites whose owners use the Site Search 360 hosted search box.

Search Engine CrawlerRelevant

Skroutz ImageBot

Skroutz ImageBot to fetch the individual product images.

AggregatorIndirect

Skype

Skype's URI Preview services fetches a page preview when someone posts a URL in a Skype message.

PreviewIndirect

Slack-ImgProxy

Slack-ImgProxy is a bot operated by Slack that fetches and caches images posted in Slack channels.

PreviewIndirectUnverifiable

Slackbot

Slackbot is Slack's default, general-purpose bot that handles various API requests and integrations.

PreviewIndirectUnverifiable

SlackLinkExpandingBot

Slackbot Link Expanding is a bot operated by Slack that fetches metadata from shared links to create rich previews.

PreviewIndirectUnverifiable

SMTnet PM Bot

SMTnet PM Bot visits partner company websites so their content can appear in SMTnet's on-site search.

Search Engine CrawlerRelevant

SnapchatAdsBot

SnapchatAdsBot is a crawler operated by Snapchat that verifies and analyzes websites for their advertising platform.

AdvertisingNot relevantUnverifiable

SnapURLPreviewBot

SnapURLPreviewBot is a crawler operated by Snap Inc. that analyzes and generates previews of URLs shared on Snapchat and other Snap platforms.

AnalyticsIndirectUnverifiable

Sogou Web Spider

Sogou Web Spider visits pages for the search index at sogou. com.

Search Engine CrawlerRelevantUnverifiable

sourcedash

sourcedash is the crawler identity for Source Dashboard, a UK property search and management platform.

Search Engine CrawlerRelevant

Spyglasses

Spyglasses accesses site content to assist with AEO capabilities. Learn more at https://www.spyglasses.io.

SEOIndirect

Stably

Stably is a QA testing bot that users run to E2E test their websites for functionality testing and protecting user flows against regressions.

MonitoringNot relevant

Statabot

Statabot searches for stata.toc files and indexes their contents.

AggregatorIndirect

StatistikAustria

Bot to collect product prices for the official consumer price index of Austria.

AggregatorIndirect

StatsDroneBot

The StatsDrone affiliate marketing statistics scraping and aggregating tool.

AggregatorIndirect

StatusCake Page Speed

StatusCake Page Speed monitors your page load and render speeds.

MonitoringNot relevant

StatusCake SSL Monitoring

StatusCake SSL monitors your website certificates for common issues

MonitoringNot relevant

StatusCake Uptime

StatusCake monitors the uptime of your website.

MonitoringNot relevant

Steam Chat

The Steam Chat bot fetches previews of URLs shared within the Steam client's chat feature.

PreviewIndirect

Stripe Webhooks

Stripe's webhook service that delivers real-time event notifications for payment processing and account updates.

WebhookNot relevant

Stripebot

Crawls Stripe merchant websites to collect data for service delivery and financial regulatory compliance.

AnalyticsNot relevant

Strivve Automation

Strivve Web Bot Authentication Agent.

AI AssistantRelevant

svix

svix is a webhook service for sending events to webhooks.

WebhookNot relevant

TangibleeBot

TangibleeBot reads retailer product pages for Tangiblee's visualization and virtual try-on services.

E-commerceRelevantUnverifiable

Telegram Bot

TelegramBot crawls websites to render a link preview when people send a message containing a URL in the Telegram messaging service.

PreviewIndirect

TermlyBot

Crawls websites to detect and categorize cookies set by first and third parties.

MonitoringNot relevant

Terracotta

The Terracotta bot scrapes websites for use in generating indices for serving searches using Ceramic's search product.

Search Engine CrawlerRelevant

TikTokSpider

TikTokSpider is a TikTok and ByteDance web crawler used to index and analyze content outside the TikTok platform.

AI CrawlerRelevantUnverifiable

Timpibot

Timpibot crawls the web to build Timpi's decentralized search and data index, which is used to supply training and grounding data for AI applications.

AI CrawlerRelevantUnverifiable

TNOThesisCrawler

Research crawler by TNO collecting publicly available academic theses from university websites for the DIAMONDS platform.

Academic ResearchIndirect

Toutiao

Toutiao is ByteDance's automated news aggregation bot collecting content across web platforms.

Search Engine CrawlerRelevant

Trellis-Services

Critical CSS Generator to Optimize Websites.

SEOIndirect

Trendiction Bot

Trendiction's web crawler that discovers and collects public web data for their social media monitoring and media intelligence platform.

AnalyticsNot relevantUnverifiable

TrustedSite

Crawl a customer website to perform basic validation.

Academic ResearchIndirect

TTD-Content

TTD-Content is a crawler operated by The Trade Desk that verifies content and quality of ad placements for their demand-side platform.

AdvertisingNot relevantUnverifiable

Tumblr

A Tumblr request usually starts with a post author pasting a URL.

PreviewIndirect

TurnitinBot

TurnitinBot collects online material for Turnitin's plagiarism prevention service.

Academic ResearchIndirect

Twilio Knowledge

Twilio Knowledge fetches website content when a Twilio customer adds a web source to a knowledge base.

AI CrawlerRelevantUnverifiable

Twilio Proxy

Twilio's proxy service that handles communications between end-users and applications through Twilio's programmable voice and messaging platform.

WebhookNot relevantUnverifiable

TwinAgent

Automate complex operations end-to-end.

AI AssistantRelevant

Twitterbot

Fetches content for shared links on X/Twitter to generate rich previews.

PreviewIndirect

upday

upday is a news aggregator app, and its bot crawls news sources. It collects and indexes articles to be recommended to users on its platform.

AggregatorIndirect

Updown.io

Performs uptime and performance checks on websites.

MonitoringNot relevant

Uptime Robot

Uptime Robot is a platform for monitoring and alerting on your applications.

MonitoringNot relevant

UsercentricsBot

UsercentricsBot is operated by Usercentrics GmbH to scan websites for data processing services and third-party technologies.

AnalyticsNot relevantUnverifiable

v0bot

v0bot is listed as an AI crawler for v0 services, with v0bot as its stable user-agent token. That is the full purpose stated in the supplied bot record.

AI CrawlerRelevant

Velen Public Web Crawler

Velen Public Web Crawler collects public web content for Webz.io's data feeds, which are licensed for AI training, market intelligence, and monitoring.

AI CrawlerRelevant

Vemetric Favicon Bot

Fetches favicons from websites in the highest quality available.

PreviewIndirect

Vercel build container

System-initiated requests made from Vercel's build container during a build

PreviewIndirect

Vercel Favicon Bot

Vercel Favicon Bot is categorized as a preview client, and its name indicates a narrow job: obtaining a site's favicon for use in a Vercel interface or preview.

PreviewIndirect

Vercel Screenshot Bot

Vercel Screenshot Bot is listed as a preview client whose name points to page-image capture.

PreviewIndirect

vercelflags

vercelflags is a monitoring operated by its operator.

MonitoringNot relevant

verceltracing

verceltracing is a monitoring operated by its operator.

MonitoringNot relevant

videootv Bot

videootv Bot works on sites that publish through Digital Green's native advertising module.

AggregatorIndirect

Visually.io Shopify Editor

Shopify theme editor alternative for live, real-time store editing via a secure iframe and controlled proxy.

AI AssistantRelevant

W3 Validator Services

W3C provides various free validation services that help check the conformance of Web sites against open standards.

SEOIndirect

WARDBot

WARDBot tracks URL status codes, helping users monitor the availability of web pages they have added to the monitoring list.

AI CrawlerRelevant

WebSpiderMount

Job wrapping data processor handling jobs distribution from employer websites to multiple endpoints, like job boards, advertisement platforms, job alerts etc.

AggregatorIndirect

WindowsForum-AI

Technology news aggregation bot for WindowsForum.com.

Search Engine CrawlerRelevant

WMF Citoid

Citoid handles the automatic citation lookup in Wikimedia's VisualEditor.

PreviewIndirect

WMF Zotero Translation Server

WMF Zotero Translation Server sits behind Citoid rather than facing Wikimedia editors directly.

PreviewIndirect

WordCountBot

WordCountBot analyzes website word count based on public pages. All words belonging to public pages and included in HTML source code.

PreviewIndirect

XY Archive Compliance Bot

XY Archive Compliance Bot visits websites selected by customers with recordkeeping requirements. The operator describes a two-part job.

ArchiverIndirect

Yahoo Ad Monitoring

A landing page listed in a Yahoo advertisement may be opened by Yahoo Ad Monitoring.

AdvertisingNot relevant

Yahoo Japan SEO Crawler

Yahoo Japan search engine crawler for SEO analysis.

SEOIndirect

Yahoo Link Preview

Yahoo Link Preview's bot fetches data from URLs shared on Yahoo platforms.

PreviewIndirect

Yahoo! JAPAN

This Yahoo! JAPAN entry uses J-DLC as its stable token.

Search Engine CrawlerRelevant

Yahoo! Slurp

Yahoo! Slurp is the web crawler (robot) used by Yahoo! Search to discover and index web pages for its search engine.

Search Engine CrawlerRelevant

YahooCacheSystem

YahooCacheSystem caches website contents as part of the Yahoo! Search Service.

PreviewIndirect

YahooMailProxy

Yahoo Mail Proxy is a content fetch proxy that retrieves the page content of URLs that are embedded within emails sent to Yahoo Mail users.

PreviewIndirect

YandexAdditional

YandexAdditional is Yandex's dedicated user agent for content used by YandexGPT and other generative AI features.

AI CrawlerRelevant

Yandexbot

YandexBot is a web crawler operated by Yandex, a major Russian search engine.

Search Engine CrawlerRelevant

Yeti

Yeti is the web crawler for Naver, a South Korean search engine. It indexes websites to provide search results and power other services on the Naver platform.

Search Engine CrawlerRelevant

YGS Group Falconer Scraper

YGS Group Falconer Scraper supports Falconer, a coverage discovery product presented by YGS Content Licensing.

AI CrawlerRelevant

YisouSpider

The directory cannot independently verify YisouSpider's identity.

Search Engine CrawlerRelevantUnverifiable

YouBot

YouBot crawls and indexes web pages to power the You.com AI search engine and its cited answers.

AI AssistantRelevant

Zoombot

ZoomBot is SEOZoom's web crawler that builds.

SEOIndirect

ZoomInfo

Zoominfobot is an indexing robot for a web search engine, similar to Google. Created by Zoom Information Inc.(www.zoominfo.

Search Engine CrawlerRelevant

ZumBot

ZumBot is a web crawler that indexes webpages for Zum Open Internet Search.

Search Engine CrawlerRelevant

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard