AI Bots & Web Crawlers Directory
435 bots
AdagioBot
Adagiobot is a web crawler that analyzes websites for advertising demand optimization, helping publishers maximize revenue through real-time bidding analysis and performance insights.
AddSearchBot
As AddSearch adds content from your site to the search, the AddSearch bot gets counted as traffic by most analytics software.
AddThis
The AddThis bot crawls websites to gather and update content for its website marketing tools. These tools include features like social sharing buttons and content recommendation widgets.
AdIdxBot
AdIdxBot is the crawler used by Bing Ads for quality control of ads and their destination websites. It has multiple user agent variants including desktop, iPhone, and Windows Phone versions.
Adsense
The AdSense crawler visits participating sites in order to provide them with relevant ads.
adsnaver
Naver's ad crawler that periodically visits registered ad landing pages to collect on-page content for effective ad matching and ranking. It ignores robots.txt for URLs registered in the ad system.
Adyen Webhook
Adyen’s webhooks (Notification API) send encrypted, real-time HTTP callbacks for key payment and account events, automating order fulfillment, settlement reconciliation, and risk-management workflows.
Agency Analytics Crawler
A web crawler by Agency Analytics that allows their clients to check their own sites for SEO.
AGI Agent
The AGI Agent is a productivity assistant that takes actions and makes purchases on behalf of users.
AhrefsBot
Powers the database for both Ahrefs, a marketing intelligence platform, and Yep, an independent, privacy-focused search engine.
AhrefsSiteAudit
Powers Ahrefs’ Site Audit tool. Ahrefs users can use Site Audit to analyze websites and find both technical SEO and on-page SEO issues.
AI Search
Cloudflare AI Search is a managed service that lets you connect your data and easily build AI-powered search.
AI Search External
Cloudflare AI Search is a managed service that lets you connect your data and easily build AI-powered search.
AI2Bot
AI2Bot is operated by the Allen Institute for Artificial Intelligence (Ai2) to crawl the web for content to train open-source AI models.
aiHitBot
aiHitBot collects and maintains historical information about companies.
AirsoftdbBot
Airsoftdb is a search engine for airsoft prod.
Algolia
The Algolia Crawler extracts content from your site and makes it searchable.
All Africa Crawler
AllAfrica Global Media produces, aggregates and distributes news from across Africa, relying on agreements with more than 140 news organizations and over 500 other institutions and individuals.
Alli AI Bot
Alli AI bot crawls customer websites to generate SEO recommendations.
Alphalens Bot
Indexes companies and their offerings to power more effective business discovery.
Amazon AdBot
Amazon AdBot is a crawler used by different advertising services at Amazon to determine a website's content in order to provide relevant and appropriate advertising.
Amazon Bedrock AgentCore Browser
AgentCore Browser provides a secure, cloud-based browser that enables AI agents to interact with websites.
Amazon Bedrock Bot
Amazon Bedrock Bot fetches web pages that customers add as data sources to Amazon Bedrock knowledge bases, so Bedrock-powered assistants can answer with that content.
Amazon Kendra
Amazon Kendra is a managed information retrieval and intelligent search service that uses natural language processing and advanced deep learning model.
Amazon Product Discovery
Amazon's web crawler used to collect publicly available product details from Amazon Selling Partner websites to help improve the accuracy and completeness of product information on Amazon.
Amazon Q
Amazon Q Business is a generative artificial intelligence (generative AI)-powered assistant that you can tailor to your business needs.
Amazon Route 53 Health Check Service
Amazon Route 53 Health Check Service
Amazon Seller Initiated Listing
Amazon's web crawler that helps sellers succeed by giving them the option to provide a URL to a website and create high-quality product pages in Amazon's store.
Amazonbot
Amazonbot is Amazon's web crawler used to improve our services, such as enabling Alexa to more accurately answer questions for customers.
Amzn-SearchBot
Amzn-SearchBot crawls for improving Amazon search experiences (Alexa, Rufus).
Anchor Browser
The Web Browser for AI Agents.
Andibot
Andibot gathers web content for Andi, a conversational AI search assistant that answers questions with summaries and sources.
Apify Website Content Crawler
Crawl websites and extract content to feed AI apps. Convert web data to Markdown or HTML, download files, and more.
APIs-Google
Crawling preferences addressed to the APIs-Google user agent affect the delivery of push notification messages by Google APIs.
Apple App Site Association
The Apple App Site Association is used to support "Universal Links" that can open in native iOS apps.
Apple Podcasts
Apple Podcasts crawler that only accesses URLs associated with registered content on Apple Podcasts. Does not follow robots.txt.
Applebot
Applebot powers search features in Apple's ecosystem (Spotlight, Siri, Safari) and may be used to train Apple's foundation models for generative AI features.
Applebot-Extended
Applebot-Extended is a control token that lets site owners opt out of having content crawled by Applebot used to train Apple's foundation models and Apple Intelligence features.
Arena Bot
Link preview bot for Arena.im's live engagement platform. Fetches page previews when URLs are shared in Arena's live chat, live blog, and community tools.
Arquivo Web Crawler
Web crawler archives the Portuguese web.
Artemis Web Crawler
Artemis is a calm web reader with which you can follow websites and blogs.
Artsdata Crawler
Web crawler that collects publicly available LOD for arts and culture in Canada.
Atlassian Jira Webhooks
Delivers webhook notifications from Jira Cloud when issues, projects, or other resources change.
Atlassian Rovo
Crawls and indexes web content for Atlassian Rovo's AI-powered search, chat, and agents.
atlassian-bot
atlassian-bot is a crawler for custom 3P websites that indexes data for rovo search.
Attracta
The Attracta bot is analyzes user website content as part of Attracta's SEO services.
Authory
The Authory bot visits websites to back up articles on behalf of journalists and other writers who use the service.
Awario Bot
Awario's web crawler used to discover and collect new and updated web data for their social media monitoring and brand mention tracking platform.
Awario RSS Bot
One of Awario's primary web crawlers specialized in collecting RSS feed data.
Awario Smart Bot
One of Awario's primary web crawlers that discovers and collects new and updated web data.
Baidu ADS Server Proxy
Baidu's scrubbing proxy.
BaiduSpider
Baiduspider is Baidu’s web crawler that indexes websites for inclusion in its Chinese-market search results.
Barkrowler
Barkrowler is Babbar's web crawler that fuels and updates their graph representation of the web, providing SEO tools for the marketing community.
BestChange Bot
The BestChange bot downloads exchange rate information from 600 websites every 5 seconds.
Better Stack
Better Stack is a platform for monitoring and alerting on your applications.
Bibliothèque nationale de France Crawler
Bibliothèque nationale de France's mission is to collect, catalog, preserve, enrich and communicate the national documentary heritage.
Big Sur AI
Big Sur AI Crawler, crawlers users websites to enable AI-infused experiences.
Bing Preview
BingPreview generates page snapshots for Bing. Note that BingPreview has desktop and mobile variants.
Bingbot
Bingbot is Microsoft's web crawler used for indexing websites for Bing Search.
BLEXBot
SEO PowerSuite Link Explorer (webmeup.com) is the world's freshest backlink index, and the primary source of backlink-related data for the SEO PowerSuite tools.
Bluesky Link Preview Service
Bluesky social pulls links in advance to render webpage previews.
BoardGamePrices Bot
Price comparison site for board games. Need to crawl store pages for participating stores. All stores give permission to be crawled.
BorderxBot
E-commerce product crawler operated by Borderxlab.
Botify
SiteCrawler, part of the Botify Analytics suite, gives enterprise SEO teams the power to evaluate the structure and content of their websites just like a search engine.
Brandwatch
The Magpie Crawler indexes content for its soical media monitoring solution.
BraveBot
BraveBot crawls and indexes web pages for the Brave Search index, which also grounds answers from Brave's Leo AI assistant.
Brightbot
Brightbot is Bright Data's crawler layer that monitors the health of websites and enforces ethical web data collection.
Browserbase
Runs headless browser automation on behalf of Browserbase customers for web scraping, form submission, and testing.
Buffer Link Preview Bot
Helps Buffer users create better social media posts by generating rich previews when they share links
Bytespider
Bytespider is ByteDance's web crawler used to gather training data for their AI large language models.
CaliberBot
Caliperbot crawls Conductor clients' and prospects' websites for HTML feature extraction to power Content Analytics features within our Searchlight web application.
Capital One Bot
Capital One Bot crawls dealer websites for getting the usage information for Capital One lead navigator button.
CCBot
CCBot is operated by the Common Crawl Foundation to crawl web content for AI training and research.
CensysInspectBot
Censys Inspect is a web crawler operated by Censys that performs internet-wide scanning to discover, monitor, and analyze publicly accessible devices and services.
Channel3Bot
Crawls product detail pages to index content for AI-powered product discovery, routing shoppers to original websites.
ChatGPT agent
Agent that can use its own browser to perform tasks for user.
ChatGPT-Operator
Handles user-initiated requests from ChatGPT operator accessing external content; not used for automated crawling or AI training.
ChatGPT-User
Handles user-initiated requests in ChatGPT, accessing external content to provide real-time information; not used for automated crawling or AI training.
Chathive crawler
Chathive Crawler enables our customers to crawl their own websites, which is used to power their AI assistants.
Checkly
Checkly is a platform for monitoring and alerting on your applications.
Chrome Lighthouse
PageSpeed Insights (PSI) reports on the user experience of a page on both mobile and desktop devices, and provides suggestions on how that page may be improved.
Chrome Privacy Preserving Prefetch Proxy
Chrome's Privacy Preserving Prefetch Proxy service that fetches /.well-known/traffic-advice to enable privacy-preserving prefetch hints.
CitibotSiteCrawler
CitibotSiteCrawler collects public data from government websites to power Citibot’s AI civic engagement tools.
ClarityBot
ClarityBot is seoClarity's web crawler that performs technical SEO audits, analyzes content, and monitors website performance.
Claude Web
Claude Web is a legacy Anthropic crawler that fetched recent web content for the Claude assistant. Its behavior has largely been folded into ClaudeBot and Claude-User.
Claude-SearchBot
Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses.
Claude-User
Claude-User supports Claude AI users. When individuals ask questions to Claude, it may access websites using a Claude-User agent.
ClaudeBot
ClaudeBot helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training.
ClearscopeBot
Clearscope is an AI-driven SEO content optimization platform developed by Mushi Labs.
Cledara SaaS Management Agent
Cledara’s agent automates customer-approved SaaS admin tasks, including invoice collection and user management.
Cloudflare Browser Run
Renders web pages in headless browsers for Cloudflare customers. Used for browser automation (screenshots, PDF generation, content extraction, etc.) and for AI agents to interact with the web.
Cloudflare Crawler
The Cloudflare Crawler is a well-behaved crawler that retrieves web content. By default, it self-identifies as a bot, honors robots.txt directives, and cannot bypass CAPTCHAs or bot protection.
CMU-Cylab
An academic research bot, conducting research on patterns in counterfeit document sales over the internet.
Cốc Cốc
Coccocbot scrapes websites that are request from the Vietnamese search engine Coc Coc.
cognitiveSEO Crawler
CognitiveSEO is an SEO toolset that crawls the web and analyzes links.
Cohere AI
Cohere AI collects publicly available web text that helps train and refine Cohere's large language models for enterprise generative AI.
ContentKingBot
ContentKing (now Conductor Website Monitoring) is a website monitoring tool that continuously audits websites to help improve their performance and visibility.
Cookiebot
Cookiebot automates compliance with cookie laws and helps you manage your cookie consent preferences.
CookieScript
A cookie scanning bot that examines websites for cookie usage to help maintain GDPR and other privacy regulation compliance.
CoostiBot
CoostiBot crawls merchant websites for Coosti.dk, a Danish price comparison platform. Respects robots.txt.
Cotoyogi
Cotoyogi is a web crawler operated by the Center for Research and Development on Data Lake, ROIS-DS (Research Organization of Information and Systems - Data Science) for collecting Japanese language...
Coveobot
Coveobot is a crawler operated by Coveo that indexes content for enterprise search, recommendations, and generative experience platforms.
Crawlson
Crawlson is a search engine crawler for the crawlson.com search engine.
creobot
This bot download assets from the origin server.
CriteoBot
CriteoBot is a crawler operated by Criteo that analyzes web content to serve relevant contextual ads.
Customer.io webhooks
Customer.io's webhook service for event-driven marketing automation and customer data platform.
CuteStat
Indexing content to be included on our search results and insights.
Cxense
The Cxensebot performs SEO monitoring and analysis of customer webpages.
Cybaa Agent
Performs user-initiated security checks on behalf of Cybaa customers, validating security headers, TLS/SSL configuration, and other domain-specific security controls to ensure website compliance and...
Dash0 Synthetic Monitoring
Dash0's Synthetic Monitoring provides proactive, automated insights into the availability and performance of your websites and APIs.
Datadog Synthetic Monitoring Robot
Datadog's automated monitoring service that performs synthetic tests to verify website availability and performance.
DataForSEO Site Auditor
DataForSEO is using RSiteAuditor to scan websites for critical on-site SEO errors and provides aggregated data in a structured form to its customer through a RESTful API.
DataForSeoBot
DataForSeoBot is a backlink checker bot operated by DataForSEO that crawls websites to build and maintain their backlink database.
Dataprovider.com
Dataprovider.com indexes the web and structures the data.
Daum
Korean search engine crawler.
DeepSeek Bot
DeepSeek Bot crawls web content used to train and improve DeepSeek's generative AI models.
Detectify
Detectify is a web security scanner that performs automated security tests on web applications and attack surface monitoring.
Devin
Devin is a collaborative AI teammate built to help ambitious engineering teams achieve more.
Diffbot
Diffbot crawls and structures web pages into a knowledge graph that is sold for AI training, retrieval, and data enrichment.
DigitalOceanUptimeBot
DigitalOcean Uptime is a monitoring service that checks the health of any URL or IP address.
Direqt Anomura
Anomura is Direqt’s search crawler, it discovers and indexes pages their customers websites.
Discord Bot
Discord's link preview bot that crawls URLs shared in Discord chats to generate rich previews.
DotBot
DotBot is a web crawler operated by Moz (formerly SEOmoz) that collects data for their Link Explorer tool and Links API.
DuckAssistBot
DuckAssistBot is a web crawler for DuckDuckGo Search that crawls pages in real-time for AI-assisted answers, which prominently cite their sources. This data is not used in any way to train AI models.
DuckDuckBot
DuckDuckBot is a web crawler for DuckDuckGo. DuckDuckBot’s job is to constantly improve search results and offer users the best and most secure search experience possible.
EasyScan
Automated scanning service that reviews online content on behalf of end users to identify potential legal issues.
Echobox Bot
We scrape full article/page content to ensure we can optimally automate the content distribution for the digital publishers we work with. Every single article a publisher releases will get scraped approx.
Element451Bot
Element451Bot is the Knowledge Hub crawler for Element451, indexing pages so its AI assistant can answer questions for students.
Embedly
Embedly link preview service (operated by Medium) that fetches metadata and embeds for URLs.
eMoney Advisor
Collects raw financial data that can later be used for financial planning and analysis.
Epivoz Crawler
News aggregator needs to crawl news/blog articles to generate short summaries for page preview of attributed links.
EzoicBot
Ezoic is a technology platform for digital publishers. You can learn more about what Ezoic does here.
Facebook Webhooks
Facebook's webhook service that delivers real-time event notifications for Meta platform events and changes.
FacebookBot
FacebookBot crawls public web content that Meta may use to improve language models and other AI products. It is distinct from the link-preview fetcher facebookexternalhit.
FacebookExternalHit
Fetches content for shared links on Meta platforms to generate rich previews.
Factset_spyderbot
Factset uses a Python Selenium Crawler for web scraping to deliver reliable, current financial data.
FalBot
fal.ai's webhook service that delivers asynchronous notifications for AI model processing and generation tasks.
Fastmail Bot
Fastmail fetch and image proxy bot.
Fedicabot
.
FedReporter Bot for FFIEC
Bot to download data from the ffiec, Active/Closed/Branches File, Holding Company Data, 002 Data.
FirmlyAI Bot
FirmlyAI Bot navigates e-commerce websites for agentic commerce workflows.
FishBot
FishBot crawls webpages to deliver Open Source AI for All.
FlipboardProxy
Fetches and prepares website content for presentation in the Flipboard application.
FlyingPress
FlyingPress bot optimizes pages by generating critical CSS, detecting above-fold images, and delaying 3rd-party scripts.
FN Legislative and Regulatory Bot
FN Bot makes targeted rate-limited requests to a variety of publically available legislative and regulatory sources.
Freespoke
Freespoke is a search engine that believes in free speech and shows you all viewpoints.
Funnelback
Funnelback is an enterprise search platform, and its crawler indexes content from an organization's websites and data repositories. This powers the organization's internal search function.
GeedoProductSearchBot
GeedoProductSearch is a web crawler operated by Geedo SIA that indexes product information from e-commerce websites.
GeedoShopProductFinder
GeedoShopProductFinder is the automated crawl.
Gemini Deep Research
Gemini Deep Research is Google's AI-powered research tool that performs comprehensive multi-step research on complex topics, analyzing web content to provide detailed insights and answers.
GhostExplore
An aggregator service for websites powered by ghost.org.
Gigabot
Gigablast is the only non-Big Tech search engine in the U.S. that uses its own search index and algorithms.
GitHub Camo
GitHub's image proxy service
GitHub Hookshot
GitHub's webhooks for events like push, pull request, etc.
Google AdMob Reward Verification
Sends server-side verification callbacks to confirm users completed rewarded ad views.
Google Ads Creatives Assistant
Fetches website content for Google Ads creative generation and enhancement tools.
Google AdsBot
Google AdsBot is Google's web crawler for quality control of Google Ads.
Google Association Service
Verifies associations between apps and websites for Digital Asset Links.
Google Business Link Verification
Verifies that business links in Google Business Profile are accessible and return valid HTTP status codes.
Google Docs
Fetches images and page content when users insert links into Google Docs.
Google Feedfetcher
Feedfetcher is used for crawling RSS or Atom feeds for Google News and PubSubHubbub.
Google Image Proxy
Google's image caching proxy service used by Gmail and other Google services to cache and serve images.
Google Images
The Google Images bot is the search engine crawler for Google Images Search.
Google NotebookLM
Google NotebookLM fetches web sources that a user adds to a notebook so the assistant can summarize, answer questions, and cite them. Because fetches are user-initiated, it may bypass robots.txt.
Google PageRenderer
Upon user request, Google Page Renderer fetches and renders web pages.
Google Publisher Center
Google Publisher Center fetches and processes feeds that publishers explicitly supplied for use in Google News landing pages.
Google Read Aloud
Upon user request, Google Read Aloud fetches and reads out web pages using text-to-speech (TTS).
Google Scholar
Google Scholar uses a bot to crawl and index scholarly literature from academic publishers, repositories, and university websites. This populates its academic search engine.
Google Site Verifier
Google Site Verifier fetches Search Console verification tokens.
Google StoreBot
Crawling preferences addressed to the Storebot-Google user agent affect all surfaces of Google Shopping (for example, the Shopping tab in Google Search and Google Shopping).
Google Videos
The Google Videos bot is the search engine crawler for Google Video Search.
Google-AdWords-Express
Google-AdWords-Express is a bot for a Google Ads product aimed at small businesses. It crawls advertiser websites to assist with ad creation and to verify site information.
Google-Adwords-Instant
Fetches advertiser landing pages when triggered by user actions in the Google Ads platform.
Google-Agent
Google-Agent navigates the web and performs actions upon user request, used by agents hosted on Google infrastructure such as Project Mariner.
Google-CloudVertexBot
Crawling preferences addressed to the Google-CloudVertexBot user agent affect crawls requested by the site owners' for building Vertex AI Agents. It has no effect on Google Search or other products.
Google-Display-Ads-Bot
Verifies site eligibility during the AdSense approval process.
Google-Extended
Google-Extended is a standalone product token that web publishers can use to manage whether their sites help improve Gemini Apps and Vertex AI generative APIs, including future generations of models...
Google-Flights-Search-Spackle
Google's user-triggered fetcher for flight search. Fetches airline and travel site content on behalf of end-user searches to compile flight pricing and availability data.
Google-InspectionTool
Crawling preferences addressed to the Google-InspectionTool user agent affect Search testing tools such as the Rich Result Test and URL inspection in Search Console.
Google-Safety
The Google-Safety user agent handles abuse-specific crawling, such as malware discovery for publicly posted links on Google properties. As such it's unaffected by crawling preferences.
google-xrawler
Google's user-triggered fetcher for merchant product feeds. Fetches XML product feeds from e-commerce sites for syncing with Google Merchant Center.
Googlebot
Crawling preferences addressed to the Googlebot user agent affect Google Search (including Discover and all Google Search features), as well as other products such as Google Images, Google Video,...
GoogleOther
Crawling preferences addressed to the GoogleOther user agent don't affect any specific product.
GoogleStackdriverMonitoringBot
GoogleStackdriverMonitoringBot is operated by Google Cloud to perform uptime checks and monitor availability of services.
GPT-Actions
Enables ChatGPT to interact with external APIs and retrieve real-time information from the web in response to user-initiated requests; allows access to up-to-date content without being used for...
GPTBot
Crawls web content to improve OpenAI's generative AI models and ChatGPT; respects 'robots.txt' directives to exclude sites from training data.
Grok DeepSearch
Grok DeepSearch performs multi-step research across the web to answer complex Grok queries with cited sources.
Grok Search
Grok Search fetches web pages in real time to power Grok's search and answer features inside X and the Grok apps.
GrokBot
GrokBot is xAI's crawler used to gather web content for training the Grok family of models. xAI publishes limited documentation for it.
GTmetrix
GTmetrix provides metrics and insights for your site's loading speed and performance.
HarkBot
Hark is building the most advanced personal intelligence in the world.
HelloWork
HelloWork is a French job board, and its bot aggregates job listings for its platform. It crawls company career pages and other sources to collect this information.
Henry Shopping Agent
Executes checkout via browser automation using a user's card and signed mandate.
HetrixTools Uptime Monitoring Bot
HetrixTools Uptime Monitoring Bot is used by HetrixTools's monitoring services to perform various checks on websites, including uptime and performance monitoring.
HEY Email Privacy Proxy
HEY email stops spy pixels and prevents user IP tracking by proxying all HTML email images, fonts, and external assets.
HIFIBot
Short Description: HIFI is a financial services company for musicians and professional creators. HIFI acts as an agent on behalf of its clients to automate the retrieval and processing of royalty earnings statements.
Hookdeck
A reliable Event Gateway for event-driven applications
HubSpot Page Fetcher
When posting to LinkedIn from Hubspot, images need to be pulled through to LinkedIn when published. The crawler performs this function.
Huckabuy Bot
Huckabot is Huckabuy’s main crawler which is utilized by almost all of Huckabuy’s products.
Hydrozen
Hydrozen is a tool for monitoring availability of your websites, Cronjobs, APIs, Domains, SSL etc.
Hype Machine
Since 2005, Hype Machine monitors music publications/blogs for posts about new artists and builds playlists using this metadata for listeners.
i-search-crawler
i-search-crawler is a web crawler operated by.
IASBot
IAS (Integral Ad Science) crawler, formerly known as AdmantX, is used for analyzing web content to ensure brand safety and suitability for advertisers.
iAskBot
iAskBot crawls and indexes web content to power iAsk.ai, an AI question-answering search engine.
IbouBot
IbouBot is the crawler of the Ibou Search Engine.
ICC Crawler
ICC-Crawler automatically crawls the Internet and collects web pages.
Idealo
Germany's largest price comparison service.
Iframely
Fetches your page metadata to generate rich link previews when users share your links across apps, blogs, and news sites, enhancing content visibility and engagement.
ImagesiftBot
ImageSiftBot is a web crawler that scrapes the internet for publicly available images to support Hive's suite of web intelligence products.
IndeedJobBot
Indeed's job crawling bot that crawls job and job related information.
Inngest
Inngest is a platform for building event-driven applications.
Instapaper
Instapaper is an app that lets people save articles to read later.
Internet Archive - Archive-It
Internet Archive’s Archive-It service preserves publicly accessible web pages for the historical record.
Internet Archive Bot
The Internet Archive bot, also known as archive.org_bot, is the web crawler for the Internet Archive's Wayback Machine. It systematically crawls and preserves publicly accessible web pages for historical record.
InternetMeasurementBot
InternetMeasurementBot is operated by driftnet.io to discover and measure services that network owners and operators have publicly exposed.
JobicyBot
JobicyBot monitors and verifies job listings.
Jobs with GPT
Crawls job-related pages to power jobswithgpt.com, a platform for discovering AI-enhanced career opportunities.
jobswithgptcom-bot
Simple crawler focussing on only job postings for job search site.
Kagi Bot
Kagi Bot is the web crawler for the Kagi search engine. It crawls the web to build its own search index, which supports its ad-free search product.
KakaoTalk Scrap
The KakaoTalk scrap server collects and processes webpage data to create optimized previews for URLs.
kb.dk_bot
Royal Danish Library collects the Danish Internet according to the Danish Legal Deposit Act for research purposes.
KeldanNewsCrawler
Icelandic news article crawler for a news search on keldan.is.
Kernel Browsers
Runs browser automation on behalf of Kernel customers for web agents, automations, and web scraping.
Kernel Search
Kernel's search crawler that indexes web pages for AI-powered search and retrieval.
KimiBot
KimiBot crawls content potentially used to train Kimi's foundation models.
KlaviyoAIBot
Klaviyo’s web crawler for its Kai Customer Agent feature.
Level9SearchBot
.
Library Of Congress Web Archiving
The Library of Congress Web Archive manages, preserves, and provides access to archived web content selected by subject experts from across the Library, so that it will be available for researchers today and in the future.
LINE OGP Scraper
A scraper to get OGP by LINE Corporation.
LinerBot
LinerBot gathers web content for Liner, an AI research and answer assistant that cites the sources behind its responses.
Linespider
Linespider is a Web crawler that provides a wide range of search results for LINE services while complying with the Robots Exclusion Protocol. https://help2.line.me/linesearchbot/web/?contentId=50006055&lang=en.
LinkCheck
The Siteimprove LinkCheck crawler analyzes and monitors websites for quality assurance, SEO, and accessibility purposes, and keeps website content in line with brand guidelines and organizational policies.
LinkCheckerBot
LinkCheckerBot is a backlink monitoring crawl.
LinkedInBot
LinkedInBot is a bot that renders links shared on LinkedIn.
LinksIndexerBot
LinksIndexerBot is an SEO bot that crawls websites to index backlinks and aggregate website summaries.
LogicMonitor SiteMonitor
LogicMonitor SiteMonitor monitors your website's uptime, performance, and availability from multiple global regions.
LogRocketBot
LogRocket Asset Cacher is a bot that captures and caches web assets (CSS, JavaScript, images) to ensure proper playback of user sessions in LogRocket's session replay feature.
Loomly Bot
LoomlyBot is used to extract metadata from web pages in order to show a social media post preview within Loomly so that clients can see what their social media posts will look like when published.
Lumar
The Lumar website intelligence platform is used by SEO, engineering, marketing and digital operations teams to monitor the performance of their site’s technical health, and ensure a high-performing,...
MagiBot
MagiBot is owned by Peak Labs which focuses on the research and development of information extraction and retrieval technology to transform knowledge in natural language into immeasurable value.
MagnetmeBot
MagnetmeBot checks the websites of our paying customers and ensures the job openings are being kept in sync.
MailRUBot
The mail.ru bot is a mail fetcher on behalf of the Mail.ru email service.
Make.com
Make.com workflow automation platform connector that fetches data from customer endpoints.
Manus Bot
Manus is the action engine that goes beyond answers to execute tasks, automate workflows, and extend your human reach.
Marfeel Audits Crawler
Marfeel's audit crawlers that periodically re-crawl traffic-receiving URLs to detect structured data, meta tags, and HTML issues.
Marfeel Flowcards Crawler
Marfeel's crawler that fetches content for Flowcards that load directly from specific URLs.
Marfeel Preview Crawler
Marfeel's previewer crawler used to render preview experiences for both mobile and desktop views.
Marfeel Social Crawler
Marfeel's crawler used for social experiences (Facebook, X/Twitter, Telegram, Reddit, LinkedIn).
Marginalia Search
Marginalia Search is a noncommercial niche search engine focusing on old websites, personal websites, and blogs that suffer crippling discoverability problems in today's fiercely SEO-optimized lanscape.
marketgoo
marketgoo provides white label SEO tools.
Mars Finder
Mars Finder is a website search service designed to utilize the maximum potential of a website. MARS FINDER has held the top share of website search service market of Japan in 2017.
MediaMonitoringBot
MediaMonitoringBot crawls and indexes news and media publishers websites for a new materials and try to match it against keywords provided by our customers (subscribers) and send them updates based on that information.
Mediatoolkitbot
The Mediatoolkitbot is a media monitoring tool that crawls the open internet looking for phrases Determ users search for, helping marketers find relevant opportunities for advertising.
MelonMesa Bot
This bot is used to aggregate data about a popular online multiplayer game from consenting hosts who have opted-in to this collection.
meta-externalads
Crawls the web to improve advertising and business-related products and services.
meta-externalagent
The Meta-ExternalAgent crawler crawls the web for use cases such as training AI models or improving products by indexing content directly.
meta-externalfetcher
The Meta-ExternalFetcher crawler performs user-initiated fetches of individual links to support specific product functions.
meta-webindexer
Crawls web content to provide search results for Meta AI users.
MicrosoftPreview
MicrosoftPreview generates page snapshots for Microsoft products. It has desktop and mobile variants, with Chrome version dynamically updated to match the latest Microsoft Edge version.
MirrorWebCrawler
We are a commercial web archiving supplier providing archival solutions for the financial and public sector.
MistralAI-Index
MistralAI-Index crawls and indexes web content for Mistral's search feature in Le Chat. Content it indexes is not used to train Mistral's generative models.
MistralAI-User
MistralAI-User fetches web pages in real time when someone asks Le Chat a question, so Mistral's assistant can answer with current information and link to sources. It is not used for AI training.
MJ12bot
MJ12bot is a web crawler operated by Majestic-12 Ltd, a UK-based company that builds a search engine focused on backlink analysis and web structure mapping.
Mojeek
Details and information for webmasters regarding Mojeekbot, the web crawler for the Mojeek search engine.
MomenticBot
Momentic is a AI-powered platform for software testing. It allows you to write reliable end-to-end tests for web apps in a simple and intuitive way using natural language.
MotoMinerBot
MotoMinerBot is MotoMiner's web crawling bot. All vehicle detail pages we index are searchable via MotoMiner's search engine.
Moz rogerbot
Rogerbot is Moz's site audit crawler for Moz Pro Campaigns.
MRGbot
Search engine aimed at generating a corpus of data to be able to aggregate data in various ways.
MSN
MSNBot was the web crawler for Microsoft's MSN Search, which has since been replaced by Bing. Its purpose was to index web pages for inclusion in the MSN search engine.
naver-blueno
Naver's preview-snippet crawler that fetches summary information (titles, descriptions, images) when users insert links in Naver services such as blogs or cafés.
naverbot
Naver's web crawler (also known as Yeti) is used by Naver, South Korea's largest search engine, to crawl and index web content.
Navu
Navu crawls the websites requested by their customers and prospects to train AIs for them.
Neevabot
Neevabot is the web crawler for the search engine neeva.com.
netEstate Imprint Crawler
The NetEstate Imprint crawler crawls websites for public contact information.
New York Times Newsgathering
Coders within NYT's newsroom collect public, non-copyright data, e.g. our U.S. Elections pages and Covid-19 trackers.
NewRelic Minions
New Relic Synthetic monitoring infrastructure that performs API checks and virtual browser instances to monitor websites and applications from global locations
NewsBank
NewsBank aggregates licensed publisher content for schools, libraries, and government research, learning, and archiving.
NewsNow
The NewsNow bot is the web crawler for the news aggregator service NewsNow.
NitroBot
We run a cloud based site speed optimization solution. As such, we need to make requests to our clients' sites in order to fetch the content that needs to be optimized.
Nostra
Nostra accelerates site speed for managed web platforms.
Notabot
Crawler to integrate Helpfeel external search engine.
Novellum AI Crawl
Novellum.ai is building out tools for building agents. This MCP tool will be used by agents to crawl sites.
OAI-AdsBot
Validates the safety of web pages submitted as ads on ChatGPT; data collected is not used to train generative AI foundation models.
OAI-SearchBot
Indexes websites for inclusion in ChatGPT's search results; does not crawl content for AI model training.
Observer
Crawler looking for broken links to help improve world wide web as a whole.
OhDearBot
OhDearBot is a monitoring bot operated by Oh Dear that performs uptime checks, broken link detection, and mixed content scanning.
Omgilibot
Omgilibot crawls public web content for Webz.io, which packages and licenses web data feeds that are commonly used to train AI models.
OMIM Litrack
OMIM Litrack is a literature tracking crawler.
OnCrawl
Enterprise SEO platform powered by the industry-leading SEO Crawler and Log Analyzer.
Online Webceo Bot
A bot associated with WebCEO, a company that provides SEO tools and services.
OnticaBot
Feed-fetching crawler that discovers and fetches public articles for curated content feeds.
OpenGraph.io Bot
Our API is used by mostly consumer facing products to preview links when sharing them on their platforms.
OpenGraphXYZBot
Bot for opengraph.xyz service that generates and previews Open Graph meta tags and dynamic social media images
Orlo Link Preview
The Orlo Link Preview bot is used by the Orlo social media management platform. It fetches previews of links that are scheduled to be published in social media posts.
Oseox
Oseox is a French SEO platform that provides.
Ozon Web Grabber
A component that serves to load previews for external and internal links.
Pagefreezer Website Archiving
Pagefreezer Website Archiving is a compliance.
PanguBot
PanguBot crawls web content used to train Huawei's Pangu family of large language models.
ParticleNewsBot
Particle is an AI powered aggregator that collects news from many sources.
Payhawk Invoice Fetching Agent
Automated browser bot that fetches invoices for users from supplier websites and attaches them to their expense records.
PayPal
PayPal delivers real-time event notifications for payments, subscriptions, and account updates.
payroll-bot
payroll-bot is an AI crawler operated by ADP, Inc. to collect publicly available legal and payroll documentation.
Perplexity-User
Handles user-initiated requests in Perplexity, accessing external content to provide real-time information; not used for automated crawling or AI training.
PerplexityBot
Indexes websites for inclusion in Perplexity's search results; does not crawl content for AI model training.
PetalBot
PetalBot is a web crawler operated by Huawei's Petal Search engine.
PhindBot
PhindBot crawls technical and developer-focused web content to power Phind, an AI answer engine aimed at programmers.
Pingdom Bot
Pingdom Bot is used by Pingdom's monitoring services to perform various checks on websites, including uptime and performance monitoring.
Pinterest Bot
Pinterest's web crawler that indexes content for their platform. It crawls websites to collect metadata for Pins, including images, titles, descriptions, and prices.
Platebreaker
Recipe nutrition search engine. Indexes schema.org/Recipe JSON-LD. Respects robots.txt. Verified via Web Bot Auth.
Polar Webhooks
Polar's webhook service delivers real-time event notifications for payment processing, including purchases, subscriptions, cancellations, and refunds.
Potions
The Potions bot fetches product feeds and crawls data from its customers' websites, used for e-commerce related services.
prerender
It's HTML pre rendering service for SPA(Single Page Application) Website SEO.
PressEngine Bot
The PressEngine Bot verifies coverage created by video games press as genuine and their own creation.
Pricey
Pricey collects and compares product prices, showing trends and helping users find the best time to buy deals online.
Promptwatch Bot
Promptwatch Bot is Promptwatch's verification crawler that validates crawler-log integrations and runs on-demand SEO and performance audits.
ProximicBot
Proximic is Comscore's web crawler that performs contextual content analysis to help advertisers determine the best matching campaigns for a page's content.
PulsePoint Crawler
A web crawler used by PulsePoint, a digital advertising technology company, for content indexing and ads.txt verification.
QA.tech
The QA.tech web agent browses the website and identifies potential test cases, and executes tests against a web application
QStash
QStash is a platform for building event-driven applications.
QualifiedBot
Bot crawls customer websites to provide information to customer hosted chatbots.
Quantcastbot
Quantcast Bot is a web crawler used for advertisement quality assurance and to understand page content for Interest-Based Audiences.
Quartr Crawler
Quartr uses a crawler to obtain and deliver investor relations material.
Qwantbot
Crawls and indexes web content for Qwant search engine.
Rakuten Image extraction bot
Rakuten uses this bot to crawl product images so that we can display cashbach deals for our merchants.
Razorpay-Webhook
Razorpay’s webhooks enable merchants to receive secure, real-time HTTP callbacks for key payment events, automating reconciliation, notifications, and downstream workflows.
RDTvlokipBot
Official crawler for RDTvlokip Search, an independent French search engine. Respects robots.txt.
Redirect pizza destination monitor
redirect.pizza's destination monitor ensures that the redirect destination URLs are reachable.
Retool
Retool platform user agent.
Revvim
Our bot crawls our customers' websites to identify SEO opportunities.
RyeBot
Powers automated checkout on behalf of shoppers with explicit consent.
Sanity Webhooks
Sanity's webhook service that delivers real-time event notifications for content changes and other events.
Sansec Security Monitor
Sansec Security Monitor is a web crawler that monitors online stores for malicious code, data breaches, and digital skimming attacks.
SBIntuitionsBot
SBIntuitionsBot is a crawler operated by SB Intuitions Corp. that collects web data for AI development and information analysis.
ScreamingFrogBot
Screaming Frog SEO Spider is a website crawler used by SEO professionals for site audits and technical SEO analysis.
Screpy
SEO Checker for Screpy Bot that SEO.
SE Ranking Backlinks
SE Ranking's backlink analysis crawler that discovers and analyzes backlink profiles for SEO research and competitive analysis.
SearchAtlas Bot
Bot used to evaluate customer's websites and provide SEO optimization strategy.
SeekportBot
SeekportBot is the web crawler for Seekport, a German search engine operated by SISTRIX. The bot crawls and indexes web content while respecting robots.txt directives and crawl delays.
Selectika AI
Selectika AI enrichment compute vision for Fashion.
SemanticScholarBot
The Semantic Scholar bot crawls domains to find academic PDFs. These PDFs are served on semanticscholar.org so researchers can discover and understand other academic accomplishments.
Sentry Uptime Monitoring Bot
Sentry's Uptime Monitoring Bot performs health checks on configured URLs to monitor the availability and reliability of web services.
SEO Audit Check Bot
SEO audit check bot is likely an automated tool within the WebCEO platform that performs comprehensive SEO audits on websites.
seo4ajax
The seo4ajax bot is used by a service that helps make single-page applications (SPAs) crawlable by search engines. It pre-renders JavaScript-heavy pages into static HTML so they can be indexed.
Seobility
Seobility is a browser-based online SEO software that helps you improve your website’s search engine rankings.
SerpstatBot
SerpstatBot is the Serpstat bot collects data for Serpstat's Backlink Analysis tool.
ServerHunterSpider
Our spider indexes the price, specifications and stock of hosting plans. We fully respect robots.txt and we have more information on https://www.serverhunter.com/spider/.
SeznamBot
SeznamBot is the web crawler operated by Seznam.cz, the leading Czech search engine.
ShapBot
Crawls and indexes web content to power Parallel's search and content extraction APIs for AI applications.
Shopify Webhooks
Shopify webhooks are useful for keeping your app in sync with Shopify data, or as a trigger to perform an additional action after that event has occurred.
Shortwave Image Fetcher
An email client that proxies all images found in HTML emails from to protect end customer's IP address and connection private.
SISTRIX Optimizer Uptime
SISTRIX Optimizer Uptime bot performs continuous monitoring of website availability by checking the startpage once per minute. It is part of SISTRIX's SEO and website monitoring platform.
Site24x7
Site24x7 Bot is used by Site24x7's monitoring services to perform various checks on websites, including uptime and performance monitoring.
Sitebulb
Sitebulb is a desktop and cloud-based website crawler used by SEO professionals for technical SEO audits.
SiteGuru
SiteGuru is an SEO auditing tool that crawls.
Siteimprove Crawl
Siteimprove content suite (i.e. Quality Assurance, Accessibility, Policy, and SEO). Crawls run on ports are 80 for HTTP and 443 for HTTPS.
SiteSearch360
Site Search 360 is a popular Google Site Search replacement. Our crawler indexes content on our customers' sites for search.
Skroutz ImageBot
Skroutz ImageBot to fetch the individual product images.
Skype
Skype's URI Preview services fetches a page preview when someone posts a URL in a Skype message.
Slack-ImgProxy
Slack-ImgProxy is a bot operated by Slack that fetches and caches images posted in Slack channels.
Slackbot
Slackbot is Slack's default, general-purpose bot that handles various API requests and integrations.
SlackLinkExpandingBot
Slackbot Link Expanding is a bot operated by Slack that fetches metadata from shared links to create rich previews.
SMTnet PM Bot
Crawls partner company's websites to include them in our on-site search engine.
SnapchatAdsBot
SnapchatAdsBot is a crawler operated by Snapchat that verifies and analyzes websites for their advertising platform.
SnapURLPreviewBot
SnapURLPreviewBot is a crawler operated by Snap Inc. that analyzes and generates previews of URLs shared on Snapchat and other Snap platforms.
Sogou Web Spider
The web crawler for sogou.com
sourcedash
UK Property Search Engine.
Spyglasses
Spyglasses accesses site content to assist with AEO capabilities. Learn more at https://www.spyglasses.io.
Stably
Stably is a QA testing bot that users run to E2E test their websites for functionality testing and protecting user flows against regressions.
Statabot
Statabot searches for stata.toc files and indexes their contents.
StatistikAustria
Bot to collect product prices for the official consumer price index of Austria.
StatsDroneBot
The StatsDrone affiliate marketing statistics scraping and aggregating tool.
StatusCake Page Speed
StatusCake Page Speed monitors your page load and render speeds.
StatusCake SSL Monitoring
StatusCake SSL monitors your website certificates for common issues
StatusCake Uptime
StatusCake monitors the uptime of your website.
Steam Chat
The Steam Chat bot fetches previews of URLs shared within the Steam client's chat feature.
Stripe Webhooks
Stripe's webhook service that delivers real-time event notifications for payment processing and account updates.
Stripebot
Crawls Stripe merchant websites to collect data for service delivery and financial regulatory compliance.
Strivve Automation
Strivve Web Bot Authentication Agent.
svix
svix is a webhook service for sending events to webhooks.
TangibleeBot
TangibleeBot is a crawler operated by Tangiblee that collects product data from e-commerce websites to power their product visualization and virtual try-on services.
Telegram Bot
TelegramBot crawls websites to render a link preview when people send a message containing a URL in the Telegram messaging service.
TermlyBot
Crawls websites to detect and categorize cookies set by first and third parties.
Terracotta
The Terracotta bot scrapes websites for use in generating indices for serving searches using Ceramic's search product.
TikTokSpider
TikTokSpider is a web crawler used by TikTok/ByteDance to index and analyze web content for their platform. It helps in content discovery, link previews, and data collection for TikTok's services.
Timpibot
Timpibot crawls the web to build Timpi's decentralized search and data index, which is used to supply training and grounding data for AI applications.
TNOThesisCrawler
Research crawler by TNO collecting publicly available academic theses from university websites for the DIAMONDS platform.
Toutiao
Toutiao is ByteDance's automated news aggregation bot collecting content across web platforms.
Trellis-Services
Critical CSS Generator to Optimize Websites.
Trendiction Bot
Trendiction's web crawler that discovers and collects public web data for their social media monitoring and media intelligence platform.
TrustedSite
Crawl a customer website to perform basic validation.
TTD-Content
TTD-Content is a crawler operated by The Trade Desk that verifies content and quality of ad placements for their demand-side platform.
Tumblr
On Tumblr, post authors can paste a URL in their post, and we'll "unfurl" that URL into a pretty Link "Block" for their post by making a request to the URL and parsing the response.
TurnitinBot
Turnitin.com offers various services to the educational community. Most prominently, we provide a widely used and effective plagiarism detection service.
Twilio Knowledge
Twilio's AI assistant crawler that gathers web content to build knowledge bases for Twilio AI Assistants, enabling conversational AI experiences with up-to-date information.
Twilio Proxy
Twilio's proxy service that handles communications between end-users and applications through Twilio's programmable voice and messaging platform.
TwinAgent
Automate complex operations end-to-end.
Twitterbot
Fetches content for shared links on X/Twitter to generate rich previews.
upday
upday is a news aggregator app, and its bot crawls news sources. It collects and indexes articles to be recommended to users on its platform.
Updown.io
Performs uptime and performance checks on websites.
Uptime Robot
Uptime Robot is a platform for monitoring and alerting on your applications.
UsercentricsBot
UsercentricsBot is operated by Usercentrics GmbH to scan websites for data processing services and third-party technologies.
v0bot
Bot for v0 services.
Velen Public Web Crawler
Velen Public Web Crawler collects public web content for Webz.io's data feeds, which are licensed for AI training, market intelligence, and monitoring.
Vemetric Favicon Bot
Fetches favicons from websites in the highest quality available.
Vercel build container
System-initiated requests made from Vercel's build container during a build
Vercel Favicon Bot
Vercel Favicon Bot
Vercel Screenshot Bot
Vercel Screenshot Bot
vercelflags
vercel flags
verceltracing
vercel tracing
videootv Bot
Crawler to extract the newest articles in the publisher's website (via feed or parsing html) to make a carrousel with images, links and text for our native ads module in order to improve recirculation in the publisher's web.
Visually.io Shopify Editor
Shopify theme editor alternative for live, real-time store editing via a secure iframe and controlled proxy.
W3 Validator Services
W3C provides various free validation services that help check the conformance of Web sites against open standards.
WARDBot
WARDBot tracks URL status codes, helping users monitor the availability of web pages they have added to the monitoring list.
WebSpiderMount
Job wrapping data processor handling jobs distribution from employer websites to multiple endpoints, like job boards, advertisement platforms, job alerts etc.
WindowsForum-AI
Technology news aggregation bot for WindowsForum.com.
WMF Citoid
Citoid is a Wikimedia service in VisualEditor that generates citations from URLs, DOIs, and ISBNs, relying on the Zotero Translation Server (see wikimedia-zotero) for accurate metadata, processed on demand from website visitors.
WMF Zotero Translation Server
The Wikimedia Foundation's Zotero Translation Server is a customized metadata extraction tool that powers Citoid (see wikimedia-citoid), retrieving citation data from URLs, DOIs, and ISBNs using Zotero translators, on demand from website visitor requests.
WordCountBot
WordCountBot analyzes website word count based on public pages. All words belonging to public pages and included in HTML source code.
XY Archive Compliance Bot
Website archiver for our customers who have archive compliance requirements to fulfill them.
Yahoo Ad Monitoring
Yahoo Ad Monitoring crawls landing pages of URLs listed with Yahoo advertising services to analyze content quality, ensure ad relevance, and improve user experience by maintaining accurate ad...
Yahoo Japan SEO Crawler
Yahoo Japan search engine crawler for SEO analysis.
Yahoo Link Preview
Yahoo Link Preview's bot fetches data from URLs shared on Yahoo platforms.
Yahoo! JAPAN
Yahoo! JAPAN manages and operates a system that accesses web pages published on the Internet for the purpose of providing services, research, development, maintenance, etc.
Yahoo! Slurp
Yahoo! Slurp is the web crawler (robot) used by Yahoo! Search to discover and index web pages for its search engine.
YahooCacheSystem
YahooCacheSystem caches website contents as part of the Yahoo! Search Service.
YahooMailProxy
Yahoo Mail Proxy is a content fetch proxy that retrieves the page content of URLs that are embedded within emails sent to Yahoo Mail users.
YandexAdditional
YandexAdditional is the crawler Yandex uses to collect web content for its YandexGPT and other generative AI products, separate from the YandexBot search crawler.
Yandexbot
YandexBot is a web crawler operated by Yandex, a major Russian search engine.
Yeti
Yeti is the web crawler for Naver, a South Korean search engine. It indexes websites to provide search results and power other services on the Naver platform.
YGS Group Falconer Scraper
A content based scraper only for partners we collaborate with who have given permission to have their website scraped.
YisouSpider
YisouSpider is a search engine crawler operated by Yisou that indexes web content for their search engine results. The crawler follows standard crawling practices and respects robots.txt directives.
YouBot
YouBot crawls and indexes web pages to power the You.com AI search engine and its cited answers.
Zoombot
ZoomBot is SEOZoom's web crawler that builds.
ZoomInfo
Zoominfobot is an indexing robot for a web search engine, similar to Google. Created by Zoom Information Inc.(www.zoominfo.
ZumBot
ZumBot is a web crawler that indexes webpages for Zum Open Internet Search.
Be the brand AI recommends
Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.
