Promptwatch Logo

Diffbot

Diffbot crawls and structures web pages into a knowledge graph that is sold for AI training, retrieval, and data enrichment.
DiffbotDiffbot
AI CrawlerAI Training

What is Diffbot?

Diffbot turns public web pages into structured records. Its crawler reads pages, identifies entities such as articles and organizations, and feeds a self-updating Knowledge Graph that customers can search or use to enrich their own data.

The Diffbot token denotes Diffbot's proactive crawl for its general search engine and Knowledge Graph. Diffbot documents a separate Diffbot-User agent for a URL fetched in response to a human using its software. Customer-run Extract and Crawl API jobs can also send a custom user agent, so not every request made through Diffbot infrastructure will carry the token covered on this page.

Diffbot says its proactive crawler uses caching, compression, conditional requests, and predictive scheduling to reduce repeat work on origin servers. Those measures describe normal behavior, not an identity check. this bot has no Cloudflare entry, verified IP range, or HTTP-signature directory for the crawler.

There is an important training distinction. Diffbot sells structured Knowledge Graph data for uses that include machine learning, retrieval, and data enrichment, but its crawler documentation says the Diffbot crawl is not used to train generative AI foundation models. Allowing it can make a page discoverable in Diffbot's graph and web-search services; it does not grant a foundation-model trainer direct access under this token.

By default, Diffbot follows Disallow and Crawl-delay in robots.txt. It does not follow the Allow directive. Diffbot also says robots rules may be overridden for a specific partnership or agreement, while customer Crawl API jobs can disable the default robots setting. The true metadata value therefore describes the standard proactive crawl, not every possible Diffbot-backed fetch.

A site should choose access based on whether it wants its public facts represented in Diffbot's products and downstream customer uses. If a page contains data that should not enter the Knowledge Graph, exclude it explicitly and watch for other Diffbot-related agents rather than assuming one rule covers customer-configured jobs.

Relevant for AI search

Is Diffbot relevant for AI search?

Yes. Diffbot collects pages for an AI product, so what it can crawl influences how AI systems describe your brand.

Diffbot gathers public web content that can end up in the training data for large language models. Once your pages are in that set, they influence how the operator's models talk about you for that model generation. Allowing it lets your own writing carry weight in those answers; blocking it means the models learn about you from third parties instead.

How to handle Diffbot

Leave Diffbot allowed if you want public pages eligible for Diffbot's Knowledge Graph and web-search services. To exclude the proactive crawler from the whole site, use:

User-agent: Diffbot
Disallow: /

Diffbot follows Disallow and Crawl-delay by default, but not Allow. Write path rules with that limitation in mind. Its documentation also permits an override where Diffbot has a site agreement, so resolve unexpected access with the party running that crawl.

This group does not necessarily govern Diffbot-User, a custom user agent, or every customer-configured Extract or Crawl API request. Enforce sensitive-path policy with authentication and edge controls, and do not trust the plain Diffbot string as proof of origin.

Examples

  • A business directory allows Diffbot to crawl public company profiles so their facts can be structured in the Knowledge Graph, but it excludes paid exports.
  • A publisher uses `Disallow` for an archive instead of relying on an `Allow` exception, which Diffbot says it does not honor.
  • A log review separates routine `Diffbot` discovery from `Diffbot-User` requests triggered by people using Diffbot software.

Frequently asked questions about Diffbot

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

Diffbot converts public web content into structured entities and facts for its Knowledge Graph and web-search services. Customers can query that graph or use it to enrich their own records.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard