Promptwatch Logo

Brightbot

Brightbot is Bright Data's crawler layer that monitors the health of websites and enforces ethical web data collection.
Brightbot
MonitoringAnalytics

What is Brightbot?

Brightbot is Bright Data's named pipeline for collecting public web data across its products and services. That is broader than a search crawler with one index and one destination. Bright Data offers data tools for uses that include AI training, but a Brightbot request does not reveal the customer, dataset, or eventual use of the page.

The crawler has a cache intended to prevent another download of the same data within 24 hours, unless Bright Data approves an exception for a business reason. Bright Data also opens a health monitor for targeted domains. If its traffic correlates with slower responses, the system applies a rate limit based on the last rate that did not hurt site performance.

Bright Data publishes two identifiers that should be checked together: the user agent Brightbot 1.0 and source addresses in 82.97.199.0/24. This makes recognized Brightbot traffic easier to separate from ordinary visitors. It does not identify every request made through Bright Data's products.

Verified site owners can use the Webmaster Console to submit collectors.txt rules for personal information, private or copyrighted material, and interactive endpoints. Bright Data reviews the file before Brightbot enforces it. The company says this identity and the file currently cover Web Unlocker traffic, not Browser API or browser-based Data Collector jobs.

Not relevant for AI searchBrightbot

Is Brightbot relevant for AI search?

No. Brightbot is not part of AI search or training, so allowing or blocking it does not change your AI visibility.

Brightbot runs synthetic monitoring or uptime checks on a schedule. It is not collecting content for any index or model, so it has no bearing on search rankings or AI answers. You usually see it because your own team or a service you use set up a check.

How to handle Brightbot

Use the Webmaster Console and an approved collectors.txt file when you want identifiable Brightbot traffic on public pages but need specific endpoints excluded. Match both the published user agent and subnet before granting an allowlist exception.

The stable name can also express a site-wide robots.txt preference:

User-agent: Brightbot
Disallow: /

This catalog records robots support, but Bright Data's current control documentation does not explain how ordinary robots directives are applied. Its documented mechanism is collectors.txt, and even that only covers Web Unlocker traffic. Keep non-public data authenticated and use edge controls for a mandatory denial. Bright Data warns that blocking URLs marked as allowed can make the Brightbot identity disappear from traffic to the domain for seven days.

Examples

  • A marketplace verifies its domain, keeps catalog pages available, and marks account and review-posting routes in `collectors.txt` before approving the file.
  • An operations team sees `Brightbot 1.0` outside the published subnet and declines to treat the header alone as genuine Brightbot traffic.

Frequently asked questions about Brightbot

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

No. Bright Data calls it the main collection pipeline for its products and services, although the named identity does not cover every Bright Data network.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard