What is Coveobot?
Coveobot is the crawler Coveo uses when a customer configures web content for a Coveo index. Coveo's platform combines material from websites and other repositories so that the customer can offer search, recommendations, customer-service experiences, or generative answers. The crawler is therefore tied to an enterprise content source rather than a universal consumer search engine.
The resulting index can support a generative experience, so pages available to Coveobot may influence answers inside that customer's Coveo implementation. This is not the same as broad visibility across public AI assistants. Coveo's crawler documentation does not state that a page fetched by Coveobot is used to train a general-purpose model.
For a Coveo Web source, an administrator supplies a starting URL and the crawler discovers pages through site navigation and links. The source configuration controls what content is indexed and how it is retrieved. Coveo can also retain source permissions so that indexed items remain visible only to users who are authorized to see them.
The standard header recorded by Coveo and Cloudflare is Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko) (compatible; Coveobot/2.0;+http://www.coveo.com/bot.html). Coveo's default robots matching name is CoveoBot. There is no Web Bot Auth key directory or verified IP identity in this bot, so the user-agent text cannot authenticate a request and the bot remains classified as unverifiable.
A Coveobot visit usually means that a particular organization chose the site as one of its content sources. If the site belongs to that organization, the requests may be refreshing its own support, commerce, or documentation index. If another party configured the URL, the destination may not have a direct relationship with the Coveo customer.
Coveo's Web source respects robots.txt page restrictions and Crawl-delay by default, which is why this bot is classified as respecting robots.txt. Administrators can enable an override that tells a Web source to ignore those directives, and Cloudflare's directory labels the agent as not following robots.txt. A rule is meaningful for the default setup, but it cannot guarantee the behavior of every customer-configured source.
