Promptwatch Logo

Omgilibot

Omgilibot crawls public web content for Webz.io, which packages and licenses web data feeds that are commonly used to train AI models.
Webz.ioomgilibot
AI CrawlerAI Training

What is Omgilibot?

Omgilibot is a Webz.io crawler for collecting public web content. Webz.io originally built it for the Omgili search engine, which the company later discontinued, and then used the crawler to supply its web data service. A request from this bot belongs to that collection pipeline, not to a person browsing the page.

Webz.io says the crawler has covered sources such as news sites, blogs, review pages, ecommerce sites, and online discussions. After retrieval, Webz.io indexes the material and makes structured records available through its APIs and licensed data feeds. Customers use those feeds for research, media monitoring, and software that needs web data without running its own crawler.

Some Webz.io datasets are sold for machine learning and large language model training. Omgilibot can therefore be an upstream source of training material, although a visit does not show which customer will receive the page or whether any particular model will use it. Omgilibot does not operate a public AI answer engine, so allowing it offers no documented promise of a citation or placement in AI search.

The stable token recorded for this crawler is omgilibot. That token is useful when reviewing access logs or writing a robots.txt group, but a user-agent value by itself does not authenticate the sender. This bot has no published request-signing method or verified IP mechanism in the record used for this page.

Webz.io states that Omgilibot honors robots.txt and accepts direct requests from site owners who do not want it to crawl. A specific rule lets you exclude the whole site or selected paths while leaving ordinary search crawlers alone. Review AI crawler access separately from general robots.txt policy because the downstream use here is licensed data distribution.

The practical choice depends on the material. A public publisher may accept collection of open articles but exclude subscriber archives, while a company with proprietary research may block the crawler entirely. Logs can show what Omgilibot requested and whether it returned after a rule changed, but they cannot reveal the eventual customer or dataset.

Relevant for AI search

Is Omgilibot relevant for AI search?

Yes. Omgilibot collects pages for an AI product, so what it can crawl influences how AI systems describe your brand.

Omgilibot gathers public web content that can end up in the training data for large language models. Once your pages are in that set, they influence how the operator's models talk about you for that model generation. Allowing it lets your own writing carry weight in those answers; blocking it means the models learn about you from third parties instead.

How to handle Omgilibot

Decide whether Webz.io's licensed data distribution fits your publishing policy. If it does, leave public pages available and use path rules for areas such as paid archives. If it does not, address the crawler by its recorded token.

To request a site-wide block, add this group to robots.txt:

User-agent: omgilibot
Disallow: /

Webz.io documents robots.txt compliance for Omgilibot. After publishing the rule, check that /robots.txt returns successfully on every relevant host and review later requests for the token. Use access controls rather than robots.txt for private content, since robots.txt does not create authentication.

Examples

  • A news publisher allows Omgilibot to fetch free reporting but disallows the subscriber archive that contains licensed investigations.
  • A forum operator finds `omgilibot` in access logs and blocks member profile paths while leaving public discussion threads available to Webz.io's data feeds.
  • A research company does not license its reports for downstream model training, so it adds the site-wide rule and confirms that later crawl attempts stop.

Frequently asked questions about Omgilibot

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

Webz.io uses Omgilibot to collect and index public web material for APIs and licensed data feeds. The crawler was first built for the discontinued Omgili search engine.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard