Promptwatch Logo

TNOThesisCrawler

Research crawler by TNO collecting publicly available academic theses from university websites for the DIAMONDS platform.
TNOTNOThesisCrawler
Academic Research

What is TNOThesisCrawler?

TNOThesisCrawler gathers publicly available master's theses from university websites for the DIAMONDS platform. TNO says the collection provides structured access to academic work for research. Its policy also names AI model training as a purpose, so this crawler has a direct training-data connection that most academic bots do not.

The stated scope is public thesis documents. TNO says the bot does not enter login pages or restricted material and does not collect private data. Repositories should still enforce those boundaries with authentication because a robots file cannot make an exposed document private.

Collection runs can be short and busy. TNO warns that one host may receive more than 10,000 requests per day for roughly a week, followed by no traffic until another scheduled run. The published ceiling is about five requests per second for each host.

Traffic identifies itself as TNOThesisCrawler/1.0 and includes a policy URL and support email. A Web Bot Auth signature directory is also published for verification. TNO's policy says the crawler respects robots.txt, but the accompanying crawler record marks it as noncompliant, so the promise and observed classification do not fully agree.

Indirectly relevant

Is TNOThesisCrawler relevant for AI search?

Indirectly. TNOThesisCrawler has no AI product of its own, but its output can end up in the systems that AI answers draw on.

TNOThesisCrawler gathers pages for research corpora and web-measurement studies. Several widely used AI training datasets began as academic crawls, so pages collected for a study can later teach commercial models what your brand is.

How to handle TNOThesisCrawler

Allow the bot only if the repository's terms permit public theses to be indexed for DIAMONDS research and AI training. You can limit collection to approved thesis paths while keeping licensed, embargoed, or withdrawn files unavailable to anonymous requests.

TNO documents support for this rule:

User-agent: TNOThesisCrawler
Disallow: /

The operator promises robots.txt compliance, while the crawler record says otherwise. Watch the next collection run to confirm that the chosen paths are skipped. If your CDN supports Web Bot Auth, validate signatures with https://diamonds.tno.nl/.well-known/http-message-signatures-directory rather than allowlisting the user-agent string.

Plan capacity for temporary bursts when you allow the crawler. TNO publishes a rate of about five requests per second per host and asks operators to contact [email protected] about crawl problems. Blocking prevents DIAMONDS from collecting the affected material for both stated uses, including AI model training.

Examples

  • A university repository records more than 10,000 TNOThesisCrawler requests in one day. It verifies the request signatures, confirms the per-second rate stays within TNO's policy, and lets the temporary collection run finish.
  • A graduate school permits `/public-theses/` but excludes `/embargoed/` for this bot. The embargoed documents also require login, so they remain protected even if a crawler ignores the robots rule.

Frequently asked questions about TNOThesisCrawler

Learn about AI visibility monitoring and how Promptwatch helps your brand succeed in AI search.

It collects publicly available master's theses for TNO's DIAMONDS platform, which provides structured access to academic work.

Be the brand AI recommends

Monitor your brand's visibility across ChatGPT, Claude, Perplexity, and Gemini. Get actionable insights and create content that gets cited by AI search engines.

Promptwatch Dashboard