What is TNOThesisCrawler?
TNOThesisCrawler gathers publicly available master's theses from university websites for the DIAMONDS platform. TNO says the collection provides structured access to academic work for research. Its policy also names AI model training as a purpose, so this crawler has a direct training-data connection that most academic bots do not.
The stated scope is public thesis documents. TNO says the bot does not enter login pages or restricted material and does not collect private data. Repositories should still enforce those boundaries with authentication because a robots file cannot make an exposed document private.
Collection runs can be short and busy. TNO warns that one host may receive more than 10,000 requests per day for roughly a week, followed by no traffic until another scheduled run. The published ceiling is about five requests per second for each host.
Traffic identifies itself as TNOThesisCrawler/1.0 and includes a policy URL and support email. A Web Bot Auth signature directory is also published for verification. TNO's policy says the crawler respects robots.txt, but the accompanying crawler record marks it as noncompliant, so the promise and observed classification do not fully agree.
