What is AI2Bot?
AI2Bot is the web crawler used by the Allen Institute for Artificial Intelligence, commonly called Ai2. The institute's crawler notice says it explores selected domains for web content that is used to train open language models.
This is a collection crawler, not a user-triggered browser. A request does not mean that someone asked an Ai2 model about the page at that moment. It means the URL was considered during a crawl for potential training material, subject to whatever selection and processing Ai2 applies later.
The training purpose gives the allow or block decision a fairly direct consequence. Allowing AI2Bot can make public material available for Ai2's open-model work. It does not guarantee that a particular page will remain in a dataset, be memorized by a model, appear in an answer, or receive a citation. Blocking the crawler affects this collection route, not copies obtained elsewhere.
Ai2 publishes one full request identity: Mozilla/5.0 (compatible) AI2Bot (+https://www.allenai.org/crawler). The stable robots token within it is AI2Bot. The operator says this string can be used to filter or reject crawler traffic.
This entry records AI2Bot as respecting robots.txt. Ai2's short crawler notice does not spell out crawl-delay support, request rates, JavaScript rendering, or other protocol details. Use a standard allow or disallow rule, then check your own access logs if path-level compliance matters.
The bot remains marked unverifiable because no Cloudflare record, published IP range, reverse-DNS suffix, or HTTP signature directory is attached to this entry. A matching user agent is useful for classification, but it is not proof that a request came from Ai2.
