What is NewsBank?
NewsBank works with publishers to build licensed collections of current and archived material. Schools, libraries, government organizations, and other researchers use those collections. Its crawler belongs to that publisher relationship rather than an open project that archives any public website it encounters.
One documented ingestion method is a customizable web harvest of staff-produced content. NewsBank says the setup can exclude syndicated or non-copyrighted material. Publishers can also deliver single-issue PDFs, content-management feeds, and legacy data, so a website crawl is only one way content reaches the archive.
Requests have been recorded as NewsBank.com/1.0 and NewsBank.comMobile/1.0. A participating publisher should compare the requested pages with its agreed harvesting method. NewsBank does not publish a crawler address range or a DNS verification process on the publisher pages, which means the header cannot establish authorization by itself.
The resulting archive is a licensed research product, not a documented generative AI dataset. Blocking an expected harvest can interrupt archive updates or staff access, but allowing it has no stated effect on AI search visibility or model training.
