Library Of Congress Web Archiving
GoodBotVerifiedBOTThe Library of Congress Web Archive manages, preserves, and provides access to archived web content selected by subject experts from across the Library, so that it will be available for researchers today and in the future. More information on the programme here: https://www.loc.gov/programs/web-archiving/about-this-program/ And information about crawling policy here: https://www.loc.gov/programs/web-archiving/for-site-owners/ Operated by United States Library of Congress.
What is Library Of Congress Web Archiving?
Library Of Congress Web Archiving is a web crawler operated by United States Library of Congress. The Library of Congress Web Archive manages, preserves, and provides access to archived web content selected by subject experts from across the Library, so that it will be available for researchers today and in the future.
More information on the programme here: https://www.loc.gov/programs/web-archiving/about-this-program/
And information about crawling policy here: https://www.loc.gov/programs/web-archiving/for-site-owners/
Identification
https://www.loc.gov/programs/web-archiving/for-site-owners/Why it crawls your site
Library Of Congress Web Archiving crawls websites to collect data for academic and scientific research. It is operated by United States Library of Congress as part of their academic research infrastructure. If you see this bot in your server logs, it is a verified crawler and generally safe.
Block or allow Library Of Congress Web Archiving
Add a Disallow rule for https://www.loc.gov/programs/web-archiving/for-site-owners/ in your robots.txt file. You can also block at the server level using your web server configuration or CDN firewall rules to filter requests matching the user-agent string.
# Block Library Of Congress Web Archiving from crawling your entire site User-agent: https://www.loc.gov/programs/web-archiving/for-site-owners/ Disallow: / # Allow Library Of Congress Web Archiving full access User-agent: https://www.loc.gov/programs/web-archiving/for-site-owners/ Allow: /
Ensure your robots.txt allows https://www.loc.gov/programs/web-archiving/for-site-owners/. Verify requests are genuine by checking the user-agent string and referring to United States Library of Congress's documentation.
Traffic & trends
No Cloudflare traffic data available.
Loading trend data…
Links & references
Data sources
This profile is compiled from the following sources.