Skip to content
Directory/Categories/Academic Research/Library Of Congress Web Archiving
LI

Library Of Congress Web Archiving

GoodBotVerifiedBOT

The Library of Congress Web Archive manages, preserves, and provides access to archived web content selected by subject experts from across the Library, so that it will be available for researchers today and in the future. More information on the programme here: https://www.loc.gov/programs/web-archiving/about-this-program/ And information about crawling policy here: https://www.loc.gov/programs/web-archiving/for-site-owners/ Operated by United States Library of Congress.

GoodBot
Assessment
Academic Research
Category
BOT
Kind
1
Data sources
AI-generated summary, drawn from Cloudflare, operator docs, and community data

What is Library Of Congress Web Archiving?

Library Of Congress Web Archiving is a web crawler operated by United States Library of Congress. The Library of Congress Web Archive manages, preserves, and provides access to archived web content selected by subject experts from across the Library, so that it will be available for researchers today and in the future.

More information on the programme here: https://www.loc.gov/programs/web-archiving/about-this-program/

And information about crawling policy here: https://www.loc.gov/programs/web-archiving/for-site-owners/

Helpful — Verified, safe crawler. Respects robots.txt and provides operator documentation.

Identification

Bot name
Library Of Congress Web Archiving
Operator
United States Library of Congress
Category
Academic Research
Kind
BOT
Verification
Verified bot
Cloudflare slug
library-of-congress-web-archiving
User-agent patterns
https://www.loc.gov/programs/web-archiving/for-site-owners/
User-agent string
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/78.0.3904.97 Safari/537.36 (+https://www.loc.gov/programs/web-archiving/for-site-owners/)
Main use cases
Web crawlingData collection

Why it crawls your site

Library Of Congress Web Archiving crawls websites to collect data for academic and scientific research. It is operated by United States Library of Congress as part of their academic research infrastructure. If you see this bot in your server logs, it is a verified crawler and generally safe.

Block or allow Library Of Congress Web Archiving

Block it

Add a Disallow rule for https://www.loc.gov/programs/web-archiving/for-site-owners/ in your robots.txt file. You can also block at the server level using your web server configuration or CDN firewall rules to filter requests matching the user-agent string.

# Block Library Of Congress Web Archiving from crawling your entire site
User-agent: https://www.loc.gov/programs/web-archiving/for-site-owners/
Disallow: /

# Allow Library Of Congress Web Archiving full access
User-agent: https://www.loc.gov/programs/web-archiving/for-site-owners/
Allow: /
Allow & verify

Ensure your robots.txt allows https://www.loc.gov/programs/web-archiving/for-site-owners/. Verify requests are genuine by checking the user-agent string and referring to United States Library of Congress's documentation.

Traffic & trends

Cloudflare Radar traffic

No Cloudflare traffic data available.

Google Trends interest

Loading trend data…

Data sources

This profile is compiled from the following sources.

CloudflareRadar
Last updated July 26, 2026