Skip to content
Directory/Categories/Archiver/Internet Archive - Archive-It
IN

Internet Archive - Archive-It

GoodBotVerifiedBOT

Internet Archive’s Archive-It service preserves publicly accessible web pages for the historical record. Operated by Archive-It.

GoodBot
Assessment
Archiver
Category
BOT
Kind
1
Data sources
AI-generated summary, drawn from Cloudflare, operator docs, and community data

What is Internet Archive - Archive-It?

Internet Archive - Archive-It is a web crawler operated by Archive-It. Internet Archive’s Archive-It service preserves publicly accessible web pages for the historical record.

Helpful — Verified, safe crawler. Respects robots.txt and provides operator documentation.

Identification

Bot name
Internet Archive - Archive-It
Operator
Archive-It
Category
Archiver
Kind
BOT
Verification
Verified bot
Cloudflare slug
internet-archive-archive-it
User-agent patterns
Archive-It
User-agent strings
Mozilla/5.0 (X11; Linux x86_64; special_archiver; Archive-It; +http://archive-it.org/files/site-owners-special.html) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36
Mozilla/5.0 (X11; Linux x86_64; archive.org_bot; Archive-It; +http://archive-it.org/files/site-owners.html) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36
Mozilla/5.0 (compatible; special_archiver; Archive-It; +@http://archive-it.org/files/site-owners-special.html)
Mozilla/5.0 (compatible; archive.org_bot; Archive-It; +@http://archive-it.org/files/site-owners.html)
Main use cases
Web crawlingData collection

Why it crawls your site

Internet Archive - Archive-It crawls websites to archive and preserve web content. It is operated by Archive-It as part of their web archival infrastructure. If you see this bot in your server logs, it is a verified crawler and generally safe.

Block or allow Internet Archive - Archive-It

Block it

Add a Disallow rule for Archive-It in your robots.txt file. You can also block at the server level using your web server configuration or CDN firewall rules to filter requests matching the user-agent string.

# Block Internet Archive - Archive-It from crawling your entire site
User-agent: Archive-It
Disallow: /

# Allow Internet Archive - Archive-It full access
User-agent: Archive-It
Allow: /
Allow & verify

Ensure your robots.txt allows Archive-It. Verify requests are genuine by checking the user-agent string and referring to Archive-It's documentation.

Traffic & trends

Cloudflare Radar traffic

No Cloudflare traffic data available.

Google Trends interest

Loading trend data…

Data sources

This profile is compiled from the following sources.

CloudflareRadar
Last updated July 26, 2026