internet-archive-mcp-server

v0.3.1 pre-1.0

Search the Wayback Machine and IA library (40M+ items), fetch archived snapshots, retrieve item metadata and full text via MCP. STDIO or Streamable HTTP.

internet-archive.caseyjhand.com/mcp
claude mcp add --transport http internet-archive-mcp-server https://internet-archive.caseyjhand.com/mcp
codex mcp add internet-archive-mcp-server --url https://internet-archive.caseyjhand.com/mcp
{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "url": "https://internet-archive.caseyjhand.com/mcp"
    }
  }
}
gemini mcp add --transport http internet-archive-mcp-server https://internet-archive.caseyjhand.com/mcp
{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "command": "bunx",
      "args": [
        "mcp-remote",
        "https://internet-archive.caseyjhand.com/mcp"
      ]
    }
  }
}
{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "http",
      "url": "https://internet-archive.caseyjhand.com/mcp"
    }
  }
}
curl -X POST https://internet-archive.caseyjhand.com/mcp \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"curl","version":"1.0.0"}}}'

Tools

5

ia_find_snapshots

open-world

Find Wayback Machine snapshots of a URL. Mode "closest" returns the single nearest capture to a given timestamp via the Availability API, falling back to a CDX closest-capture query when the Availability API has no answer (one result). Mode "history" returns the full capture list via the CDX API — filterable by date range, HTTP status code, and MIME type, collapsed by default to one capture per day (collapse=timestamp:8). Use history mode to survey how a page changed over time; use closest mode when you need the snapshot nearest a specific date. history mode supports resume-key pagination for URLs with very large capture histories.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "ia_find_snapshots",
    "arguments": {
      "url": "<url>",
      "mode": "<mode>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "minLength": 1,
      "description": "The URL to look up in the Wayback Machine."
    },
    "mode": {
      "type": "string",
      "enum": [
        "closest",
        "history"
      ],
      "description": "Lookup mode: \"closest\" returns the single nearest snapshot to the given timestamp; \"history\" returns the paginated full capture list via the CDX API."
    },
    "timestamp": {
      "description": "Target timestamp in YYYYMMDDHHMMSS format (or any prefix thereof). Required for mode=closest. Example: \"20200101\" for January 1, 2020.",
      "type": "string"
    },
    "from": {
      "description": "Start of date range filter in YYYYMMDD format (history mode only). Example: \"20200101\".",
      "type": "string"
    },
    "to": {
      "description": "End of date range filter in YYYYMMDD format (history mode only). Example: \"20231231\".",
      "type": "string"
    },
    "status_filter": {
      "description": "Filter CDX results to a specific HTTP status code (history mode only). Must be a 3-digit numeric code. Example: \"200\" to return only successful captures.",
      "anyOf": [
        {
          "type": "string",
          "const": ""
        },
        {
          "type": "string",
          "pattern": "^\\d{3}$",
          "description": "3-digit HTTP status code."
        }
      ]
    },
    "limit": {
      "default": 100,
      "description": "Maximum number of CDX records to return (history mode only, default 100).",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000
    },
    "collapse": {
      "description": "CDX collapse parameter (history mode only). Format: \"timestamp:N\" where N is 1–14 (4=year, 6=month, 8=day, 10=hour, 14=exact). Default \"timestamp:8\" collapses to one per day. Pass an empty string or omit to get uncollapsed results.",
      "anyOf": [
        {
          "type": "string",
          "const": ""
        },
        {
          "type": "string",
          "pattern": "^timestamp:\\d{1,2}$",
          "description": "Collapse precision string."
        }
      ]
    },
    "resume_key": {
      "description": "Opaque pagination key from a previous history mode response. Pass this to continue from where the last page left off.",
      "type": "string"
    }
  },
  "required": [
    "url",
    "mode",
    "limit"
  ],
  "additionalProperties": false
}
view source ↗

ia_get_snapshot

open-world

Fetch the archived content of a URL at a specific Wayback Machine timestamp. Resolves to the nearest available capture when the exact timestamp has no snapshot. Returns the archived page as readable plain text (HTML stripped) and the replay URL of the capture Wayback served, for browser access. Use ia_find_snapshots first to discover valid timestamps for a URL.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "ia_get_snapshot",
    "arguments": {
      "url": "<url>",
      "timestamp": "<timestamp>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "minLength": 1,
      "description": "The URL whose archived content to retrieve."
    },
    "timestamp": {
      "type": "string",
      "description": "Target Wayback timestamp in YYYYMMDDHHMMSS format (or any prefix). The nearest available snapshot will be resolved and fetched. Example: \"20200101120000\" for noon on January 1, 2020."
    }
  },
  "required": [
    "url",
    "timestamp"
  ],
  "additionalProperties": false
}
view source ↗

ia_search_items

open-world

Search the Internet Archive library (40M+ items) using the Advanced Search (Solr) API. Filter by media type (texts, audio, movies, software, image, etc.), collection, creator, date range, and language. Sort by relevance, date, or downloads. Supports pagination via page/rows. Returns identifiers, titles, creators, media types, dates, download counts, total_found, and current page/rows for pagination context. Use ia_get_item with a returned identifier to get full metadata and file manifests.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "ia_search_items",
    "arguments": {
      "query": "<query>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "minLength": 1,
      "description": "Solr query string. Supports field prefixes such as title:\"war and peace\", creator:dickens, subject:history. Plain keywords search all fields."
    },
    "mediatype": {
      "description": "Filter by media type: texts, movies, audio, software, image, data, web, collection, etree (live concert recordings), or account. Case-insensitive; text, book, and books mean texts, movie, video, and videos mean movies, images means image, and collections means collection. Any other value is rejected.",
      "type": "string"
    },
    "collection": {
      "description": "Filter to items within a specific collection identifier, e.g. \"gutenberg\" or \"librivoxaudio\".",
      "type": "string"
    },
    "creator": {
      "description": "Filter by creator name, e.g. \"Charles Dickens\".",
      "type": "string"
    },
    "date_from": {
      "description": "Start of date range filter in YYYY-MM-DD format. Example: \"1900-01-01\".",
      "type": "string"
    },
    "date_to": {
      "description": "End of date range filter in YYYY-MM-DD format. Example: \"1999-12-31\".",
      "type": "string"
    },
    "language": {
      "description": "Filter by language code or name, e.g. \"eng\" or \"English\".",
      "type": "string"
    },
    "sort": {
      "description": "Sort order in Solr format. Examples: \"downloads desc\", \"date asc\", \"titleSorter asc\". Default: \"downloads desc\".",
      "type": "string"
    },
    "rows": {
      "default": 50,
      "description": "Number of results per page (default 50, max 200).",
      "type": "integer",
      "minimum": 1,
      "maximum": 200
    },
    "page": {
      "default": 1,
      "description": "Page number (1-indexed). Combine with rows for pagination.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "query",
    "rows",
    "page"
  ],
  "additionalProperties": false
}
view source ↗

ia_get_item

open-world

Retrieve metadata and the file manifest for an Internet Archive item by identifier. Returns title, creator, description, subjects, collections, license, language, the total file count, and a page of files, each with its format, size, and direct download URL. Pages hold up to max_files files (default 50) starting at file_offset; set format to keep a single file type (e.g. "DjVuTXT", "Text PDF", "VBR MP3"). The primary tool to act on a search result from ia_search_items. Use ia_get_text to retrieve the readable text of a text item.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "ia_get_item",
    "arguments": {
      "identifier": "<identifier>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "identifier": {
      "type": "string",
      "minLength": 1,
      "description": "Internet Archive item identifier, e.g. \"prideprejudice00aust\" (Pride and Prejudice). Obtain from ia_search_items results."
    },
    "format": {
      "description": "Keep only files whose format equals this value, ignoring case — e.g. \"DjVuTXT\", \"Text PDF\", \"VBR MP3\". Applied before paging. Omit to list every format.",
      "type": "string"
    },
    "max_files": {
      "default": 50,
      "description": "Maximum number of files to return (1–500, default 50).",
      "type": "integer",
      "minimum": 1,
      "maximum": 500
    },
    "file_offset": {
      "default": 0,
      "description": "Zero-based position in the (format-filtered) file list to start from (default 0). To read the next page, add max_files to the previous file_offset.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "identifier",
    "max_files",
    "file_offset"
  ],
  "additionalProperties": false
}
view source ↗

ia_get_text

open-world

Retrieve the readable text content of a text item (OCR DjVuTXT or plain-text file) from the Internet Archive, with length-aware truncation and a continuation pointer for pagination. Suited for public-domain books, documents, scanned periodicals, and transcripts. Use max_chars and char_offset to page through long documents. To confirm the item has a text file first, call ia_get_item with format "DjVuTXT"; its response also gives the mediatype.

read
invocation
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "ia_get_text",
    "arguments": {
      "identifier": "<identifier>"
    }
  }
}
schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "identifier": {
      "type": "string",
      "minLength": 1,
      "description": "Internet Archive item identifier, e.g. \"prideprejudice00aust\" (Pride and Prejudice). Obtain from ia_search_items results."
    },
    "max_chars": {
      "description": "Maximum number of characters to return in this response. Defaults to the server-configured maximum (IA_MAX_SNAPSHOT_CHARS, typically 50 000). Lower values reduce token usage.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "char_offset": {
      "default": 0,
      "description": "Character offset to start reading from (default 0). To read the next page, add max_chars to the previous char_offset.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "identifier",
    "char_offset"
  ],
  "additionalProperties": false
}
view source ↗

Resources

1

Metadata snapshot for an Internet Archive item — title, creator, mediatype, description, subjects, collections, date, license, and file count. Provides stable injectable context for agents that need to reference an item by URI without fetching the full file manifest.

uri ia://item/{identifier} mime application/json