Find Wayback Machine snapshots of a URL. Mode "closest" returns the single nearest capture to a given timestamp via the Availability API, falling back to a CDX closest-capture query when the Availability API has no answer (one result). Mode "history" returns the full capture list via the CDX API — filterable by date range, HTTP status code, and MIME type, collapsed by default to one capture per day (collapse=timestamp:8). Use history mode to survey how a page changed over time; use closest mode when you need the snapshot nearest a specific date. history mode supports resume-key pagination for URLs with very large capture histories.
Fetch the archived content of a URL at a specific Wayback Machine timestamp. Resolves to the nearest available capture when the exact timestamp has no snapshot. Returns the archived page as readable plain text (HTML stripped) and the replay URL of the capture Wayback served, for browser access. Use ia_find_snapshots first to discover valid timestamps for a URL.
Search the Internet Archive library (40M+ items) using the Advanced Search (Solr) API. Filter by media type (texts, audio, movies, software, image, etc.), collection, creator, date range, and language. Sort by relevance, date, or downloads. Supports pagination via page/rows. Returns identifiers, titles, creators, media types, dates, download counts, total_found, and current page/rows for pagination context. Use ia_get_item with a returned identifier to get full metadata and file manifests.
Retrieve metadata and the file manifest for an Internet Archive item by identifier. Returns title, creator, description, subjects, collections, license, language, the total file count, and a page of files, each with its format, size, and direct download URL. Pages hold up to max_files files (default 50) starting at file_offset; set format to keep a single file type (e.g. "DjVuTXT", "Text PDF", "VBR MP3"). The primary tool to act on a search result from ia_search_items. Use ia_get_text to retrieve the readable text of a text item.
Retrieve the readable text content of a text item (OCR DjVuTXT or plain-text file) from the Internet Archive, with length-aware truncation and a continuation pointer for pagination. Suited for public-domain books, documents, scanned periodicals, and transcripts. Use max_chars and char_offset to page through long documents. To confirm the item has a text file first, call ia_get_item with format "DjVuTXT"; its response also gives the mediatype.