Batch WebP/JPEG/PNG optimization for image URL lists (or downloader dataset/KV). KV + bytesBefore/After, savedPct, EXIF strip. Failed free. AVIF: 「runtime optional / currently unavailable on this runtime」— not a selling point. No crawl.
Download image URL lists or extract img/og:image from pages into KV store with metadata (contentType, bytes, sha256, dims). Default under 1 GB; failed/empty free. For RAG/archives — not spam scraping.
Batch domain availability verdicts via RDAP (with DNS/WHOIS fallback) — available|registered|unknown per domain; 256 MB default; failed/unknown free by default. Not lead-gen: inventory & availability checks only.
Convert PDF and Word (DOCX) files to LLM-ready Markdown. Tables stay real Markdown tables (76% of rows exact on unseen statistical PDFs vs 31% without our repair), plus headings, page markers and optional RAG chunks with page range and section path. Batch URLs or uploads. No AI keys.
Parse RSS 2.0 and Atom feeds into structured JSON with Markdown body text for RAG and LLM pipelines. Optional feed discovery from a page, content hash, and heading-aware chunks. Failed feeds are reported, not fatal. 256 MB default.
OCR scanned PDFs and images into Markdown with RapidOCR (CJK-capable, onnxruntime). Page markers, optional RAG chunks, geometric table assembly. No external AI API keys.
Bulk extract Open Graph, Twitter Cards, JSON-LD, microdata, RDFa and basic meta tags from a URL list. Pure structured JSON — not an SEO audit score. Chain from sitemap / URL status. Default 256 MB. No browser, no AI keys.
Discover every URL in XML sitemaps via robots.txt and indexes; tag PDF, DOCX, XLSX and more for SEO audits and RAG ingest. Auto-finds nested indexes and .xml.gz; exports DOC_TO_MARKDOWN_INPUT for the PDF/DOCX Actor.
Bulk-check HTTP status with HEAD→GET fallback, final URL, timing, and error class — chain from Sitemap Actor output. Error class dns/timeout/ssl/http; accept URL list, sitemap datasetId, or DOC_TO_MARKDOWN_INPUT.
Query the Wayback Machine CDX index, fetch selected snapshot HTML, and extract clean Markdown with trafilatura for RAG pipelines. Optional heading-aware chunks. Competitors often stop at CDX URL lists — this Actor returns page body Markdown.
Batch WHOIS, DNS (A/AAAA/MX/NS/TXT) and TLS cert lookup for domain lists — 256 MB default; failed domains free by default. One row per domain: registrar/dates + DNS + SSL days-left; no paid WHOIS API keys.