PDF to Markdown converter and PDF text extractor for RAG and AI agents, including scanned PDFs via OCR (7 Latin-script languages). Convert a public PDF to text or page-aware Markdown with metadata, links and simple tables; OCR reads image-only pages instead of only flagging them.
Visual regression testing for public web pages: capture screenshots and compare them with another page or a PNG/JPEG baseline. Get pixel-level diff images and change percentages for QA, monitoring, CI, and AI agents.
Broken link checker and technical SEO audit for public sites: find broken internal links, duplicate titles, missing meta descriptions, canonical, heading, robots and hreflang issues. Crawl from a URL or XML sitemap for per-page fixes and a site summary. HTTP-only SEO crawler, no JavaScript.
Check and validate XML sitemaps: find them via robots.txt, follow sitemap indexes and gzip files, extract every URL with lastmod/changefreq/priority, diff against a previous run, and optionally check HTTP status. Broken sitemaps are reported, never skipped.