Main text of the article and blog URLs you list, as plain text and Markdown, with title, author, date and language. Article-body F1 0.959 on a held-out half of a public benchmark. One request per URL, robots.txt obeyed; blocked pages reported, not bypassed.
Greenhouse jobs plus Lever, Ashby, Workable and Personio postings from public job boards in one flat schema with salary, location and remote fields. Change mode returns only new, changed or closed postings since your last run. No recruiter fields; emails and phone numbers removed from descriptions.
Measures audiobook chapter files against ACX's published technical specs (RMS, peak, noise floor, room tone, MP3 format) and returns mastered 192 kbps CBR 44.1 kHz MP3s with a before/after PASS/FAIL report per chapter.
Matches messy Australian addresses to the official G-NAF address file: returns the clean address, its G-NAF persistent identifier (PID), a confidence score and a match status for every input line. Paste addresses or give a CSV URL.
Parses X12 850 purchase orders and UN/EDIFACT ORDERS messages into plain JSON and CSV, with a per-message check of envelopes, control numbers, segment counts, required segments, dates and numbers. No EDI software needed.
Batch-checks HTML email templates for machine-checkable accessibility problems: missing alt text, low colour contrast, missing lang and title, layout tables without role=presentation, heading jumps, vague link text, tiny or all-caps text. One row and one report per template.
Google Trends data for a keyword list: interest over time, interest by region and related queries (top and rising), alone or in comparisons of up to 5. Paced, honest requests with back-off; every keyword gets a status row (ok, no data, rate limited). Not affiliated with Google.
Checks .h5p e-learning packages against the published H5P file specification: h5p.json and library.json fields, bundled library versions, content.json, file types, missing image alt text and oversize files. One report row per package, plus HTML/Markdown report.
Extracts the text from images (receipts, scanned pages, screenshots) with an open OCR model that runs inside the actor: your images are not sent to any other service. Returns plain text plus lines and words with boxes and confidence.
Checks a Google Merchant Center product feed (TSV, CSV, RSS 2.0 or Atom XML) against Google's published product data specification. One row per item with its problems, plus a corrected feed with only safe mechanical fixes and a change log. Never invents product data.
Checks ONIX for Books 3.0/3.1 XML feeds against EDItEUR's schema and code lists, plus rules the schema misses (ISBN-13 check digits, real dates, price, currency, territory, contributor). One row per issue with record, field path and line, plus an HTML/Markdown report.
PageSpeed Insights scores and real-user Core Web Vitals (Chrome UX Report p75 LCP, INP, CLS) for a list of URLs, one row per URL and device, from Google's official APIs with your own free Google API key. URL-level field data, origin fallback labelled.
Adds accessibility tags (headings, paragraphs, lists, tables, figures, links), language and title to untagged born-digital PDFs, then checks the input and output with veraPDF's PDF/UA-1 profile. Returns the tagged PDF and a before/after report.
Crawls a website and converts each page to clean Markdown for AI and RAG use. Plain HTTP by default (256 MB), spend cap on, priced per page.
Check your sitemap URLs, any list of URLs, or every link on your pages: status, redirect chain, noindex, canonical, robots.txt, and broken links with the page and anchor text they sit on. On-page SEO rules (title, meta description, H1, canonical, alt text, schema, hreflang) with fix hints.
Cleans a survey export (CSV, SPSS .sav or Qualtrics CSV) with fixed, documented rules and returns the cleaned file, a reconstructed codebook, and a row-level report of every change. Speeders, straight-liners and out-of-range answers are flagged, never deleted.
Reads the tables shown on screen in your own video file (screen recordings, slide talks, demos) and returns them as CSV and Excel. Every row cites the frame time it was read from and links to that frame's image. Scrolled tables are stitched into one.
Lists every PDF a website links to with the facts that decide accessibility work: tagged or not, title, language, text layer or scanned, pages, referring pages. Also lists Word files and Google Drive/Docs links. Machine checks only, not an audit. Obeys robots.txt.
Website screenshots of the URLs you list (first screen or full page, desktop or mobile, PNG/JPEG), each checked for blank page, HTTP error, bot-check page and height cut. Cookie banners hidden, never accepted. Honours robots.txt by default; blocked pages are reported, not bypassed.
Tech stack detector for websites: finds the CMS, ecommerce platform, JavaScript framework, analytics, CDN and hosting a site uses, with evidence for each finding. One polite HTTP request per URL. Accepts URLs or bare domains; works from the API and AI agents (MCP). Honours robots.txt.
Transcribe audio and video files from direct links, or a podcast RSS feed, to text, SRT, WebVTT and JSON with timestamps. Speech to text by an open Whisper model inside the run, so audio is not sent to an outside AI service. Language auto-detect and a max-minutes limit.