Transform your Apify dataset or a CSV, JSON or JSONL file from a URL: filter rows, remove duplicates by any fields, sort, pick and rename fields, flatten nested JSON. Output to a dataset, plus an optional CSV or JSONL file (JSON to CSV in one step). Pay per output row.
Speech to text for audio and video: transcribe files from URLs and podcast RSS feeds into text, SRT subtitles and WebVTT captions with Whisper, run inside the actor. Auto-detects 99 languages, can translate to English, and can return only new episodes. Pay per audio minute; files that fail are free.
Search academic and research papers in OpenAlex and Crossref (incl. arXiv preprints) by keywords, title, author, DOI, arXiv id, date range and open access. One clean record per paper: authors, venue, DOI, citation count, open-access PDF link, and the abstract where its licence allows.
Turn web pages, or another Actor's dataset, into clean JSON with an LLM. Give a JSON schema or list the fields in plain English. Output is validated against your schema; pages that can't be fetched or don't match aren't charged. $5 per 1,000 pages plus the model's tokens.
Extract news and blog articles from any URL, or find the newest via a site's RSS/Atom feed or sitemap. One clean row each: title, author, published and modified dates, lead image, language, main text as plain text and Markdown, word count, canonical URL. Only-new-articles mode for monitoring.
Type company names or paste Ashby job-board URLs and get every open job from those companies' Ashby career pages: full descriptions, structured salary where published, one clean format. Filter by keyword, place, department, salary or date; get only new, changed or closed jobs since your last run.
Scrape ATS jobs from company career pages: enter company names or careers URLs and get every open job, live from each company's Greenhouse, Lever, Ashby, Recruitee, Personio or Teamtailor board, in one format. Filter by keyword, place, department, job type, salary or date, or get only new jobs.
Download all images from web pages you list, or from direct image links, into your key-value store, optionally as a ZIP. One row per image: page, image URL, alt text, width, height, type, bytes, SHA-256. Duplicates saved once; icons and tracking pixels skipped.
Official company data without an API key: search French companies by name, SIREN, SIRET or VAT, and any company worldwide by name or LEI (GLEIF), or monitor new company registrations in France daily (BODACC) by département and NAF code. One clean schema, LEI<->SIREN links. Companies only.
Convert PDF to text, and Word (DOCX) and Excel (XLSX) files too: give document URLs, get plain text, Markdown for LLMs and RAG, tables as arrays, and metadata. Scanned PDF pages and images are read with OCR in 32 languages. Pay per document, plus per OCR page.
Search EU and UK public procurement tenders in one format: TED (Tenders Electronic Daily), the EU's official procurement journal, and the UK's Find a Tender, via their official public APIs. Filter by keyword, CPV code, country, buyer, notice type and date; get deadlines, values, buyers and links.
Type company names or paste Greenhouse job-board URLs and get every open job from those companies' Greenhouse career boards: full descriptions, salary where published, one clean format. Filter by keyword, place, department, salary or date; get only new, changed or closed jobs since your last run.
Turn a list of companies into sales and recruiting signals: new job postings matching your skills or tools (Rust, Kafka, Snowflake), hiring surges, and technologies added to or dropped from their website. Reads Greenhouse, Lever, Ashby, Recruitee, Personio and Teamtailor. Pay only for signals.
Extract text from images in bulk with Tesseract OCR: image URLs (PNG, JPEG, WebP, TIFF, GIF, BMP or one-page scanned PDFs) in, text out, with a confidence score, image size and optional line and word boxes. 32 languages incl. Chinese, Japanese and Arabic. Pay per image.
Type company names or paste Lever job-board URLs (US and EU) and get every open job from those companies' Lever career pages: full descriptions, salary where published, one clean format. Filter by keyword, place, department, salary or date; get only new, changed or closed jobs since your last run.
Get price, currency, stock, GTIN (UPC/EAN), SKU, MPN, brand, name and image from any shop's product pages, read from their schema.org Product data. Paste product URLs or whole stores (Shopify, WooCommerce and more, via sitemaps). Monitor mode returns only changed prices and stock.
Remote and work-from-home job listings from We Work Remotely, Remote OK and Python.org Jobs in one clean format, with salary, employment type and categories. Filter by keyword, region, category or date, merge duplicates across boards, or get only jobs new since your last run. Links to each source.
Read and monitor RSS 2.0, RSS 1.0 (RDF), Atom and JSON Feed from feed or website URLs (feeds found automatically). One clean JSON row per item: title, link, UTC dates, author, plain-text and Markdown summary and full text, categories, enclosures. Keyword and date filters, only-new-items mode.
Get SEC EDGAR filings by ticker, CIK or company name, or every new filing of a form type: 10-K, 10-Q, 8-K, Form 4, 13F, Form D, S-1. Filter by form and date, add the document text as Markdown, and schedule it with only-new mode. Live from the SEC's official APIs.
SEO audit and website audit for your own site: crawl it from a start URL or sitemap and get a 0-100 SEO score per page and site from a published formula, plus broken links and images, redirects, title and meta description length, H1s, canonical, robots, hreflang, Open Graph and JSON-LD.
Extract all URLs from a website's XML sitemaps: finds them via robots.txt, follows sitemap indexes, reads .xml.gz. Returns lastmod, changefreq, priority, image, news, video and hreflang data. Filter by URL pattern and date, or monitor sitemap changes: only new and removed URLs since the last run.
Detect any website's tech stack: CMS, e-commerce platform, frameworks, analytics, tag managers, CDN, hosting and JavaScript libraries, with version, confidence and evidence, from the MIT-licensed Wappalyzer fingerprints plus our own. Monitor a list: get only sites that added or dropped a technology.
Turn web pages or whole sites into clean Markdown for RAG and LLMs: main content only, with title, links and word count. Whole-site mode reads the sitemap first, can return only changed pages, and splits pages into heading-based chunks. Respects robots.txt and AI opt-outs. No browser.
Check broken links and bulk URL status codes: for each URL, the HTTP status code, every redirect hop, final and canonical URL, response time, content type and size. Paste URLs, check a whole sitemap, or scan your pages for broken links. HEAD with GET fallback, robots.txt respected.