OCR images and scanned PDFs into text or Markdown with Tesseract. Supports 100+ languages, preprocessing, word confidence and bounding boxes, PDF quality summaries and low-confidence flags, plus input from Apify datasets, CSV/JSON/text manifests and public Google Sheets.
Booking.com reviews scraper for hotel intelligence. Discover properties by name or destination, export scores, text, guest/stay/room data, photos and partner replies, filter by traveler/language/date/stay details, run summary-only benchmarks, and monitor only new reviews.
Apple App Store reviews scraper API for ASO and product research. Export ratings, titles, review text, dates, app versions and helpful votes across multiple countries; filter before billing, include app metadata, batch apps and use retry-safe Apple public feeds.
Broken link checker and technical SEO crawler for dead links, 4xx/5xx errors, redirects, assets, canonicals and hreflang. Crawl sites or sitemaps, check URLs concurrently, flag slow responses, rank issues, and persist snapshots to surface new, regressed, resolved and changed links.
Data Cleaner API for JSON/CSV normalization, deduplication and CRM cleanup. Trim or collapse whitespace, lowercase emails, normalize phones and ISO dates, remove empty fields or rows, deduplicate by key and process nested objects with bounded batch limits.
Google Ads Transparency Center scraper API for competitor ad intelligence. Search brands, domains or advertiser IDs; resolve verified advertisers, paginate creatives, filter country, date and format, capture previews and optionally enrich creative details with proxy and retry controls.
Google Autocomplete keyword scraper for long-tail SEO research. Expand seeds with A–Z, numbers, questions, prepositions, comparisons and custom modifiers, recurse through suggestions, run country/language matrices, cap requests/results, and monitor new or disappeared suggestions.
Google News scraper for brand, competitor and market monitoring. Search by country, language and date; filter publishers, resolve URLs and optionally extract article text. Free analytics report publisher/domain share, daily/query volume, enrichment coverage and tracked keyword mentions.
Google Play reviews scraper for ASO and product research. Export filtered reviews, developer replies, helpful votes and app versions, plus free rating distributions, reply-rate analytics and rating-based positive/neutral/negative sentiment summaries.
Google Play reviews scraper for Android feedback intelligence. Find apps by name or package, export ratings, dates, text, helpful votes and developer replies; use multi-keyword/date/rating filters, app metadata, version/month/topic insights, proxy fallback and only-new monitoring.
Google Trends scraper and API alternative for daily trending searches and keyword research. Export interest over time and by region, related queries and topics, current trends, geo/time/category/property filters, retries and optional proxy routing.
HTML to Markdown converter and cleaner for LLM/RAG pipelines. Convert static or JavaScript pages, preserve headings, tables, links and images, remove noisy selectors, output clean text, and create overlap-aware chunks with stable IDs, SHA-256 fingerprints, word counts and token estimates.
Product Hunt scraper and launch monitor for startup research. Collect launches, ranks, upvotes, comments, makers, topics, media and links; optionally enrich public product pages with social, follower, review, funding, team and YC signals when exposed.
Schema markup extractor and structured data API for static or JavaScript-rendered pages. Extract JSON-LD, Microdata, RDFa, Schema.org types, metadata, E-E-A-T and LocalBusiness signals, plus technical SEO checks. Includes auto browser fallback, selector waits, proxy support and SSRF protection.
Shopify App Store reviews scraper plus competitor ranking intelligence. Export reviews, merchant usage, ratings and developer replies; discover apps by keyword, category or developer profile, capture app rankings and metadata, filter feedback, and monitor only newly appearing reviews.
Extract URLs from XML and gzipped sitemaps, nested indexes and robots.txt discovery. Export lastmod, priority, hreflang, image/video data, filter URLs, audit status/redirects and page indexability, fall back to crawling, and compare persistent snapshots for URL changes.
Steam reviews scraper for game intelligence. Filter by language, sentiment, purchase type, date, keywords, playtime, reviewer history, helpful votes, refunds, free copies, Early Access and developer responses; include game metadata, off-topic periods, weighted sentiment and only-new monitoring.
TripAdvisor reviews scraper for hotels, restaurants, attractions, activities, airlines and cruises. Discover places by search or location ID, filter by rating, language, date, traveler type, keyword, helpful votes and owner replies, export summaries, use proxy fallback and monitor only new reviews.
Trustpilot reviews scraper for deep reputation research. Go beyond the usual ~200-review anonymous window, filter by rating/language/date/verification/replies, discover brands, export business insights, and run incremental only-new review pulls with persistent monitoring.
Website screenshot API for URLs or raw HTML. Capture PNG, JPEG, WebP or PDF with full-page, element or region modes, device presets, retina, dark mode, lazy-load scrolling, cleanup, auth, locale, timezone, geo and proxy controls. Monitor hashes and pixel-change visual diffs with SSRF protection.
YouTube comments scraper for videos, Shorts, channels, playlists and search. Export comment/reply threads, authors, likes, dates, creator/hearted/pinned flags and video metadata; filter results, add insights, use proxy fallback and monitor only new comments.
Bulk YouTube transcript scraper for videos, Shorts, channels, playlists and search. Export clean text, paragraphs, timed segments, SRT, VTT and RAG chunks; choose or translate caption languages, use proxy/retry fallback and monitor transcript changes.
PDF parser and converter for structured JSON, CSV-ready rows and Markdown. Extract reading-order text, page-level Markdown, heuristic tables and document metadata with bounded file size, page, timeout and retry controls.
Recursive text splitter and RAG chunker for embeddings, vector databases and LLM context. Split on paragraph, line, sentence, word or hard boundaries, add overlap, stable chunk IDs and previous/next chunk links.