Scrape Grailed designer and streetwear resale listings — active and SOLD — via the Algolia search API. Returns full listing data including sold price, seller score, original price, and arbitrage-ready comparables.
Scrape job listings from Gupy — Brazil's #1 ATS used by 3,000+ enterprises including Itaú, Petrobras, and Magalu. Search by keyword, workplace type, or state. Returns job title, company, location, remote status, PCD flag, description, and apply URL.
Scrape the full Project Gutenberg catalog via the Gutendex JSON API. Filter by search, language, subject, author era, and download count. Returns EPUB, Kindle, plain-text, and HTML download URLs — built for AI training corpora, NLP datasets, and TTS pipelines.
Scrape Indeed job listings by keyword and location. Extract job titles, companies, locations, salaries, job types, posting dates, and full descriptions. Perfect for market research, recruiting intelligence, and hiring trend analysis.
Scrapes the US insurance carrier directory from the NAIC Consumer Information Source — ~5,200 licensed carriers with NAIC code, carrier name, licensed states, lines of business, address, phone, and website. Covers PC, life, health, surplus-lines, and captive carriers.
Scrape Kenbiya (建美家 / kenbiya.com) — Japan's #2 investment-property portal
after Rakumachi. Extracts yield, monthly rent, occupancy, structure, and
broker data from ~20-40K active listings. Natural companion to
rakumachi-investment-scraper for full Japan investor-market coverage.
Search Legacy.com's obituary and death-notice database by name or by newspaper, or pull the most recently posted notices across all outlets. Returns structured records: full name, age, dates, city/state, newspaper affiliate, and funeral home name + website.
Scrape public room metadata from LINE OpenChat — Japan's largest public chat platform, also popular in Taiwan and Thailand. Extracts room name, description, member count, hashtags, region, and more from the public OpenChat directory. No LINE account required. Supports JP, TW, and TH markets.
Scrape live listings from magi (magi.camp) — Japan's leading C2C trading-card marketplace. Search by keyword (Pokémon, Yu-Gi-Oh!, One Piece, MTG, etc.) and collect listing prices, sold signals, favorite counts, badges, and image URLs.
Search Meta Ad Library ads by keyword, advertiser page, country and ad type. Returns spend and impressions ranges, reach estimates, per-country reach, payer byline, state-media and AI-media disclosure flags, and creative-variant collation counts the standard scrapers omit.
Bulk scraper for Mexico's DENUE business registry — 5.5M+ establishments with geo coordinates, SCIAN industry codes, employee ranges, contact info, and addresses. Filter by state, activity keyword, and entity type. Requires a free INEGI API token (email registration at inegi.org.mx).
Search and look up music metadata from MusicBrainz — the open music encyclopedia. Retrieve artists, release groups, recordings, labels, and works with full relationship graphs, genre tags, and cross-platform ID resolution (Spotify, Discogs, ISRC, Wikidata, and more). No API key required.
Scrape the official NFL schedule from nfl.com. Returns every game for a season -- home/away teams, kickoff time, venue, broadcast network (CBS/FOX/NBC/ESPN/TNF), game status, primetime flags, and international game markers.
Scrapes Texas Railroad Commission (RRC) oil and gas well data. Returns API number, operator, lease, field, county, district, well type, on-schedule status, and drilled depth. Covers all 13 RRC districts and over 300k wells. Built for mineral rights research, energy analytics, and landman databases.
Search Pastebin, GitHub Gist, Ideone, Paste.org and Textbin for public pastes mentioning your keywords - leaked credentials, API keys, hostnames, or any other exposure. Returns each match's title, URL and snippet, per keyword and site.
Scrape board-certified plastic surgeons from the ASPS find-a-surgeon directory. Returns name, board certs, practice address, phone, email, website, languages, NPI, and lat/lon — structured for B2B lead-gen in aesthetic medicine.
Normalizes published prompt-injection and jailbreak datasets from HuggingFace and GitHub research repos into one labeled corpus: technique, target model, defense bypassed, license, cross-source dedup. Defensive only — aggregates public data for guardrail/eval testing, never targets a live LLM.
Extract government surplus auctions from PublicSurplus.com across every browse category - vehicles, heavy equipment, computers, furniture, industrial gear and more - from roughly 9,000 US and Canadian government sellers, with agency, location, bid, and sold-price data.
Scrape the QueryTracker literary agent directory — the #1 querying-author platform. Extracts agent name, agency, genres, query status, query method, social links, and last-updated date from every public agent profile. Ideal for query-CRM tools and literary data pipelines.
Scrapes the complete product catalog from REP Fitness and Titan Fitness — the two dominant direct-to-consumer strength-equipment retailers on Shopify. Returns normalized product records with titles, variants, pricing, availability, and images from both sources in a single dataset.
Scrape running shoe catalogs from Brooks, Saucony, Altra, On, and Running Warehouse — normalized records with runner-relevant specs (drop, stack, foam, plate, weight).
Scrape ranked startup lists from SeedTable — startups across 4,000+ city lists worldwide. Returns startup name, industry, location, profile URL, and logo for every ranked company, and filters to just the cities you care about.
Scrape curated highlight stories from public Snapchat profiles. Provide usernames and get direct media URLs, thumbnails, story titles, and creator metadata — one record per story snap.
Scrape Social Blade's YouTube top-creator leaderboard. Extracts rank, username, subscribers, video views, and 30-day subscriber growth for the top creators.