Point at any self-hosted Harbor container registry and extract its public projects, repositories, image artifacts (with vulnerability-scan severity) and instance fingerprint via the open /api/v2.0 REST API. One actor, any Harbor — no per-host scraper.
Point at ANY IIIF manifest or collection (Wellcome, Bodleian, Harvard, Library of Congress and thousands more) and get flat metadata rows: label, description, metadata pairs, rights, thumbnail, canvas count, plus per-canvas image info. Handles IIIF Presentation API 2.x and 3.x. Pay per manifest.
Point at ANY IVOA Simple Cone Search service (VizieR, HEASARC, GAVO, IRSA, NOIRLab and thousands of observatory catalogs) and pull every catalog source within a sky circle: give RA, Dec and a radius, get one row per object. One actor, the whole Virtual Observatory.
Point at ANY IVOA Simple Image Access (SIA) service — IRSA, SkyView, GAVO, HEASARC, ESO, CADC and thousands of observatory image archives — and list every astronomical image over a sky region: give RA, Dec and a size, get one row per image with its access URL. One actor for the Virtual Observatory.
Point at ANY IVOA Simple Spectral Access (SSA) service — ESO, GAVO and observatory spectral archives — and list every astronomical spectrum over a sky region: give RA, Dec and a size, get one row per spectrum with its access URL. One actor for the Virtual Observatory.
Point at ANY Lemmy instance (lemmy.world, sh.itjust.works, lemm.ee…) and pull instance stats, posts, comments or communities into clean flat rows. One actor reads every federated /api/v3 server. Sort + community filters, auto-paged, lossless _raw. Pay per item.
Point at any apt (Debian Packages) or yum/dnf (repomd/primary.xml) repository and extract every binary package's metadata — name, version, deps, checksums, size, license. Generic repo-index parser for SBOM, supply-chain audit, and mirror monitoring.
Point at ANY MediaWiki wiki's Action API (Wikipedia, Fandom, wiki.gg, or any corporate/OSS wiki) and pull structured data: full-text search, enumerate pages, list category members, fetch page text + revisions, or read site info. Handles continue-token pagination. Pay per record.
Point at ANY Mobilizon instance and extract events + groups over its public GraphQL API. Browse or full-text-search the federated events/groups directory, or fetch one event/group by id. One clean row each: title, dates, category, location, organizer, tags, member counts.
Point at ANY OGC API - Coverages endpoint (the JSON/CoverageJSON successor to WCS) — pygeoapi, GNOSIS, ldproxy, GeoServer. List coverage collections, describe their schema (axes, params, CRS), or summarize CoverageJSON (axis shape + value count) as flat rows. Pay per record.
Point at ANY OGC API - Tiles endpoint (the JSON-native successor to WMTS) — ldproxy, GNOSIS, GeoServer, pygeoapi. Inventory vector/map tileset offerings, list tile matrix sets (CRS, zoom levels, axes), or discovery-scan a server's tiles capabilities as flat rows. Pay per record.
Point at ANY OGC API - EDR (Environmental Data Retrieval) service — Met Office, FMI, pygeoapi. List collections, then query environmental data by position, area, radius, trajectory, or cube. Flattens CoverageJSON / GeoJSON to rows (parameter, value, unit, lon/lat, time). Pay per record.
List the data sources behind a public or authorized OHDSI ATLAS / WebAPI server: source names, keys, SQL dialects and CDM/vocabulary/results daimons, one flat row each. Metadata only, never patient records. Optional modes: cohort definitions, concept sets, vocabulary search. Pay per row.
Point at ANY German OParl council-information system (Ratsinformationssystem) for flat rows — papers (Drucksachen), meetings, committees, people. One actor, every vendor (STERNBERG, more-rubin, CC e-gov): reads list URLs from the Body, follows links.next paging, skips dead endpoints. Pay per record.
Point at ANY OpenActive dataset site, data catalog, or RPDE feed and harvest UK sports & activity opportunity data — classes, sessions, facility slots and events from 100s of leisure operators (GLL, Everyone Active, Halo...). Follows the RPDE cursor, handles tombstones. Pay per item.
Scrape scholarly works (papers, preprints, books), authors, institutions & journals from the free public OpenAlex API. Search + filter by author, institution, concept, year, type & open-access; sort; auto-paginate. Clean columns + raw. Pay per record.
Point at ANY OpenDataSoft portal (Paris, RTE, data.opendatasoft.com + thousands more) and pull real ROW data via the Explore API v2.1. List the dataset catalog or extract records with ODSQL where/select/order_by/group_by + facets. Auto-paginates to flat rows + geo. Pay per record.
Point at ANY openEO back-end (the standard API for Earth-observation cloud back-ends — VITO/Terrascope, Copernicus Data Space, EODC). Extract collections, the process registry, capabilities, file formats, service types or UDF runtimes as flat rows + raw JSON. No auth. Pay per record.
Point at any PeerTube instance and export its videos, channels and instance metadata from the public /api/v1 REST API. Federated, self-hosted YouTube alternative — one actor over thousands of instances.
Point at any Prometheus HTTP API v1 server (Prometheus, Thanos, Mimir, VictoriaMetrics, Grafana Agent) and export PromQL query results, range series, label values, scrape targets, rules, metric metadata and build info as clean rows. One actor over every Prometheus-compatible server.
Point at ANY W3C Reconciliation (OpenRefine) service and pull results: fetch the capability manifest, batch-match query strings to candidate entities with scores, autocomplete (suggest), and fetch property values (extend). Works with Wikidata, GND, Getty and any conforming endpoint.
Point at ANY RPKI validator's VRP export (rpki-client, Routinator, FORT, Cloudflare) and extract validated ROAs: origin ASN, prefix, maxLength, trust anchor, expiry. Optional ASN / prefix / trust-anchor filters. Auto-detects the array key. Seeded on Cloudflare & rpki-client.org.
Point at ANY RSS or Atom feed (news, blogs, podcasts, GitHub releases, YouTube) and get clean structured rows. Auto-detects RSS 2.0 / RSS 1.0 / Atom and normalizes every entry to one schema: title, link, content, author, categories, ISO-8601 dates, enclosures. Pay per item.
Scrape U.S. SEC EDGAR company filings + metadata from the official public JSON APIs. Query by ticker, CIK, or full-text keyword; filter by form type (10-K, 10-Q, 8-K…) and filing-date range. One clean record per filing with the index + document URLs. Pay per filing.