Point at ANY public SPARQL endpoint (Wikidata, DBpedia, UniProt, government & museum Linked-Open-Data) and run a query over the W3C SPARQL 1.1 Protocol. SELECT flattens to dynamic columns per binding, ASK to a boolean, CONSTRUCT/DESCRIBE to triples. One actor, any knowledge graph. Pay per binding.
Point at any Taiga instance (open-source agile PM) for one structured row per public project: name, slug, fans, watchers, activity, tags, owner, timestamps. Also fetch a project by slug, per-project agile stats, and fleet counts via the documented REST API.
Export public YouTube comments and replies from 1–10 videos. Compare videos with an equal per-video limit and a per-video coverage CSV; skip already-seen comment IDs on repeat runs. $0.0005 first delivered comment, then $0.000165 each; zero delivered rows = no Actor event charge.
Fetch and parse the IAB Tech Lab ad-tech transparency triad — ads.txt, app-ads.txt and sellers.json — for any domain or list of domains. One flat, normalized row per seller relationship, seller entry or variable, with a lossless raw copy. The authorized-digital-seller graph as data. Pay per record.
Collect public reviews from Google Play and the Apple App Store: rating, text, author, date. Apple may return only ~10 public-page reviews per app per country. Repeat monitoring: each run saves a Next run input that skips reviews you already have, plus a per-app coverage CSV. Pay per new review.
Point at any Bugzilla server (Mozilla, Red Hat, LibreOffice, Gentoo) and get one structured row per bug: summary, status, resolution, product, component, priority, severity, reporter, timestamps, keywords. Also lists products and server version.
Point at any Buildbot CI master (CPython, WebKit, LLVM, GNOME, self-hosted) and get one structured row per builder, build, or worker: names, tags, build results and status, timestamps, and worker inventory. One actor for every Buildbot data API v2 server.
Scrape clinical-trial records from the official ClinicalTrials.gov v2 REST API. Query by condition, intervention, sponsor, status, phase or free text. One record per study — NCT id, title, status, phase, sponsor, conditions, interventions, enrollment, dates, locations + the raw study. Pay per study.
Point at any Conda channel's repodata.json and get one structured row per package: name, version, build, subdir, license, md5/sha256, size, timestamp, depends and constrains.
Export validator data from a Cosmos-SDK chain's public LCD (REST) endpoint you choose: moniker, status, jailed, tokens, commission. The ready-to-run sample lists 20 bonded Cosmos Hub validators. Also reads governance proposals, supply, balances, delegations and inflation. Amounts stay raw strings.
Point at ANY Datasette instance and pull structured data via its uniform JSON API. Discover every database + table, page a table's rows into a clean dataset, or run a read-only SQL query. Works on any datasette publish site with no per-site scraper. Pay per record.
Download earthquake lists from USGS, EMSC, INGV, GEOFON, IRIS and other FDSN seismic data centers. Filter by date, magnitude, depth, area or radius; get one row per quake with time, lat/lon, depth, magnitude and place name. Seismic station metadata mode too. Pay per record.
Point at ANY HL7 FHIR R4 server (HAPI, SMART-on-FHIR, Azure/Google/AWS health APIs) and page any resource type (Patient, Observation, Condition, Encounter) into clean flat rows. Follows the Bundle next-link cursor, keeps the full resource in _raw. Discovery mode maps a server. Pay per record.
Point at any Gerrit code-review server (Android, Chromium, OpenStack, Wikimedia, GerritHub, self-hosted) and get one structured row per change: subject, status, owner, review labels, insertions/deletions, comments and merge history. Also lists projects and server info.
Point at any Grafana instance and extract its dashboard + folder inventory, version and health via the open HTTP API. One row per dashboard/folder with title, uid, tags, folder; plus health and discovery modes. No per-host scraper.
Monitor Hacker News for keywords, brands, domains or authors via the public Algolia HN API. Search stories + comments, filter by points, comments, date & type. One record per hit — title, author, points, URL, text + raw. Pay per item.
Point at any Helm chart repository's index.yaml and get one structured row per chart version: appVersion, digest, dependencies, maintainers, keywords, sources, deprecation and more.
Point at ANY IVOA TAP service (VizieR, Gaia, GAVO, MAST, ESO, CADC and thousands of observatory archives) and run an ADQL query over the Table Access Protocol. Query mode returns dynamic columns per your SELECT; tables mode discovers the full schema. One actor, every astronomical catalog.
Point at any Jenkins or Hudson controller (anonymous-read) and get one structured row per job or build: status, results, timings, build history, job graph. Works on any controller via the standard /api/json Remote Access API.
Point at ANY Matrix homeserver and extract its public room directory over the Client-Server API. One row per room: room ID, name, topic, member count, canonical alias, avatar, join rule, room type. Opaque next_batch pagination; optional token unlocks search/space filters.
Point at ANY Memento (RFC 7089) TimeMap and extract every archived snapshot as flat rows — memento URI, datetime (ISO 8601), rel, original URI, collection, raw link line. Works across the Internet Archive, arquivo.pt and every Memento-compliant web archive. Pay per memento.
Point at any open Netdata monitoring agent and export node info, the chart catalog, active alarms and fleet inventory from the public unauthenticated REST API. One actor over thousands of self-hosted Netdata nodes.
Point at ANY fediverse instance (Mastodon, Lemmy, Misskey, Pixelfed, PeerTube…) and extract its NodeInfo: software name & version, protocols, services, open-registrations, user counts, local posts & comments, plus the raw NodeInfo JSON. One generic runner for every fediverse server.
Point at ANY OAI-PMH endpoint (arXiv, DSpace, EPrints, Invenio, OJS and thousands more) and get flat metadata rows — identifier, datestamp, title, creators, subjects, publisher, date, rights — plus lossless raw XML. All 6 verbs, incremental from/until, resumptionToken paging. Pay per record.