API

Read-only JSON API over the PrivacyWatch dataset: 96 LLM provider surfaces with data privacy, retention, model-training, and zero-data-retention fields, each backed by verbatim evidence quotes. No API key, no account, no writes. The same reference is available as docs/API.md in the repository.

#Quick start

# Providers that don't train on API data by default and offer self-serve ZDR
curl "https://privacywatch.wyrdwerk.com/api/v1/providers?training=off&zdr=self-serve"

# One full provider row, including evidence quotes and incidents
curl "https://privacywatch.wyrdwerk.com/api/v1/providers/anthropic-api"

# Dataset version, counts, and the exact filter vocabularies
curl "https://privacywatch.wyrdwerk.com/api/v1/meta"

Every response is JSON. List endpoints wrap rows in { "meta": { version, lastUpdated, count, license }, "data": [...] }.

#Endpoints

GET /api/v1/providers

Summary rows for all 96 provider surfaces, filtered in memory. Use this for comparisons, dashboards, and “show me providers that…” queries. Parameters are listed in the filters table; all are optional. Add fields=full for complete rows (same shape as /api/v1/providers/{id}) or ids=a,b,c to compare a few providers side by side.

Example response (?training=off&zdr=default&rating=clean, trimmed to 1 of 14 matches):

{
  "meta": { "version": "1.13.0", "lastUpdated": "2026-09-30", "count": 14, "license": "CC-BY-4.0" },
  "data": [
    {
      "id": "fireworks-ai", "name": "Fireworks AI", "surface": "API", "surfaceType": "api",
      "category": "inference", "rating": "clean", "incident": false,
      "training": "off", "trainingOptOut": "not-needed",
      "retention": "fixed", "retentionDays": 30, "zdr": "default",
      "regions": ["GLOBAL", "US", "CA", "EU"], "sourceDate": "2026-09-30",
      "url": "/api/v1/providers/fireworks-ai"
    }
  ]
}

Errors:

GET /api/v1/providers/{id}

One full provider row, including evidence[] (source URL + verbatim quote + retrieval date), incidents[], and compliance fields. IDs are stable and lowercase, e.g. openai-api, anthropic-api, deepseek-api.

Example response (anthropic-api, trimmed):

{
  "meta": { "version": "1.13.0", "lastUpdated": "2026-09-30", "license": "CC-BY-4.0" },
  "data": {
    "id": "anthropic-api", "name": "Anthropic", "surface": "API", "surfaceType": "api",
    "category": "us-frontier", "rating": "clean", "incident": false,
    "training": { "label": "Off by default", "default": "off", "optOut": "not-needed",
                  "detail": "No training on commercial API data. Opt-in via feedback submission only." },
    "retention": { "label": "30 days auto-delete", "kind": "fixed", "days": 30,
                   "detail": "Inputs/outputs auto-deleted within 30 days. ZDR = immediate post-response." },
    "zdr": { "label": "Yes — enterprise approval", "status": "full", "access": "approval",
             "detail": "ZDR available for enterprise API customers. Requires Anthropic account team + per-org enablement." },
    "location": { "label": "US storage; global inference", "regions": ["GLOBAL", "US"],
                  "detail": "Data stored in the US. Inference may run in any geography by default; US-only inference available at 1.1x pricing." },
    "compliance": { "dpa": "public", "soc2": "type2", "hipaaBaa": "available" },
    "incidents": [],
    "evidence": [
      {
        "field": "training",
        "url": "https://privacy.claude.com/en/articles/7996868",
        "quote": "By default, we will not use your inputs or outputs from our commercial products (e.g. Claude for Work, Anthropic API, Claude Gov, etc.) to train our models.",
        "retrieved": "2026-09-30"
      }
    ],
    "sourceUrl": "https://www.anthropic.com/legal/commercial-terms",
    "sourceDate": "2026-09-30"
  }
}

Errors:

GET /api/v1/meta

Dataset version, licence, counts by rating and category, the exact filter vocabularies, and an endpoint map. Use it to validate filter values client-side or to detect dataset updates.

{
  "version": "1.13.0",
  "lastUpdated": "2026-09-30",
  "count": 96,
  "license": "CC-BY-4.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/4.0/",
  "attribution": "PrivacyWatch by WyrdWerk (https://privacywatch.wyrdwerk.com)",
  "counts": { "rating": { "clean": 21, "guarded": 23, "caution": 38, "high-risk": 13, "unverified": 1 },
              "category": { "us-frontier": 8, "chinese": 10, "inference": 69, "coding": 9 } },
  "endpoints": { "list": "/api/v1/providers", "provider": "/api/v1/providers/{id}",
                 "meta": "/api/v1/meta", "openapi": "/api/v1/openapi.json",
                 "search": "/api/search?q=", "dataset": "/providers.json" }
}

Errors: 502 — {"error": "dataset unavailable"}

GET /api/v1/openapi.json

OpenAPI 3.1 spec for the API, with the dataset version injected as info.version. Machine-readable; import it into tooling or code generators.

GET /api/search?q=

Semantic search over the archived policy corpus (privacy policies, ToS, DPAs behind the dataset). Returns the top matching text chunks with source URLs and provider IDs. Not a provider filter — it searches documents, not the structured fields.

Parameters: q (required, free text, max 512 chars), k (optional, results 1–25, default 8).

Example response (?q=zero+data+retention, illustrative and trimmed):

{
  "query": "zero data retention",
  "count": 42,
  "results": [
    {
      "providerIds": ["anthropic-api"],
      "url": "https://platform.claude.com/docs/en/manage-claude/api-and-data-retention",
      "score": 0.822,
      "snippet": "ZDR is enabled per organization; each new organization requires ZDR to be enabled separately by your account team…"
    }
  ]
}

Errors:

GET /providers.json

The full raw dataset — meta plus all 96 complete provider rows. Same data the API reads, served as a static file. Use it for bulk processing or offline analysis; use the API for filtered/point queries. CORS-enabled, cached for an hour.

#Filters

All filters are optional and combine with AND. Within one filter, comma-separated values or repeated parameters mean OR. The vocabularies come from scripts/lib/v2-fields.mjs; /api/v1/meta returns the same lists at runtime.

ParameterMatchesAllowed values
categoryrow.categoryus-frontier, chinese, inference, coding
ratingrow.ratingclean, guarded, caution, high-risk, unverified
surfaceTyperow.surfaceTypeapi, consumer
trainingtraining.defaultoff, on, opt-in, tier-dependent, silent, conflicting
zdrzdr.accessdefault, self-serve, approval, enterprise, none, silent
retentionretention.kindnone, transient, fixed, until-deleted, indefinite, silent
regionvalue appears in location.regionsISO 3166-1 alpha-2 code (e.g. US, DE), EU, or GLOBAL
maxRetentionDaysretention.kind is fixed and retention.days ≤ Nnon-negative integer. Only fixed retention can match; silent, indefinite, etc. never do
incidentrow.incidenttrue, false
idsrow.id in the listcomma-separated provider IDs (use to compare rows side by side)
fieldsresponse shapesummary (default), full

Invalid values return 400 with the allowed list, so a client can self-correct from the error alone:

GET /api/v1/providers?rating=great
→ 400 {"error": "rating \"great\" is not valid; allowed: clean, guarded, caution, high-risk, unverified"}

#Field glossary

Summary row

FieldMeaning
id, name, surface, surfaceTypeStable ID, display name, product surface (“API”, “Consumer”…) and its type (api/consumer)
categoryus-frontier, chinese, inference, coding
ratingOverall assessment — see Ratings below
incidenttrue when a confirmed breach, regulatory action, or government ban is on record
trainingtraining.default: what a new account gets
trainingOptOuttraining.optOut: how to turn training off when it's on
retention / retentionDaysretention.kind; retentionDays is only set when kind is fixed, otherwise null
zdrzdr.access: whether a zero-data-retention option exists and what unlocks it
regionslocation.regions
sourceDateWhen this row was last verified against the primary source
urlPath to the full row for this provider

Structured fields (full rows)

FieldValuesMeaning
training.defaultoff · on · opt-in · tier-dependent · silent · conflictingTraining use for a new account by default. tier-dependent = free/paid/enterprise differ (details in training.detail); conflicting = the provider's own documents disagree
training.optOutsetting · api-param · email · contract · not-needed · none · silentHow to opt out. not-needed pairs with training.default: off; none = training is on and there is no opt-out
retention.kindnone · transient · fixed · until-deleted · indefinite · silentHow long inputs/outputs are kept. transient = only to serve the request
retention.dayspositive integerPresent only when retention.kind is fixed
zdr.accessdefault · self-serve · approval · enterprise · none · silentZero data retention option: default = on for everyone, self-serve = toggle in settings/console, approval = requires provider sign-off, enterprise = contract-gated, none = no ZDR
location.regionsISO 3166-1 alpha-2 codes, EU, GLOBAL; [] = silentWhere data can be processed/stored. GLOBAL = provider states global processing. Empty array = read the docs, they don't say
compliance.dpapublic · on-request · enterprise · none · silentData processing addendum availability
compliance.soc2type1 · type2 · none · silentSOC 2 attestation level
compliance.hipaaBaaavailable · enterprise · none · silentHIPAA business associate agreement availability
incidents[]{ date, type, confirmed, summary, sourceUrl }type: breach, regulatory, ban, allegation. date is YYYY-MM or YYYY-MM-DD. Only confirmed non-allegation entries set the row's incident flag
evidence[]{ field, url, quote, retrieved }The backing for a structured field: verbatim quote from the source URL, with the date we retrieved it. field is one of training, retention, zdr, location, compliance, incidents. Quotes are omitted only when the value is silent

“silent” vs missing

silent — we read the provider's documents and they don't address this field. A finding of opacity, not an absence of research, and not evidence the provider trains or retains anything.
Missing — not researched yet. Never conflate the two.

Ratings

🟢 Clean 🟡 Guarded 🟠 Caution 🔴 High Risk ⚫ Unverified

Clean: training off by default, ZDR documented, clear policy. Guarded: generally safe with documented caveats. Caution: training on by default, vague retention, or docs that stay opaque after a live check. High Risk: training on with weak opt-out, China storage, confirmed breach or government bans. Unverified: the cited document never addresses API data at all. The 🚩 incident flag is additive on top of the base rating. Full definitions with caveats: README rating system.

#Versioning policy

meta.version (also in every response's meta) is the dataset version; /api/v1/meta's lastUpdated is the research date.

#Licence & attribution

Not legal advice. This is a good-faith summary of public policy documents. Policies change — verify with primary sources (each row's evidence[] links them) before making compliance decisions.