Skip to content

New paid orders and subscriptions are retired. Existing records, settlement and refunds remain available through your account.

Hospitality in codes — for web scrapers

We’d rather give you the JSON.

HTML scraping is a bad contract for both of us: layout can change without notice, your parser breaks, our compute pays ~10× the cost of serving JSON. The JSON API at /api/v1/* is machine-readable, but each route has its own stability and rights boundary. Public reach is not a blanket CC0 grant. Start with its response metadata.

If you’re here to scrape, try these instead

Structural card lookup (prices withheld)

/api/v1/universal/card/[sku]

Card list (per set)

/api/v1/universal/set/[code]

Bulk publication status (0 card rows)

/data/catalog.jsonl

Date-shaped structural compatibility view

/api/at/[YYYY-MM-DD]/card/[sku]

Sets per game

/api/v1/universal/sets/[game]

Every game

/api/v1/universal/games

All publicly reachable without auth. Reuse rights vary by response: Cambridge-authored structure may be CC0; mixed catalog data is NOASSERTION. Preserve _meta.license and _meta.source_license; absence of a license is not permission. See /api/v1/welcome for the full menu, or the mirror-the-catalog guide for the polite refresh discipline.

If you must scrape HTML

Some legitimate use cases require scraping the rendered surface (e.g. archive crawlers, web-of-trust verifiers, accessibility audits). Here are the substrate primitives we publish to make that easier.

/robots.txt

What’s allowed; Crawl-delay: 2; sitemap pointer; per-bot opt-outs for training-only crawlers.

/sitemap.xml

Structured listing of every crawlable URL with lastModified + changeFrequency + priority.

/.well-known/cambridge-tcg.json

Machine-readable manifest of every public surface with status, auth, methodology links.

/.well-known/ai-plugin.json

OpenAI-style plugin discovery; LLM platforms reading this register us as a tool.

/.well-known/mcp.json

MCP (Model Context Protocol) discovery with suggested read-tools per endpoint.

Crawl etiquette

  • • User-Agent: send <project>/<version> (<contact-email>). A contact makes communication possible; it does not promise outreach before an infrastructure or security control acts.
  • • Honour Crawl-delay: 2 from robots.txt. Per-resource cadence at /api/v1/rate-limits.
  • • Cache when a response permits it: respect any Cache-Control and freshness metadata that is present.
  • • Honour HTTP 429: response body may include a Retry-After header or retry detail. Back off on repeated responses even when no machine-readable delay is present.
  • • Don’t bulk re-export data tagged internal-only in _meta.source_license. License boundary. Legacy CardRush-derived values are currently withheld rather than exposed behind an authentication gate.

Structured-data markup on HTML pages

Cambridge TCG’s HTML pages emit schema.org markup where applicable (Product, Offer, BreadcrumbList, DefinedTermSet for the glossary). If you’re a structured-data crawler, parse the application/ld+json blocks instead of CSS selectors.

schema.org coverage is ongoing; gaps are tracked in the substrate-honesty audit (substrate-honesty-audit.md).