API v0.2 · comparison boundaries corrected October 7, 2026

Look up evidence in a defined scope.

The API returns a best recorded result only when the benchmark, metric, version and evaluation protocol can be matched. It does not compare FPS, accuracy, Elo or scores from unrelated benchmarks. CodeSOTA supplies evidence; inference happens at your provider.

Open evidence index

Quickstart

Available tasks, result dates and missing-date counts
curl https://www.codesota.com/api/sota
Source-scoped OCR evidence record
curl "https://www.codesota.com/api/sota/ocr?tier=sota"
Explicit matched scope, same OCR record
curl "https://www.codesota.com/api/sota/document-ocr?benchmark=omnidocbench&metric=composite&benchmark_version=1.5&protocol=glm-ocr-publisher-report&evaluation_scope=end-to-end-document-parsing"

Short aliases remain accepted: ocr, code, asr, stt, tts, vqa, caption, t2i and t2v. A registered task does not imply that a defensible ranked scope exists.

Illustrative response fields

This shortened example shows the shape. The source reports GLM-OCR’s 94.62 score on OmniDocBench v1.5, but the database records no result date. The response keeps that date unknown. Publisher source.

OCR record; snapshot hash and response time omitted
{
  "task": "ocr",
  "benchmark": "omnidocbench",
  "benchmark_version": "1.5",
  "protocol": "glm-ocr-publisher-report",
  "evaluation_scope": "end-to-end-document-parsing",
  "as_of": null,
  "pick": {
    "model_id": "glm-ocr",
    "score": 94.62,
    "metric_id": "composite",
    "source_url": "https://github.com/zai-org/GLM-OCR",
    "result_date": null
  },
  "runners_up": [],
  "coverage": {
    "candidate_count": 1,
    "latest_result_date": null,
    "undated_result_count": 1,
    "release_coverage": "limited historical source-scoped records; not exhaustive current model coverage"
  }
}
FieldMeaning
benchmarkCanonical dataset ID. OCR defaults to OmniDocBench document parsing.
benchmark_versionSource-backed benchmark version; user parameters cannot assign missing metadata.
protocol / evaluation_scopePublisher report or harness identity, plus evaluated task scope.
pick / runners_upOnly rows sharing benchmark, metric, version, protocol and scope. One candidate is a record, not a global SOTA ranking.
pick.source_urlDirect evidence source for the score.
pick.result_date / as_ofActual recorded result date, or null when unknown. Neither is inferred from access time.
coverageCandidate count, known date range, undated count and release-coverage limitation.
snapshot_idStable hash of scope and compared evidence. Undated records use reg-undated rather than today’s date.
retrieved_atTime the response was generated. This is not evidence freshness.
provider_hints / cost_per_1k_usd / cost_basisNull when no defensible unified provider or pricing evidence is available.

TTS constraints and unavailable evidence

Polish local commercial constraint request — no established measured winner
curl "https://www.codesota.com/api/sota/tts?language=pl&deployment=local&commercial_use=true&max_params=500M&metric=hardtext_entity_accuracy"

The former example incorrectly returned Kokoro v1.0 as a measured Polish winner. Its official v1.0 voice list does not establish Polish support, and the published hard-text records used English prompts. Their placeholder manifests have also been withdrawn from measured recommendations. There is no inspectable measured Polish comparison here. Official Kokoro voices.

Evidence-unavailable response (HTTP 409)
{
  "error": "No comparable measured TTS evidence for those constraints.",
  "hint": "No Polish winner is currently established. Browse candidate metadata and verify exact checkpoint and voice support.",
  "see": "https://www.codesota.com/text-to-speech/registry"
}

TTS supports language, deployment, commercial use, exact license, streaming, cloning, voice design, parameter and RAM limits, minimum sample rate and code-switching filters. Latency and cost limits require known measurements; missing values cannot pass. Supported measured metrics are hardtext_entity_accuracy, wer, cer, ttfb_p95_ms, cost_per_1m_chars_usd and utmos_mean. Reported MOS from separate listening studies is not a fallback ranking.

Errors and transport

HTTPMeaning
400Invalid filter value or unsupported measured metric.
404Unknown task or no registered matching candidates.
409No verified comparable scope or no measured evidence for the requested constraints. Follow the hint; changing parameters cannot manufacture evidence.
501Unsupported tier. Only sota is implemented.
503 / 500Registry unavailable / query failure.

GET and OPTIONS use open CORS. Responses declare public, max-age=300, s-maxage=300. Cache and response refresh times do not change evidence dates. Paths and aliases remain stable; the 409 comparison guard is a deliberate correction to unsafe earlier rankings.

Evidence methodology