[gold]rails: how a benchmark run flows

From pinned source rows to a published result. Arrows name the artifact that moves. Coral steps are committed evidence gates.

01 / Data 02 / Declare & freeze 03 / Run 04 / Evaluate 05 / Check & publish Public sources pinned revisions 20 datasets Release v1.3 8,807 rows sha 818c9e00 Subset tune / test split grouped rows Declaration systems + questions git commit Tuning run tune rows only per arm Freeze manifest thresholds git commit Ledgers JSONL per call row hash + times Systems 7 guardrails 3 backends Harness goldrails_bench retry policy Contract v1.1 rules + weights signed Evaluator leaderboard.py approved sha Leaderboard final JSON scores + CIs Hugging Face dataset v0.0.1 public GitHub code + records public Release checks 27 page checks gate Results page results.json local pinned rows licence-checked sampled rows hash-pinned question sets committed first tune rows held-in thresholds tuning only frozen config git commit test rows never tuned on same prompt every system raw answers timestamped call records row hashes scoring rules signed scores + CIs paired bootstrap leaderboard frozen inputs results.json local export scanned full-text package rows re-verified Legend main path evidence gate publish step stored artifact

Systems under test

  • • Jev 1.13.0 through the TypeSafe API
  • • Bedrock through ApplyGuardrail and InvokeGuardrailChecks
  • • Kev 0.8B, 4B, 9B, Open-Jev 2B and Laya on one GCP g2-standard-24 (2x L4)

Evidence gates

  • • Declaration and freeze manifest are committed before the first test call
  • • Every ledger record carries its row hash and attempt timestamps
  • • The evaluator refuses scoring code or contracts that no committed approval names

What the evaluator computes

  • • Task score = 100 x (recall + benign pass rate) / 2
  • • 2,000 paired group-bootstrap draws for every interval
  • • Cost from tokens or GPU time x dated tariffs; serial latency p50 and p95