About
Who runs this
LedgerRate is run by John Carlson. Questions, corrections, and challenges to the methodology all go to jjccarlson@outlook.com — a named person answers, not a support queue.
How this is funded
LedgerRate is self-funded by its operator. If and when it takes revenue, it will come from readers and users of the data — subscriptions and results-data licensing — never from the companies being measured. Concretely, that means:
- No paid placement. No model vendor can pay to be included, excluded, re-run, or presented differently (independence policy).
- No vendor money. No sponsorships, grants, or advertising from AI model providers.
- Numbers are never revised silently. Corrections are dated and published on the errata page; the original stays visible.
- Methodology changes are versioned and announced, never applied retroactively. Trendlines break rather than pretend.
If you're a model vendor and dispute a result
Email with the edition, workload, and cell you dispute, and the specific claim. The raw logs for every published number are already public — start there. If the dispute identifies a harness defect or an infrastructure failure misclassified as a model failure, the cell is re-run, and both the original and corrected results are published with an errata note. If the dispute is "the model can do better with different prompting," the answer is no: prompts are identical across models by design, published verbatim per workload. Vendor responses are published unedited alongside the edition they concern.
Red-team the methodology
The fastest way to make this benchmark better is to break it. If you can show a way the current design rewards the wrong thing — grading loopholes, dataset artifacts, statistical sleight of hand — write in. Substantive critiques get published with credit, including the ones we can't fix yet.
Verify the numbers yourself
Every edition ships with its raw graded logs, frozen pricing, and dataset hashes. The harness includes a one-line verification command that recomputes every published aggregate from those logs and diffs it against the published data — see any edition page for the exact command. Don't trust us; run it.