HumanEval.org

Data downloads

Nightly anonymized dumps of every ratings-eligible battle and vote — the exact inputs the official leaderboards are computed from. Stable URLs: /downloads/files/latest/* always points at the newest dump. Prefer machine access? See the public JSON API.

Available dumps

2026-10-10latest

generated 2026-10-10T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B284637ec09002dd090004278eec12a7bdad43c9e2664d7684dce3654f9bd9850
votes.jsonl079 Bd116c1f53216f72a87347ace1893993049f7d002cb356270810f68f6293d7fe9

Manifest: manifest.json

2026-10-09

generated 2026-10-09T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Ba7e47b53d3cc9b70a8b975cc1c0b2a4abf1abc16500b170818fbd10496cf852d
votes.jsonl079 B8d2e7065341feee2c17ba47c506ae12b549908e4d3d07d894be89adbe54f4c04

Manifest: manifest.json

2026-10-08

generated 2026-10-08T03:15:01Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bf76dd610d32fa0a38b47e5e71371b1778e8d0b9a0a1255a89ea81441a22d0263
votes.jsonl079 B4d4bf3b072543dd9f3cf6452fb44adb8e25f2f16cda93d0f563b6f30f449bf71

Manifest: manifest.json

2026-10-07

generated 2026-10-07T03:15:04Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bd595ec445dfd07c07b928b3a7efcb70b4a54dcfa33bf85257d1b907da4931159
votes.jsonl079 B111629a042a677f430144c4717601688d0748fec8eddbb95edc94cbc5d2a4ebc

Manifest: manifest.json

2026-10-06

generated 2026-10-06T03:15:10Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B2802f6301706fb506123a2c0184a1c96eaf013cda327e29ca6e1a8c6337c6c99
votes.jsonl079 B8c444f14466f073db34fcdc680a7a17fdf789bf7e1926db88220d650a4e8c4ab

Manifest: manifest.json

2026-10-05

generated 2026-10-05T03:15:07Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bb67441a52115e6219479dc1e1b662f319abb6a3a514ef8df8d31387b925ac974
votes.jsonl079 Ba5000523a846d454777edb3eba7fd71ce40b3c50864955cb229bbb347ab5e0dc

Manifest: manifest.json

2026-10-04

generated 2026-10-04T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B935939af7f309820c59de565535c30572d450c6647172244de59e524ac9e5309
votes.jsonl079 B2d32f24862431f1f4f0af94ad7634d17858f7ad1b8bd64e430e0455cc32f73d8

Manifest: manifest.json

2026-10-03

generated 2026-10-03T03:15:02Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Ba1410a5f328431eedfe420d33463a08282aaa93d3dfef6e6472da8419d0e8373
votes.jsonl079 B181467fa228848164628e2cc810b3f687db8603e2036a87489a87f0a12004cca

Manifest: manifest.json

2026-10-02

generated 2026-10-02T03:15:00Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bc98d400356db0dcb61d87c49fc239f80f8f66c0179e7c12b5aa3e25d36c7836c
votes.jsonl079 B5ad09225e28e8c6b944d96be99dd73861d6f3f87b8a6633d5d70ee03dc9c1278

Manifest: manifest.json

2026-10-01

generated 2026-10-01T03:15:01Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B0c60c9397fbd711da24913c5018b36366374bbd7b025f5cb77e8fe3f2d9fb72f
votes.jsonl079 Bd630f59e7b57026dbb8793ba7ff4e6c61b538c8b995c7f7463c8c172c1af92a8

Manifest: manifest.json

2026-09-30

generated 2026-09-30T03:15:02Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B2905371d23e01702e89c0829360d5ca6a9143e98882ea0426d2e0a885f55af94
votes.jsonl079 B35b1bbcb0c16c9960aec8ab810aa0655973b8111e888e55de0e8cecd6587fdd6

Manifest: manifest.json

2026-09-29

generated 2026-09-29T03:15:04Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bc198b0f895c68549a89dc6155c7cbf51738644c7912995e311dae5d5c263e723
votes.jsonl079 B849ec673d3006bd50c8427ae1bc023cc26a413d6f73a21128b0e3d45a5f5a05a

Manifest: manifest.json

2026-09-28

generated 2026-09-28T03:15:03Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B5d2307a2d138e69901c3cadb6706eae05d9b8c614611f2c80e14e5b864211119
votes.jsonl079 B0ac828ef525e5903e1e905d4e38496a8e8cbe693270fc5b684420ba067bc07ac

Manifest: manifest.json

2026-09-27

generated 2026-09-27T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bd663a95ac45382be139d78374b5c1296b465d574802289263ca8039e78ec11b9
votes.jsonl079 B08254f116124cbdb4ca499b6b2f365128ffce2452124cfcafc544c437a83b8c6

Manifest: manifest.json

2026-09-26

generated 2026-09-26T03:15:07Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B4d66b3324db2b8cf3c492b693ac7bfaa66a2d0d41451e4fc2adcab5d1e12a475
votes.jsonl079 Bebb9bc10560f6800a9e203391b884cbe8b8d8932618fabe11c1d7c60b23fe009

Manifest: manifest.json

2026-09-25

generated 2026-09-25T03:15:04Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B85a648e44ccd319fb9d1fa9bddbf7387a5e7cdcfb788d4d19d4c123b99fa1f9c
votes.jsonl079 Bc617d3b4c87aaf5e83c98ede57e4813d06c89138fda357726a562ca3ba5fc797

Manifest: manifest.json

2026-09-24

generated 2026-09-24T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B183ecfa9914a287cc60984bc1a1150ba0f1a44e77168407225853994b12d75a3
votes.jsonl079 Be88c8358dcc718a79922477cbcacd262b779c3ff642a6ed7ba8f431ae5197685

Manifest: manifest.json

2026-09-23

generated 2026-09-23T03:15:10Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B517f0e4bbe03846ae3a197f68629e6ff04f3fcdeb37b98bffb80657ee9b39622
votes.jsonl079 Bdf7861f03b4be56abb605693cc96b78425042fe78ce6009f0ece0ab26db5ce49

Manifest: manifest.json

2026-09-22

generated 2026-09-22T03:15:09Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B81f58878905df95687f27e7766369af6eabde87bb7f95d345266fe831b21478b
votes.jsonl079 Bf2649b1a68e76adab22156c6ab6ac2715ccefc0b4b4da37cfe87aa0317f8d67a

Manifest: manifest.json

2026-09-21

generated 2026-09-21T03:15:21Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Be90e062c9e695f4da373d0f080423e0cb99a5116ab1f3de72a8431cbe10597a3
votes.jsonl079 B40c42b82c3c7775c73080be10423f229104b258fff78de465e5c00311394a9a8

Manifest: manifest.json

2026-09-20

generated 2026-09-20T03:15:01Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B2b5eadc948bf8021cad1a1badc22ac630d2d062c0c9c8eae43403e95c7206091
votes.jsonl079 Bac5b163c0a24373149efaa83bc965b629f34acdae2d80ad50ea2be3d09bb1418

Manifest: manifest.json

2026-09-19

generated 2026-09-19T03:15:00Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Ba9ef0a898827a9816f73cf762616609b9ee5f40efb62f2c5bd0d2d4586ab2622
votes.jsonl079 B796defcc3730cfef971b70d3d34bfd4966dd60b585dabfb540b0d068a0f1a96b

Manifest: manifest.json

2026-09-18

generated 2026-09-18T03:15:02Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B7173defc5983c48ca2a224d49edbf9aaa0f3367e88857249948780e3fa04da7b
votes.jsonl079 Bacb76a724c14ed6021ff3dac21098592990394f34299f37a2e1a9aa86545fbf3

Manifest: manifest.json

2026-09-17

generated 2026-09-17T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Be91f817f800191b33fb4b50c69ed7df466147684e5625d69b9a4de3004197c73
votes.jsonl079 Bac15d831ff6e8c24cb01a14de289fd4b853887ac8cee7e79c5a64d79b5de0e5f

Manifest: manifest.json

2026-09-16

generated 2026-09-16T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bdd15d5572c14fb858cf85f741256999c8c483ace66190cf5a2b1c323c037a705
votes.jsonl079 Bec3cc6e0ab9984db3079d7072ed106012263648c6d3c311ce0ae33583e4dbdf6

Manifest: manifest.json

2026-09-15

generated 2026-09-15T03:15:02Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B8ec98ccfce807166bff264710f44a600cbe1002a38ff24c82c76233655d7fcf2
votes.jsonl079 B23dd643a5be360a4bba49a23b6e5c0fe37630e59de03d044dd62bdd4881fc3a5

Manifest: manifest.json

2026-09-14

generated 2026-09-14T03:15:06Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 B474769b88044c0d4d735189418524a9c4edc0d8ab6eafb0a26aa39a3788accff
votes.jsonl079 B45dd86c45171ddd58b81f258b9f62ec9fe0b57687615c119669fcfb4553f544b

Manifest: manifest.json

2026-09-13

generated 2026-09-13T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Be316405de66fc78d5f43e9dbe61811da0aaf3dcd551f69591d8ff8621050c76b
votes.jsonl079 B9cf694d2fc599fc330bd30b8d468faf7e2bb39a0080fea6c3508276cdb259db3

Manifest: manifest.json

2026-09-12

generated 2026-09-12T03:15:16Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bd9289a2ac9deff301f678920cc75d7552f5429ec052742a03ffd3cabd9e6a741
votes.jsonl079 Ba08a4f56ef627e0c37ff5a54ca4d08b53343aaac6819a5d0573f0abc0ad134e1

Manifest: manifest.json

2026-09-11

generated 2026-09-11T03:15:29Z · categories: (empty)

FileRecordsSizeSHA-256
battles.jsonl099 Bc0fb4437b8145fe41431800ef69b1fdde49e02d77cf9fd51cc823c61b30b2e23
votes.jsonl079 B4200dd00b1e2792136078bc9590e76f28b456d86a5c507b0b08d08b5fc9e6e0f

Manifest: manifest.json

Reproduce the official leaderboard

The official numbers are produced by running the open humaneval-ratings package on these dumps with the published seed — no database, no network, no hidden inputs. Byte-identical reproduction is guaranteed under the package's pinned environment (requirements-lock.txt):

# 1. Download the dump you want to verify
curl -O https://humaneval.org/downloads/files/latest/battles.jsonl
curl -O https://humaneval.org/downloads/files/latest/votes.jsonl
curl -O https://humaneval.org/downloads/files/latest/manifest.json

# 2. Check the digests against manifest.json
sha256sum battles.jsonl votes.jsonl

# 3. Install the rating engine with its exact pins
python3 -m venv .venv
.venv/bin/pip install -r requirements-lock.txt   # from the package source
.venv/bin/pip install humaneval-ratings           # or: pip install <source dir>

# 4. Compute — the official run uses seed 42, 100 bootstrap rounds
.venv/bin/humaneval-ratings compute \
    --battles battles.jsonl --votes votes.jsonl \
    --out leaderboard.json --seed 42

# 5. Compare with the published snapshot / API values
sha256sum leaderboard.json

Every leaderboard.json records the seed, bootstrap rounds, package version and the SHA-256 of its exact input files in its metadata, so third parties can verify each other's runs as well. Vote weights in the dump come from a hidden quality pipeline (see methodology) — their values are public, their derivation is not.

Dump format specification

The text below is the canonical DUMP_FORMAT.md shipped with the rating engine (single source of truth, rendered as-is).

HumanEval.org public data dump format — v2

This document is the canonical, normative specification of the HumanEval.org public data dump and of the leaderboard.json file the humaneval-ratings package produces from it. The platform's nightly export is built to match this spec; the official leaderboard is produced by running this package on the published dump, so anyone can reproduce the official numbers byte-for-byte (see README.md).

Version note. Dumps published from 2026-09-08 on carry schema_version: 2 (section 3.1 adds recording battles for the computer-use category). v2 is a superset of v1: every v1 record is a valid v2 text record. The engine (humaneval-ratings ≥ 1.1.0) reads both versions; v1-only readers must be updated before consuming new dumps (the header version is exactly what they check).

The rating mathematics applied to a dump is specified in METHODOLOGY.md.

1. General conventions

  • A dump consists of exactly two files: battles.jsonl and votes.jsonl.
  • Encoding: UTF-8, Unix line endings (LF), one JSON object per line, no blank lines, file ends with a trailing LF.
  • Header line: the first line of each file is a metadata object (section 2). All subsequent lines are data records.
  • IDs are JSON strings (battle_id, vote_id, voter_id). The platform's internal keys are 64-bit integers which can exceed JavaScript's 2^53 safe-integer range; strings keep IDs safe and opaque. IDs are unique within their file and stable across dumps.
  • Models and categories are identified by their public slugs (models.slug, categories.slug on the platform side), never by internal numeric ids.
  • Timestamps are ISO 8601 UTC with a trailing Z, second precision: 2026-08-22T19:31:04Z.
  • Forward compatibility: readers MUST ignore unknown object fields. Additive changes (new optional fields) do not bump schema_version; breaking changes do.
  • Privacy: battles and responses are public content. No user data is ever exported; voter_id is a pseudonymous opaque token (see 4.1).

2. Header lines

First line of battles.jsonl:

{"kind": "battles", "schema_version": 2, "generated_at": "2026-09-08T04:00:00Z", "categories": ["browser-use", "chat-writing", "chat-coding"]}

First line of votes.jsonl:

{"kind": "votes", "schema_version": 2, "generated_at": "2026-09-08T04:00:00Z"}
FieldTypeRequiredMeaning
kindstringyes"battles" or "votes"; must match the file's content
schema_versionintegeryesThis spec is version 2 (v1 dumps remain readable). Readers reject versions they don't support
generated_attimestampyesWhen the export was produced
categoriesarray of stringsbattles file onlyThe full category-slug set covered by this dump; every battle's category_slug MUST be in it

Headers may carry extra fields (e.g. a source URL); readers ignore them.

3. battles.jsonl — data records

One record per voted battle. Battles without eligible votes MAY be included (the engine ignores them) but the export SHOULD omit them.

{"battle_id": "1024", "category_slug": "chat-writing", "model_a": "gpt-alpha", "model_b": "claude-beta", "created_at": "2026-08-22T19:31:04Z", "normalization_mode": "rendered", "modality": "text", "response_a": "# Draft\n\nDear team, ...", "response_b": "Dear team, ..."}
FieldTypeRequiredMeaning
battle_idstringyesUnique battle identifier
category_slugstringyesCategory the battle was fought in; must appear in the header categories list
model_astringyesSlug of the model in slot A (as shown to the voter, post position-randomization)
model_bstringyesSlug of the model in slot B; must differ from model_a
created_attimestampyesBattle creation time
normalization_modestringyesOne of rendered, raw, noformat — how responses were displayed to voters (raw for recordings; not meaningful there)
modalitystringv2: optional, default texttext or recording (section 3.1). Absent in v1
response_astringtext battles: yesFull response TEXT produced by the slot-A model. MUST be absent for recording battles
response_bstringtext battles: yesFull response TEXT produced by the slot-B model. MUST be absent for recording battles

Notes for the platform exporter:

  • Response bodies live in object storage as text artifacts (artifacts.storage_ref); the exporter MUST resolve and inline the raw text (the stored artifact content, not the rendered HTML) into response_a/response_b. Style covariates are computed from these texts (METHODOLOGY.md §7), so their exact bytes matter.
  • v1 exported text-modality battles only. v2 adds the recording modality below; any further artifact modality is another bump.

3.1 recording battles (v2 — computer-use category)

Computer-use ("Browser Use") battles have no response text: each contestant drove a real browser through the same task on live websites, concurrently from an identical starting scene, and voters judged wall-clock-faithful recordings side by side (METHODOLOGY.md §10). A recording battle is exported as:

{"battle_id": "2048", "category_slug": "browser-use", "model_a": "gpt-alpha", "model_b": "claude-beta", "created_at": "2026-09-08T10:00:00Z", "normalization_mode": "raw", "modality": "recording", "recording_a": {"video_key": "battles/2048/a.mp4", "trace_key": "battles/2048/a.trace.jsonl", "screenshot_key": "battles/2048/a.png", "assisted": false}, "recording_b": {"video_key": "battles/2048/b.mp4", "trace_key": "battles/2048/b.trace.jsonl", "screenshot_key": null, "assisted": false}, "telemetry_a": {"duration_ms": 184000, "time_to_first_action_ms": 6200, "steps": 14, "median_step_ms": 9100, "model_ms_total": 121000, "harness_ms_total": 4900}, "telemetry_b": {"duration_ms": 152000, "time_to_first_action_ms": 4100, "steps": 11, "median_step_ms": 8700, "model_ms_total": 88000, "harness_ms_total": 3900}}
FieldTypeRequiredMeaning
recording_a, recording_bobjectyesPer-slot artifact references (below)
recording_*.video_keystringyesObject key of the constant-frame-rate video whose duration equals the run's wall clock (t=0 = the shared start tick)
recording_*.trace_keystringyesObject key of the JSONL action/event trace (cursor moves, clicks, typing with secrets masked, per-step model/harness latency) — the source of truth for what the agent did
recording_*.screenshot_keystring or nullyesObject key of the final screenshot, if captured
recording_*.assistedbooleanyestrue if a human intervened in THIS slot's run. Assisted battles are never exported (they are ratings-ineligible); the field exists so the record is self-describing
telemetry_a, telemetry_bobjectyesRun telemetry; every member is an integer or null
telemetry_*.duration_msinteger/nullyesWall-clock run length in ms
telemetry_*.time_to_first_action_msinteger/nullyesms from the shared start tick to the contestant's first dispatched action
telemetry_*.stepsinteger/nullyesNumber of observation→action steps
telemetry_*.median_step_msinteger/nullyesMedian per-step latency (model + harness)
telemetry_*.model_ms_totalinteger/nullyesTotal time spent waiting on the model
telemetry_*.harness_ms_totalinteger/nullyesTotal harness overhead (identical scaffold for every model; published so equal treatment is checkable)

Object keys are relative to the platform's artifacts bucket; the media themselves are large binary objects and are not part of the dump (they are served through the platform, not published in bulk). The rating engine does not read them: it uses model_a/model_b/ category_slug and the votes exactly as for text battles, and treats the style covariates of a recording battle as zero (METHODOLOGY.md §7.5).

4. votes.jsonl — data records

One record per ratings-eligible vote.

{"vote_id": "5001", "battle_id": "1024", "voter_id": "v_9f2c7a1e", "choice": "win_a", "created_at": "2026-08-22T19:32:40Z", "weight": 1.0}
FieldTypeRequiredMeaning
vote_idstringyesUnique vote identifier
battle_idstringyesThe battle voted on; MUST exist in battles.jsonl
voter_idstringyesPseudonymous voter token (see 4.1)
choicestringyeswin_a, win_b, or tie
created_attimestampyesVote time
weightnumberyesEffective vote weight, finite and > 0; 1.0 is the default full weight

4.1 Voter pseudonyms

voter_id is an opaque token derived by the platform (e.g. a keyed hash of the internal user id). It is stable within a dump (and across dumps, so long-term voter behavior is analyzable) but cannot be linked back to an account. The derivation is internal and never published.

4.2 Vote weights and the quality pipeline

The platform runs a hidden vote-quality pipeline (consensus agreement, gold-standard checks, behavioral signals, provisional-account windows). That pipeline runs before export:

  • Votes it excludes never appear in the dump.
  • Votes it down-weights appear with their reduced effective weight.
  • Full-quality votes carry weight: 1.0.

The dump therefore contains exactly the votes that count, each with the weight it counts at. Ratings computed from the dump MUST use these weights (METHODOLOGY.md §3). How weights are derived is deliberately unpublished (anti-gaming); that they are applied, and their values, are fully public here. The exporter also guarantees at most one vote per (battle, voter) pair; the engine does not re-check this.

5. Validation rules (normative for the engine)

The engine rejects a dump (non-zero exit, no output) when:

  1. A header line is missing, has the wrong kind, or an unsupported schema_version.
  2. Any required field is missing or has the wrong JSON type (for v2 recording battles: recording_a/recording_b missing or not objects, or a response_* text present).
  3. choice, normalization_mode or (v2) modality has an unknown value.
  4. weight is not a finite number > 0.
  5. A battle_id (in battles) or vote_id (in votes) is duplicated.
  6. A vote references a battle_id not present in battles.jsonl.
  7. model_a == model_b in any battle.
  8. A battle's category_slug is absent from the header categories list.

Unknown fields are ignored everywhere. Battles with zero votes are ignored. Categories with zero votes produce no leaderboard section. Models enter a category's leaderboard only via voted battles there.

6. Output: leaderboard.json

Produced by humaneval-ratings compute. Top-level shape:

{
  "metadata": {
    "package": "humaneval-ratings",
    "package_version": "1.0.0",
    "leaderboard_schema_version": 1,
    "dump_schema_version": 1,
    "dump_generated_at": "2026-08-23T04:00:00Z",
    "seed": 42,
    "bootstrap_rounds": 100,
    "min_votes": 30,
    "style_min_votes": 50,
    "input_digests": {
      "battles_sha256": "9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08",
      "votes_sha256": "60303ae22b998861bce3b28f33eec1be758a213c86c93c076dbe9f558c11c752"
    }
  },
  "categories": [
    {
      "category": "chat-writing",
      "vote_count": 8231,
      "style_control": true,
      "style_coefficients": {
        "length_chars": 24.1830,
        "markdown_density": 6.0021,
        "list_count": 3.1187,
        "header_count": -0.4402
      },
      "entries": [
        {
          "model": "claude-beta",
          "rating": 1041.2211,
          "ci_low": 1027.0110,
          "ci_high": 1055.9024,
          "rating_style_controlled": 1030.5470,
          "ci_low_sc": 1016.2001,
          "ci_high_sc": 1046.0193,
          "vote_count": 4110,
          "provisional": false,
          "rank": 1,
          "order": 1
        }
      ]
    }
  ]
}

6.1 Metadata

Every knob that affects the numbers is recorded: the seed, bootstrap round count, thresholds, package version, and the SHA-256 digests of the exact input files. There are no timestamps generated at compute time — dump_generated_at is echoed from the battles-file header (dump_schema_version likewise) — so the output is a pure function of (inputs, CLI parameters, pinned environment). Verifiers check both digests, run the same command, and compare sha256(leaderboard.json).

6.2 Category objects

Sorted by category slug (ascending, bytewise). Fields:

FieldMeaning
categoryCategory slug
vote_countTotal eligible votes in the category (unweighted count)
style_controlfalse when the style-control fallback triggered (METHODOLOGY.md §7.4); then SC fields mirror the raw fields
style_coefficientsFitted shared style coefficients in rating points per +1 SD of each normalized feature difference; null when style_control is false
entriesModel rows, sorted by order

6.3 Entry fields

FieldMeaning
modelModel slug
ratingRaw Bradley-Terry rating (center 1000, Elo-equivalent scale; METHODOLOGY.md §4)
ci_low, ci_high95% bootstrap CI of rating (2.5th/97.5th percentiles)
rating_style_controlledRating from the style-controlled fit (METHODOLOGY.md §7)
ci_low_sc, ci_high_sc95% bootstrap CI of the style-controlled rating
vote_countUnweighted count of eligible votes on battles involving this model in this category
provisionaltrue when vote_count < min_votes; provisional models keep their estimates but get rank: null
rankDisplayed rank band (integer, non-provisional models only, else null); overlapping CIs share a band (METHODOLOGY.md §6)
order1-based sort position within the category (all models, including provisional)

Sorting and rank bands are computed from the rounded, published values (see 6.4), so the ordering and bands are verifiable from leaderboard.json alone: entries sort by rating descending, ties broken by model slug ascending; order is that position. Rank bands use the published ci_low/ci_high (METHODOLOGY.md §6).

6.4 Determinism and precision policy (normative)

  • Every floating-point value in leaderboard.json is rounded to 4 decimal places (round-half-even, Python round(x, 4); matches the platform's numeric(10,4) snapshot columns) and serialized with exactly four decimal digits (-?d+.dddd).
  • Object keys are sorted (bytewise ascending); output is ASCII-only (non-ASCII escaped as in JSON \uXXXX); separators are ", " / ": " with 2-space indent; LF newlines; single trailing LF.
  • The RNG and every iteration order in the engine are deterministic (METHODOLOGY.md §8). Same input files + same CLI parameters + the pinned environment (requirements-lock.txt) → byte-identical output.

7. Worked example

A minimal, valid dump (v1 headers — still accepted; a v2 dump would say "schema_version": 2 and may add "modality": "text"). battles.jsonl:

{"kind": "battles", "schema_version": 1, "generated_at": "2026-08-23T04:00:00Z", "categories": ["chat-writing"]}
{"battle_id": "1", "category_slug": "chat-writing", "model_a": "gpt-alpha", "model_b": "claude-beta", "created_at": "2026-08-22T10:00:00Z", "normalization_mode": "rendered", "response_a": "## Plan\n\n- step one\n- step two", "response_b": "Do step one, then step two."}
{"battle_id": "2", "category_slug": "chat-writing", "model_a": "claude-beta", "model_b": "gpt-alpha", "created_at": "2026-08-22T11:00:00Z", "normalization_mode": "raw", "response_a": "Short answer.", "response_b": "A considerably longer answer with more detail."}

votes.jsonl:

{"kind": "votes", "schema_version": 1, "generated_at": "2026-08-23T04:00:00Z"}
{"vote_id": "10", "battle_id": "1", "voter_id": "v_a1", "choice": "win_a", "created_at": "2026-08-22T10:05:00Z", "weight": 1.0}
{"vote_id": "11", "battle_id": "1", "voter_id": "v_b2", "choice": "tie", "created_at": "2026-08-22T10:06:00Z", "weight": 0.5}
{"vote_id": "12", "battle_id": "2", "voter_id": "v_a1", "choice": "win_b", "created_at": "2026-08-22T11:09:00Z", "weight": 1.0}

Reading it: battle 1 was won by its slot-A model (gpt-alpha) at full weight, plus a half-weight tie (which counts as half a win for each side, METHODOLOGY.md §2). Battle 2 was won by its slot-B model (gpt-alpha again — note the slots swapped). Running

humaneval-ratings compute --battles battles.jsonl --votes votes.jsonl \
    --out leaderboard.json --seed 42

yields a chat-writing leaderboard where both models are provisional (3 votes < the default --min-votes 30) with rank: null, gpt-alpha at order 1, and style_control: false (3 votes < the default --style-min-votes 50), so the SC fields mirror the raw ones.

8. Version history

  • v1 (2026-08-23): initial spec — battles + votes JSONL with header lines; leaderboard schema v1.
  • v2 (2026-09-08, humaneval-ratings 1.1.0 / humaneval-exporter 1.1.0): battle records gain modality; new recording modality (section 3.1) with per-slot recording_* object references, assisted flags and telemetry_* for computer-use battles, no response_* texts. Text records are unchanged apart from modality: "text". Exporter rule 8: assisted battles are excluded. leaderboard.json schema stays 1 (only metadata.package_version and dump_schema_version change); the engine's numbers for v1 dumps are byte-identical to 1.0.0.