Skip to content

The ladder: how the bot is measured against past versions of itself

You get one honest answer out of this page: whether today's bot is actually better than last hour's bot, and how much of that answer you should believe.

Everything here lives under royalelearn/ladder/. You do not have to configure any of it to start a run. You do have to read it before you trust a number.

The deep version is docs/harness-spec.md section 11, and the reasoning behind the whole harness is in docs/design.md. This page is the short one.

Why you cannot judge the bot by its reward

Your reward function is a small piece of Python that scores what just happened in a battle. The training loop tries to make the reward go up. So the reward going up tells you the loop is working. It does not tell you the bot is good.

Two reasons, and both bite in practice.

The reward is zero sum. The bot plays itself. Both seats of a battle are in the same batch of training data. Every shipped reward gives one seat exactly what it takes from the other, so the average reward across the batch is zero by construction and stays zero forever. It is a flat line no matter how good or bad the bot gets. The comment in royalelearn/metrics/records.py around line 170 explains why two numbers are published per reward term instead of one, for exactly this reason.

The opponent moves. Even with a reward that was not zero sum, the bot's win rate in training is a win rate against itself. A bot that gets worse in a way that its own copy cannot punish still wins half its battles. Half is what self-play always reports, at every skill level, including none.

So the only way to ask "is it better" is to make it play something that does not move. That is what the ladder is.

The pieces

File What it holds
ladder/snapshots.py The archive of frozen copies of the bot, on disk
ladder/pool.py Who is in the ladder, who the champion is
ladder/matchmaker.py Who each battle plays
ladder/evaluate.py Playing a measured set of battles between two of them
ladder/gate.py Whether a new copy is allowed to join
ladder/rating.py Turning a pile of results into a number per player
ladder/eviction.py Which copies the matchmaker stops drawing, to keep it cheap
ladder/results.py The append-only log of every measured battle

In your run folder they write to ladder/games.jsonl, ladder/gates/*.json, ladder/ratings/*.json and ladder/eval_seeds.json.

The frozen pool, and why old versions are kept

A snapshot is a frozen copy of the bot's decision-making network, saved to disk and never trained again. The pool is the set of snapshots the bot still plays against.

The code takes a snapshot at a fixed cadence measured in game steps, not in training iterations. royalelearn/coordinator.py does this in _maybe_gate. The shipped values are in examples/configs/:

Profile candidate_every_env_steps
laptop.json 4,000,000
workstation.json 2,000,000
smoke.json 32

Steps rather than iterations, because an iteration changes size when you change the batch size or the number of workers, and then a cadence in iterations quietly means something else.

Old snapshots are kept rather than replaced for one reason. A bot that only ever plays its most recent self forgets how to beat the things it already beat. It drifts into a corner of the game where its recent ancestors are weak, wins there, and loses to something it crushed an hour ago. Keeping the weak old versions in the pool is what stops that.

Snapshots are kept forever in the archive. What is bounded is how many the matchmaker draws from, because that draw runs in Python once per battle. pool_working_size in royalelearn/config.py defaults to 48. When the sampler grows past that, ladder/eviction.py drops members, and it does not drop the oldest. It keeps a spread across the rating range, plus the two scripted anchors, the run's very first snapshot, every snapshot that was ever champion, and every seed snapshot (next section). Dropping the oldest would throw away exactly the weak opponents the pool exists to preserve. Evicted snapshots leave the sampler but stay in the archive and in the result log, so their results still count toward everyone's rating.

Starting the pool with a policy you already have

ladder.seed_snapshots puts frozen policies in the pool before the first battle. Each one joins as seed:<name>, is drawn like any other member, and is never evicted. It is not the run's first snapshot, so rating_above_v0 still measures from the run's own start.

A seed is a folder holding actor.safetensors and spec.json. That is what a run writes under runs/<run>/snapshots/<digest>/, and what a RoyaleImitate actor artifact holds. The config pins the weights by their sha256, so the same folder holding other weights is a different run:

"seed_snapshots": [
  {"name": "frozen", "path": "runs/<old-run>/snapshots/<digest>", "sha256": "<64 hex>"}
]

To get the digest:

python -c "import hashlib,sys; print(hashlib.sha256(open(sys.argv[1],'rb').read()).hexdigest())" runs/<old-run>/snapshots/<digest>/actor.safetensors

The run refuses to start if the digest does not match, and prints the one it found. It also refuses a policy built for a different network or observation, and names each field that differs. The run copies the weights into its own snapshots/ folder, so a resume does not need the original folder.

To train against one fixed opponent and nothing else: one seed, mix set to [0, 1, 0], and candidate_every_env_steps and floor_admit_every_env_steps set past the run's end, so the gate never adds anything. tests/test_seed_snapshots.py checks that every pool battle of such a run meets the seed.

A seed can also carry a weight, which multiplies how often it is drawn (default 1). With two seeds weighted 2 and 3 and pfsp_weighting set to "uniform", the pool meets them 40% and 60% of the time. Every other pool member keeps a weight of 1.

Training one deck against a field of others

ladder.seat_decks trains one deck, and only that deck, against the decks it will meet:

"seat_decks": {
  "deck": ["Giant", "Musketeer", "MiniPekka", "Archer", "Fireball", "Zap", "Goblins", "Minions"],
  "field": [["Knight", "Valkyrie", "HogRider", "Cannon", "Log", "Skeletons", "Arrows", "Musketeer"]]
}

It sets the decks each battle deals, from the battle's role:

  • A mirror battle deals deck to both seats. The learner plays both, so both seats train deck.
  • A pool or scripted battle deals deck to the learner's seat and a deck from field to the other. A frozen snapshot, a seed or a scripted opponent plays that seat, and its rows are not trained on. So the learner meets a fixed player of the field deck, never a copy of itself that is learning it.

field is drawn uniformly; list a deck twice to weight it, or leave field out for random decks. In a pool or scripted battle the learner's seat is fixed per battle, blue in even battles and red in odd ones, so the seats stay balanced. mix still decides how many battles are mirror battles. With seed_snapshots and a pool share, the field seat can be one policy you already have, and opponent_mode sets how that player picks its moves: it samples by default, "argmax" takes its most likely move, and "plugin:<module>:<function>" uses a decode you write. Without seat_decks, battles deal whatever the environment's own state mutator deals. tests/test_seat_decks.py checks the deal and the seats.

Who the bot plays in a given battle

Every battle draws its opponent from a three-way mixture. LadderConfig.mix in royalelearn/config.py defaults to (0.50, 0.35, 0.15), and all three shipped configs use it unchanged:

  • 50% mirror. The bot plays its live self. Both seats are trainable, which is where most of the training data comes from.
  • 35% pool. The bot plays one frozen snapshot.
  • 15% scripted. The bot plays one hard-coded opponent from scripted_opponents. By default that is one of the two anchors below.

Before the first snapshot exists there is nothing in the pool, so a pool draw falls back to scripted (matchmaker.py, in assign). That keeps the number of trainable rows per battle exactly what the mixture says it is, which the loop checks every iteration.

LadderConfig.scripted_opponents picks the hard-coded opponents your bot trains against. With the default, the scripted share only meets the two anchors: one never plays a card and the other plays a random legal card on one decision in ten. Set it to, for example, ["scripted:push", "scripted:defend", "scripted:patient"] and each scripted battle draws one of those, uniformly, every episode. So does a pool battle until the first snapshot exists. push plays a card as soon as one is playable, as far up the board as the rules allow. patient does the same once three of its cards are playable. The list changes training only: the anchors the gate and the rating use do not change. A name that is not a scripted opponent, a name listed twice, or an empty list while the mixture has pool or scripted battles is refused when the config loads.

At most max_resident_opponents snapshots are loaded on the GPU at once. That defaults to 2. Those two are redrawn only when the pool changes, never in the middle of an iteration.

How opponents are weighted

The config asks for PFSP weighting. PFSP means "prioritised fictitious self-play", and in plain words it means picking opponents that still beat you rather than picking uniformly at random, because those are the ones there is something to learn from. pfsp_weighting defaults to "hard".

The weight of an opponent depends on how well the live bot is predicted to score against it. The live bot has no fitted rating of its own, because it is never measured under its own name (see the broken rows section below). So its newest snapshot stands in for it: the last copy of the bot that a gate admitted. That copy can be up to one snapshot cadence behind the live bot.

Until 2026-09-27 nothing stood in. Every prediction was 0.5, and every run drew its opponents uniformly. Before the first snapshot has been rated, the draw is still uniform. Each metrics row says which happened: ladder/pfsp_effective is 1 when the draw used a rating and 0 when it fell back to uniform, and ladder/pfsp_learner_step is the env step of the snapshot that stood in.

You can see the weighting in one command. The last argument is the stand-in's rating:

python -c "from royalelearn.config import LadderConfig; from royalelearn.api.ladder import RatingTable; from royalelearn.ladder.matchmaker import MixMatchmaker; m=MixMatchmaker(1,LadderConfig()); t=RatingTable(rating={'snap:v0':0.0,'snap:v1':200.0,'snap:v2':400.0,'snap:v3':-200.0},se={},anchor='scripted:noop',draw_nu=None,n_games={},transitivity_residual=0.0,converged=True,iterations=3); print(m.weights(('snap:v0','snap:v1','snap:v2','snap:v3'),t,'hard',t.rating['snap:v0']))"

Measured 2026-09-27, that prints [0.16686453 0.31982403 0.43632903 0.0769824 ]: the strongest snapshot, snap:v2, is drawn most. Leave off the last argument and it prints [0.25 0.25 0.25 0.25], the uniform fallback.

The two scripted anchors

An anchor is an opponent whose strength never changes. They matter because everything else in the ladder is improving, so a rating measured only against other snapshots is a rating on a scale that is itself sliding. royalelearn/rollout/scripted.py ships six scripted opponents. These two are the anchors:

  • scripted:noop does nothing at all. It never plays a card.
  • scripted:random_legal plays a legal random move, but only 10% of the time. The other 90% it does nothing. RANDOM_LEGAL_NOOP_PROB = 0.9 in that file. Uniform over every legal move would play a card the instant one was affordable, which is a strange thing to train against.

They exist before the first snapshot does, so the rating scale has a fixed point from the first battle. scripted:noop is pinned at rating 0 and every other number is relative to it.

They are deliberately weak. If your bot is not beating scripted:noop comfortably, nothing else on this page is worth reading yet.

The gate: what a new snapshot has to clear to join

The champion is the current best snapshot. A candidate is the fresh snapshot just taken. The gate is the test that decides whether the candidate joins the pool and whether it takes the champion's place. It is in ladder/gate.py, class WilsonGate.

Three conditions. Defaults from GateConfig in royalelearn/config.py.

1. It beats the champion. 1,000 battles against the champion. A score rate is wins plus half the draws, over battles played. The gate does not use the raw score rate. It uses the bottom of a 95% confidence interval around it, which is the lowest score rate still consistent with what was observed. That bottom has to reach 0.52. Check the threshold yourself:

python -c "from royalelearn.ladder.rating import wilson_interval, elo_of_score; [print(p, round(wilson_interval(p,1000)[0],5)) for p in (0.550,0.551,0.552)]; print('elo at 0.552:', round(elo_of_score(0.552),1))"

Measured 2026-09-22: an observed 55.0% does not clear it, 55.1% does, and 55.2% is about 36 Elo of improvement. A threshold at 0.50 would admit anything merely not worse, and the pool would fill with sideways moves.

2. No regression against the anchors. 200 battles against each anchor. The candidate has to score within 2 percentage points of what the champion scores against the same anchor. This catches the bot that found a blind spot its recent ancestors share, wins by exploiting it, and has quietly forgotten how to play.

3. No collapse against the wider pool. 100 battles against each of 8 snapshots chosen to span the rating range. The candidate's average score has to reach what the fitted model predicts the champion would score against those same 8, less one standard error.

Early in a run the pool may hold nothing besides the champion and the two anchors. Then there is nothing to sample, and the gate records condition 3 as skipped, with that reason. The verdict rests on conditions 1 and 2, so a candidate that passes both is admitted and becomes champion. Until the fix of 2026-09-24 (9468b86) the same case was recorded as a pass with n of 0.

That is 1,000 + 400 + 800 = 2,200 battles per gate at most. A gate stops as soon as its outcome is settled and records the conditions it did not play as skipped: when condition 1 fails the gate costs 1,000 battles, and when condition 2 fails it costs 1,400 (ladder/gate.py).

The first gate of a run can cost more. The run's first snapshot joins the pool without a gate, so as champion it has no record against the anchors. When condition 2 needs that record, the champion plays each anchor itself, 200 battles on the same seeds, and later gates read the result instead of playing it again.

The verdict is not a single yes or no:

Conditions 1 and 2 Condition 3 Result
pass pass Admitted to the pool, and becomes champion
pass fail Admitted to the pool, flagged cycle, does not become champion
fail anything Discarded, never retried

The middle row is the interesting one. A candidate that beats the champion but falls apart against the rest of the pool has found a rock-paper-scissors loop, not an improvement. It is a useful opponent to keep, and ladder/eviction.py protects flagged snapshots when it prunes, because they are the ones a single rating number describes worst.

A failed candidate is thrown away and not retried. Consecutive failures are a signal, not an error: the gate_starved alarm fires at 5 in a row (AlarmConfig.gate_failures). Separately, floor_admit_every_env_steps (50,000,000 in all three configs) admits a snapshot unconditionally if nothing has passed in that long, so a plateau cannot starve the pool of fresh opponents. A floor admission never promotes.

Every decision is written to ladder/gates/<candidate>.json with the numbers behind it. A real one, trimmed, from a smoke-sized run on 2026-09-22:

{"candidate":"snap:v2","champion":"snap:v0","admit":false,"conditions":{
  "beats_champion":{"passed":false,"n":2,"observed":0.5,"bound":0.0945,"reference":0.52},
  "anchors:scripted:noop":{"passed":true,"n":2,"observed":1.0,"bound":0.6119},
  "anchors:scripted:random_legal":{"passed":false,"n":2,"observed":0.0,"bound":0.1498}}}

That file predates the early stop. Today a gate whose first condition fails plays nothing more, so its file carries beats_champion and records the anchors and the pool as skipped.

You can re-run a gate from stored snapshots without touching the training run:

python -m royalelearn gate --run runs/<your-run> --candidate snap:v3

How the battles are counted

Evaluation battles are kept separate from training battles on purpose (ladder/evaluate.py). They record no training data. They use a full match with no step cap, because a cut-short match is scored on crowns nobody has taken yet and almost every game comes out a draw. And they use a set of seeds drawn once at the start of the run and reused forever, so every player is measured on the same starting positions.

Each pairing plays each seed twice, with the sides swapped. Blue and red are not symmetric, so playing both sides removes the side advantage exactly instead of averaging over it. The unit of measurement is the seed, not the battle, which is why champion_games: 1000 means 500 seeds and eval_seed_count defaults to 500. ladder/paired_rho in your metrics file says how correlated a seed's two sides were, which is how much of a battle the starting position decided rather than the two players.

The rating, and what it is not

Two different numbers in the metrics file, and they are not interchangeable.

ladder/elo_readout is the dashboard figure. Elo is the chess-style rating where a higher number means stronger and a 400-point gap means roughly a 10-to-1 favourite. This one updates live as training battles finish, starting from 1200 (EloReadout in ladder/rating.py, k_factor = 32). It depends on the order the results arrive in, which under parallel workers is not reproducible. It is for watching. No decision reads it. It is fed by training games, so it moves with scripted_opponents: two runs with different lists do not compare on it.

The fitted rating is the real one. It refits from scratch over the entire stored result log every refit_every_iterations iterations, which defaults to 10. Same games in, same numbers out, in any order, on any machine. It is a Bradley-Terry-Davidson fit, which means it finds the one rating per player that best explains all the recorded wins, losses and draws at once, with draws modelled explicitly rather than counted as half a win.

You can run it yourself on any finished run. It only reads the log, so it needs no GPU and no environment:

python -m royalelearn rate --run runs/<your-run>

Real output, 2026-09-22, from a 10-iteration smoke-sized run with 24 evaluation battles in total:

member                              rating      se           95% interval   games
scripted:random_legal                493.4   177.7   [   145.2,    841.6]       8
snap:v3                              303.9   180.8   [   -50.5,    658.3]       6
snap:v4                              303.9   180.8   [   -50.5,    658.3]       6
snap:v1                              167.0   170.0   [  -166.2,    500.1]       6
snap:v2                              167.0   170.0   [  -166.2,    500.1]       6
snap:v0                              129.6   160.2   [  -184.4,    443.7]       8
scripted:noop                          0.0     0.0   [     0.0,      0.0]       8
anchor scripted:noop at 0; transitivity residual 0.0000

Read that table as a warning, not as a result. The intervals are roughly 700 Elo wide, because each player has 6 or 8 battles behind it. The random_legal anchor sitting on top is noise. This is what the output looks like when there is not nearly enough evidence, and it is the shape you will see from any short run.

What the rating means: one player is stronger than another inside this pool, under this deck protocol, on these seeds.

What it does not mean: anything about real players, any in-game trophy count, or any ladder rank. There is no bridge in this repository between these numbers and a human opponent. A run that climbs 400 Elo has climbed 400 Elo against its own history. That is genuine progress and it is also the only claim the number supports.

When even the within-pool reading breaks: ladder/transitivity_residual is the share of well-measured pairs that a single number per player fails to explain. It is near zero on a population where strength is a straight line. It gets large when A beats B beats C beats A, which one number per player cannot describe at all. The transitivity alarm fires above 0.10 (AlarmConfig.transitivity_residual). Above that, treat the whole rating column as unreliable rather than as a slightly noisy truth.

Reading the ladder rows in your metrics file

Every iteration writes one flat row to runs/<your-run>/metrics.jsonl. To see just the ladder part of the last row:

python -c "import json,sys; rows=[json.loads(l) for l in open(sys.argv[1],encoding='utf-8') if l.strip()]; print(json.dumps({k:v for k,v in rows[-1].items() if k.startswith('ladder/')},indent=1))" runs/<your-run>/metrics.jsonl

The rows that are worth your attention, built in royalelearn/metrics/records.py, ladder_fields:

Key What it tells you
ladder/champion_id Which snapshot the next candidate must beat
ladder/pool_size Everything in the archive, including the 2 anchors
ladder/sampler_size Snapshots the matchmaker can actually draw. Anchors are not counted here
ladder/gate_attempts, ladder/gate_passes How many gates ran, how many admitted
ladder/consecutive_gate_failures The plateau signal. 5 trips the gate_starved alarm
ladder/gate_failed_condition Which of the three conditions failed, or none
ladder/gate_observed_rate The candidate's raw score rate against the champion. Appears once a gate decision carries its champion condition
ladder/gate_lower_bound The bottom of its interval, which is what the gate compares to 0.52. Same condition as the row above
ladder/rating/<member> One pool member's fitted rating, only after the first refit
ladder/rating_ci95_lo/<member> and ..._hi/<member> That rating's interval. A wide one means you do not know yet
ladder/transitivity_residual Above 0.10, stop trusting the rating column. Appears once a rating fit exists
ladder/paired_rho How much the starting position decides, rather than the players, in the last gate's champion comparison. Appears once a gate has played the champion and the correlation is defined. Runs before 2026-09-24's fix wrote 0.0 in every row, so ignore it there
ladder/draw_rate_eval Draw rate in evaluation battles
ladder/gate_seconds_frac Share of this run's wall clock spent gating rather than training, so far. On every row, and 0 until the first gate. Absent only after resuming a checkpoint written before runs recorded the gate total
ladder/elo_readout The live dashboard Elo. Never a decision. Appears once a training game has been scored this run

A row that is missing a key is not a bug. The harness uses absence to say "nothing to report" rather than publishing a zero that reads like a measurement. So do not read the first rows of a run and conclude the ladder is broken: most of the table above only starts once a gate has run or a rating has been fitted, which takes a while. royalelearn/metrics/schema.py lists every key that may be absent in CONDITIONAL, each with the condition that makes it appear. Print them for yourself:

python -c "from royalelearn.metrics import schema; [print(k, '->', v) for k, v in schema.CONDITIONAL.items()]"

Rows you should not trust today

Three of the ladder numbers do not mean what their names say. All three come from the same root cause: the live bot is never measured under its own name. Every evaluation battle is a snapshot against something else, because the gate hands snapshot ids to the evaluator. So the id learner never appears in the evaluation results, and anything looked up under that id comes back empty.

ladder/score_vs_noop and ladder/score_vs_random_legal appear on probe iterations, and nowhere else. Until commit 72c1215 on 2026-09-22 they were published as 0.5 in every row, which was indistinguishable from a genuine even contest; then they were omitted, because the pair they asked about had no games and never could. Measured 2026-09-22 across the metrics rows in runs/: 127 rows written before that fix carry a flat 0.5. If you are reading an older run's file, ignore those two columns entirely.

The decision that was open is now taken. Set ladder.probe_every_iterations and every ladder.probe_every_iterations iterations the run plays the LIVE bot against each rung in ladder.probe_opponents and publishes, for each of them:

key what it is
ladder/score_vs/{rung} the live bot's score against that rung, 1.0 for a win
ladder/score_vs_n/{rung} how many SEEDS it was measured over, each played from both sides
ladder/score_vs_ci95_lo/{rung}, ..._hi/{rung} the interval around it

ladder/score_vs_noop and ladder/score_vs_random_legal are the same numbers under their old names, for the two rungs that have always been the anchors. It is off by default, because it costs battles nobody was paying for: at the shipped 40 battles a rung it is under 4% of one gate. Read the interval before the score. At 20 seeds the instrument's own repeat spread against scripted:random_legal is about 18 points with the policy unchanged, so a smaller move than that is the instrument. scripted:noop does not act, so it does not have that noise.

ladder/rating_above_v0 was minus the first snapshot's rating, and is now absent instead. It is meant to be the live bot's rating above the run's first snapshot. The live bot has no fitted rating, because every evaluation game is a snapshot against something, so the lookup returned a default of 0.0 and the row published 0 - rating[v0]. On one run that read -93.9, -146.2, -191.7 and -129.6, which looks exactly like a bot falling behind its own opening snapshot and is nothing of the kind. Since 2026-09-22 the key is published only when the fit holds both sides (royalelearn/metrics/records.py, the guard above fields["ladder/rating_above_v0"]), so it is simply absent until the live bot is evaluated under its own id. If you are reading a run's file from before that, ignore the column. To see the old shape for yourself:

python -c "import json,glob; [print(r.get('ladder/rating_above_v0'), r['ladder/rating/snap:v0']) for p in glob.glob('runs/*/metrics.jsonl') for r in map(json.loads, open(p,encoding='utf-8')) if 'ladder/rating/snap:v0' in r]"

Read ladder/rating/<member> directly instead. Those columns are correct.

ladder/eval_games_total was structurally zero and has just been fixed. It used to read a counter that nothing ever incremented, because the only caller of the counting method hard codes the training label. The same commit changed it to count the evaluation log directly. It is still 0 in every metrics row on disk, including runs made after the fix, but that is now the truthful answer: those runs never played an evaluation battle. Count the log yourself:

python -c "import json,sys; print(sum(1 for l in open(sys.argv[1],encoding='utf-8') if l.strip() and json.loads(l)['kind']=='eval'))" runs/<your-run>/ladder/games.jsonl

What has actually been run, and what has not

This matters more than any of the above.

The shipped gate has only just been run. Measured 2026-09-22: of the 50 run folders under runs/ with a metrics file, 24 ran at least one gate, and every one used champion_games: 2, which is the smoke setting. After that, a run used the same 1,000-battle gate that laptop.json and workstation.json ship. Reading its first real gate showed that conditions 2 and 3 could not fail there: the champion had no record against the anchors, and the pool held nothing to sample. Both were fixed on 2026-09-24. No numbers from a real gate are written up on this page yet.

The ladder has never had more than a handful of members. On 2026-09-22 the largest pool in any run on disk was 3, which is 2 anchors plus 1 snapshot. Eviction, the stratified sample in condition 3, and the anti-forgetting argument for keeping old snapshots are all untested at the pool sizes they were designed for.

What a gate costs is measured per battle, not per gate. On 2026-09-23 one evaluation battle took 8.02 s network against network and 0.195 s scripted against scripted on this laptop, the median of three full matches (EVAL_BATTLE_SECONDS in ladder/gate.py). So a full 2,200-battle gate is hours, and preflight prints a gate's battles and hours at that rate before a run starts (rollout/preflight.py). ladder/gate_seconds_frac tells you what gating is costing a run, and the schema flags it above 0.1.

What is measured, on a 4-core laptop with an RTX 3050 4 GB and 7.8 GB of RAM shared with six other jobs, on 2026-09-22: in a diagnostic run at 2 workers by 24 battles and 8,192 timesteps per iteration with minibatch 256, an iteration took 30 to 35 seconds. About 25 seconds of that was the learning update and 5 to 8 seconds was collecting battles, with the update taking 71 to 83% of every iteration across 7 iterations. That geometry is not the laptop profile, which uses 3 workers by 32 battles and 32,768 timesteps. It does say that on this machine training time is GPU time rather than simulator time.

Where to go next

  • The dense version of all of this: docs/harness-spec.md, section 11.
  • Why the harness is shaped this way: docs/design.md.
  • A single head-to-head between two members, outside the training loop: python -m royalelearn eval --run runs/<your-run> --a snap:v3 --b snap:v0 --seeds 50
  • The raw evidence behind every rating: runs/<your-run>/ladder/games.jsonl, one JSON line per battle.