The ladder: how the bot is measured against past versions of itself¶
You get one honest answer out of this page: whether today's bot is actually better than last hour's bot, and how much of that answer you should believe.
Everything here lives under royalelearn/ladder/. You do not have to configure any of it
to start a run. You do have to read it before you trust a number.
The deep version is docs/harness-spec.md section 11, and the reasoning
behind the whole harness is in docs/design.md. This page is the short one.
Why you cannot judge the bot by its reward¶
Your reward function is a small piece of Python that scores what just happened in a battle. The training loop tries to make the reward go up. So the reward going up tells you the loop is working. It does not tell you the bot is good.
Two reasons, and both bite in practice.
The reward is zero sum. The bot plays itself. Both seats of a battle are in the same
batch of training data. Every shipped reward gives one seat exactly what it takes from the
other, so the average reward across the batch is zero by construction and stays zero forever.
It is a flat line no matter how good or bad the bot gets. The comment in
royalelearn/metrics/records.py around line 170 explains why two numbers are published per
reward term instead of one, for exactly this reason.
The opponent moves. Even with a reward that was not zero sum, the bot's win rate in training is a win rate against itself. A bot that gets worse in a way that its own copy cannot punish still wins half its battles. Half is what self-play always reports, at every skill level, including none.
So the only way to ask "is it better" is to make it play something that does not move. That is what the ladder is.
The pieces¶
| File | What it holds |
|---|---|
ladder/snapshots.py |
The archive of frozen copies of the bot, on disk |
ladder/pool.py |
Who is in the ladder, who the champion is |
ladder/matchmaker.py |
Who each battle plays |
ladder/evaluate.py |
Playing a measured set of battles between two of them |
ladder/gate.py |
Whether a new copy is allowed to join |
ladder/rating.py |
Turning a pile of results into a number per player |
ladder/eviction.py |
Which copies the matchmaker stops drawing, to keep it cheap |
ladder/results.py |
The append-only log of every measured battle |
In your run folder they write to ladder/games.jsonl, ladder/gates/*.json,
ladder/ratings/*.json and ladder/eval_seeds.json.
The frozen pool, and why old versions are kept¶
A snapshot is a frozen copy of the bot's decision-making network, saved to disk and never trained again. The pool is the set of snapshots the bot still plays against.
The code takes a snapshot at a fixed cadence measured in game steps, not in training
iterations. royalelearn/coordinator.py does this in _maybe_gate. The shipped values are
in examples/configs/:
| Profile | candidate_every_env_steps |
|---|---|
laptop.json |
4,000,000 |
workstation.json |
2,000,000 |
smoke.json |
32 |
Steps rather than iterations, because an iteration changes size when you change the batch size or the number of workers, and then a cadence in iterations quietly means something else.
Old snapshots are kept rather than replaced for one reason. A bot that only ever plays its most recent self forgets how to beat the things it already beat. It drifts into a corner of the game where its recent ancestors are weak, wins there, and loses to something it crushed an hour ago. Keeping the weak old versions in the pool is what stops that.
Snapshots are kept forever in the archive. What is bounded is how many the matchmaker draws
from, because that draw runs in Python once per battle. pool_working_size in
royalelearn/config.py defaults to 48. When the sampler grows past that,
ladder/eviction.py drops members, and it does not drop the oldest. It keeps a spread across
the rating range, plus the two scripted anchors, the run's very first snapshot, every
snapshot that was ever champion, and every seed snapshot (next section). Dropping the oldest
would throw away exactly the weak opponents the pool exists to preserve. Evicted snapshots
leave the sampler but stay in the archive and in the result log, so their results still count
toward everyone's rating.
Starting the pool with a policy you already have¶
ladder.seed_snapshots puts frozen policies in the pool before the first battle. Each one
joins as seed:<name>, is drawn like any other member, and is never evicted. It is not the
run's first snapshot, so rating_above_v0 still measures from the run's own start.
A seed is a folder holding actor.safetensors and spec.json. That is what a run writes
under runs/<run>/snapshots/<digest>/, and what a RoyaleImitate actor artifact holds. The
config pins the weights by their sha256, so the same folder holding other weights is a
different run:
"seed_snapshots": [
{"name": "frozen", "path": "runs/<old-run>/snapshots/<digest>", "sha256": "<64 hex>"}
]
To get the digest:
python -c "import hashlib,sys; print(hashlib.sha256(open(sys.argv[1],'rb').read()).hexdigest())" runs/<old-run>/snapshots/<digest>/actor.safetensors
The run refuses to start if the digest does not match, and prints the one it found. It also
refuses a policy built for a different network or observation, and names each field that
differs. The run copies the weights into its own snapshots/ folder, so a resume does not
need the original folder.
To train against one fixed opponent and nothing else: one seed, mix set to [0, 1, 0], and
candidate_every_env_steps and floor_admit_every_env_steps set past the run's end, so the gate
never adds anything. tests/test_seed_snapshots.py checks that every pool battle of such a run
meets the seed.
A seed can also carry a weight, which multiplies how often it is drawn (default 1). With two
seeds weighted 2 and 3 and pfsp_weighting set to "uniform", the pool meets them 40% and 60%
of the time. Every other pool member keeps a weight of 1.
Training one deck against a field of others¶
ladder.seat_decks trains one deck, and only that deck, against the decks it will meet:
"seat_decks": {
"deck": ["Giant", "Musketeer", "MiniPekka", "Archer", "Fireball", "Zap", "Goblins", "Minions"],
"field": [["Knight", "Valkyrie", "HogRider", "Cannon", "Log", "Skeletons", "Arrows", "Musketeer"]]
}
It sets the decks each battle deals, from the battle's role:
- A mirror battle deals
deckto both seats. The learner plays both, so both seats traindeck. - A pool or scripted battle deals
deckto the learner's seat and a deck fromfieldto the other. A frozen snapshot, a seed or a scripted opponent plays that seat, and its rows are not trained on. So the learner meets a fixed player of the field deck, never a copy of itself that is learning it.
field is drawn uniformly; list a deck twice to weight it, or leave field out for random decks.
In a pool or scripted battle the learner's seat is fixed per battle, blue in even battles and red in
odd ones, so the seats stay balanced. mix still decides how many battles are mirror battles. With
seed_snapshots and a pool share, the field seat can be one policy you already have, and
opponent_mode sets how that player picks its moves: it samples by default, "argmax" takes its
most likely move, and "plugin:<module>:<function>" uses a decode you write. Without seat_decks, battles deal whatever the environment's own state mutator deals.
tests/test_seat_decks.py checks the deal and the seats.
Who the bot plays in a given battle¶
Every battle draws its opponent from a three-way mixture. LadderConfig.mix in
royalelearn/config.py defaults to (0.50, 0.35, 0.15), and all three shipped configs use
it unchanged:
- 50% mirror. The bot plays its live self. Both seats are trainable, which is where most of the training data comes from.
- 35% pool. The bot plays one frozen snapshot.
- 15% scripted. The bot plays one hard-coded opponent from
scripted_opponents. By default that is one of the two anchors below.
Before the first snapshot exists there is nothing in the pool, so a pool draw falls back to
scripted (matchmaker.py, in assign). That keeps the number of trainable rows per battle
exactly what the mixture says it is, which the loop checks every iteration.
LadderConfig.scripted_opponents picks the hard-coded opponents your bot trains against. With
the default, the scripted share only meets the two anchors: one never plays a card and the other
plays a random legal card on one decision in ten. Set it to, for example,
["scripted:push", "scripted:defend", "scripted:patient"] and each scripted battle draws one of
those, uniformly, every episode. So does a pool battle until the first snapshot exists. push
plays a card as soon as one is playable, as far up the board as the rules allow. patient does
the same once three of its cards are playable. The list changes training only: the anchors the
gate and the rating use do not change. A name that is not a scripted opponent, a name listed
twice, or an empty list while the mixture has pool or scripted battles is refused when the
config loads.
At most max_resident_opponents snapshots are loaded on the GPU at once. That defaults to
2. Those two are redrawn only when the pool changes, never in the middle of an iteration.
How opponents are weighted¶
The config asks for PFSP weighting. PFSP means "prioritised fictitious self-play", and in
plain words it means picking opponents that still beat you rather than picking uniformly at
random, because those are the ones there is something to learn from. pfsp_weighting
defaults to "hard".
The weight of an opponent depends on how well the live bot is predicted to score against it. The live bot has no fitted rating of its own, because it is never measured under its own name (see the broken rows section below). So its newest snapshot stands in for it: the last copy of the bot that a gate admitted. That copy can be up to one snapshot cadence behind the live bot.
Until 2026-09-27 nothing stood in. Every prediction was 0.5, and every run drew its opponents
uniformly. Before the first snapshot has been rated, the draw is still uniform. Each metrics row
says which happened: ladder/pfsp_effective is 1 when the draw used a rating and 0 when it fell
back to uniform, and ladder/pfsp_learner_step is the env step of the snapshot that stood in.
You can see the weighting in one command. The last argument is the stand-in's rating:
python -c "from royalelearn.config import LadderConfig; from royalelearn.api.ladder import RatingTable; from royalelearn.ladder.matchmaker import MixMatchmaker; m=MixMatchmaker(1,LadderConfig()); t=RatingTable(rating={'snap:v0':0.0,'snap:v1':200.0,'snap:v2':400.0,'snap:v3':-200.0},se={},anchor='scripted:noop',draw_nu=None,n_games={},transitivity_residual=0.0,converged=True,iterations=3); print(m.weights(('snap:v0','snap:v1','snap:v2','snap:v3'),t,'hard',t.rating['snap:v0']))"
Measured 2026-09-27, that prints [0.16686453 0.31982403 0.43632903 0.0769824 ]: the
strongest snapshot, snap:v2, is drawn most. Leave off the last argument and it prints
[0.25 0.25 0.25 0.25], the uniform fallback.
The two scripted anchors¶
An anchor is an opponent whose strength never changes. They matter because everything
else in the ladder is improving, so a rating measured only against other snapshots is a
rating on a scale that is itself sliding. royalelearn/rollout/scripted.py ships six scripted
opponents. These two are the anchors:
scripted:noopdoes nothing at all. It never plays a card.scripted:random_legalplays a legal random move, but only 10% of the time. The other 90% it does nothing.RANDOM_LEGAL_NOOP_PROB = 0.9in that file. Uniform over every legal move would play a card the instant one was affordable, which is a strange thing to train against.
They exist before the first snapshot does, so the rating scale has a fixed point from the
first battle. scripted:noop is pinned at rating 0 and every other number is relative to it.
They are deliberately weak. If your bot is not beating scripted:noop comfortably, nothing
else on this page is worth reading yet.
The gate: what a new snapshot has to clear to join¶
The champion is the current best snapshot. A candidate is the fresh snapshot just
taken. The gate is the test that decides whether the candidate joins the pool and whether
it takes the champion's place. It is in ladder/gate.py, class WilsonGate.
Three conditions. Defaults from GateConfig in royalelearn/config.py.
1. It beats the champion. 1,000 battles against the champion. A score rate is wins plus half the draws, over battles played. The gate does not use the raw score rate. It uses the bottom of a 95% confidence interval around it, which is the lowest score rate still consistent with what was observed. That bottom has to reach 0.52. Check the threshold yourself:
python -c "from royalelearn.ladder.rating import wilson_interval, elo_of_score; [print(p, round(wilson_interval(p,1000)[0],5)) for p in (0.550,0.551,0.552)]; print('elo at 0.552:', round(elo_of_score(0.552),1))"
Measured 2026-09-22: an observed 55.0% does not clear it, 55.1% does, and 55.2% is about 36 Elo of improvement. A threshold at 0.50 would admit anything merely not worse, and the pool would fill with sideways moves.
2. No regression against the anchors. 200 battles against each anchor. The candidate has to score within 2 percentage points of what the champion scores against the same anchor. This catches the bot that found a blind spot its recent ancestors share, wins by exploiting it, and has quietly forgotten how to play.
3. No collapse against the wider pool. 100 battles against each of 8 snapshots chosen to span the rating range. The candidate's average score has to reach what the fitted model predicts the champion would score against those same 8, less one standard error.
Early in a run the pool may hold nothing besides the champion and the two anchors. Then there is
nothing to sample, and the gate records condition 3 as skipped, with that reason. The verdict
rests on conditions 1 and 2, so a candidate that passes both is admitted and becomes champion.
Until the fix of 2026-09-24 (9468b86) the same case was recorded as a pass with n of 0.
That is 1,000 + 400 + 800 = 2,200 battles per gate at most. A gate stops as soon as its
outcome is settled and records the conditions it did not play as skipped: when condition 1 fails
the gate costs 1,000 battles, and when condition 2 fails it costs 1,400 (ladder/gate.py).
The first gate of a run can cost more. The run's first snapshot joins the pool without a gate, so as champion it has no record against the anchors. When condition 2 needs that record, the champion plays each anchor itself, 200 battles on the same seeds, and later gates read the result instead of playing it again.
The verdict is not a single yes or no:
| Conditions 1 and 2 | Condition 3 | Result |
|---|---|---|
| pass | pass | Admitted to the pool, and becomes champion |
| pass | fail | Admitted to the pool, flagged cycle, does not become champion |
| fail | anything | Discarded, never retried |
The middle row is the interesting one. A candidate that beats the champion but falls apart
against the rest of the pool has found a rock-paper-scissors loop, not an improvement. It is
a useful opponent to keep, and ladder/eviction.py protects flagged snapshots when it prunes,
because they are the ones a single rating number describes worst.
A failed candidate is thrown away and not retried. Consecutive failures are a signal, not an
error: the gate_starved alarm fires at 5 in a row (AlarmConfig.gate_failures). Separately,
floor_admit_every_env_steps (50,000,000 in all three configs) admits a snapshot
unconditionally if nothing has passed in that long, so a plateau cannot starve the pool of
fresh opponents. A floor admission never promotes.
Every decision is written to ladder/gates/<candidate>.json with the numbers behind it. A
real one, trimmed, from a smoke-sized run on 2026-09-22:
{"candidate":"snap:v2","champion":"snap:v0","admit":false,"conditions":{
"beats_champion":{"passed":false,"n":2,"observed":0.5,"bound":0.0945,"reference":0.52},
"anchors:scripted:noop":{"passed":true,"n":2,"observed":1.0,"bound":0.6119},
"anchors:scripted:random_legal":{"passed":false,"n":2,"observed":0.0,"bound":0.1498}}}
That file predates the early stop. Today a gate whose first condition fails plays nothing more,
so its file carries beats_champion and records the anchors and the pool as skipped.
You can re-run a gate from stored snapshots without touching the training run:
How the battles are counted¶
Evaluation battles are kept separate from training battles on purpose
(ladder/evaluate.py). They record no training data. They use a full match with no step cap,
because a cut-short match is scored on crowns nobody has taken yet and almost every game comes
out a draw. And they use a set of seeds drawn once at the start of the run and reused forever,
so every player is measured on the same starting positions.
Each pairing plays each seed twice, with the sides swapped. Blue and red are not symmetric,
so playing both sides removes the side advantage exactly instead of averaging over it. The
unit of measurement is the seed, not the battle, which is why champion_games: 1000 means 500
seeds and eval_seed_count defaults to 500. ladder/paired_rho in your metrics file says how
correlated a seed's two sides were, which is how much of a battle the starting position decided
rather than the two players.
The rating, and what it is not¶
Two different numbers in the metrics file, and they are not interchangeable.
ladder/elo_readout is the dashboard figure. Elo is the chess-style rating where a
higher number means stronger and a 400-point gap means roughly a 10-to-1 favourite. This one
updates live as training battles finish, starting from 1200 (EloReadout in
ladder/rating.py, k_factor = 32). It depends on the order the results arrive in, which
under parallel workers is not reproducible. It is for watching. No decision reads it. It is
fed by training games, so it moves with scripted_opponents: two runs with different lists do
not compare on it.
The fitted rating is the real one. It refits from scratch over the entire stored result
log every refit_every_iterations iterations, which defaults to 10. Same games in, same
numbers out, in any order, on any machine. It is a Bradley-Terry-Davidson fit, which means it
finds the one rating per player that best explains all the recorded wins, losses and draws at
once, with draws modelled explicitly rather than counted as half a win.
You can run it yourself on any finished run. It only reads the log, so it needs no GPU and no environment:
Real output, 2026-09-22, from a 10-iteration smoke-sized run with 24 evaluation battles in total:
member rating se 95% interval games
scripted:random_legal 493.4 177.7 [ 145.2, 841.6] 8
snap:v3 303.9 180.8 [ -50.5, 658.3] 6
snap:v4 303.9 180.8 [ -50.5, 658.3] 6
snap:v1 167.0 170.0 [ -166.2, 500.1] 6
snap:v2 167.0 170.0 [ -166.2, 500.1] 6
snap:v0 129.6 160.2 [ -184.4, 443.7] 8
scripted:noop 0.0 0.0 [ 0.0, 0.0] 8
anchor scripted:noop at 0; transitivity residual 0.0000
Read that table as a warning, not as a result. The intervals are roughly 700 Elo wide,
because each player has 6 or 8 battles behind it. The random_legal anchor sitting on top is
noise. This is what the output looks like when there is not nearly enough evidence, and it is
the shape you will see from any short run.
What the rating means: one player is stronger than another inside this pool, under this deck protocol, on these seeds.
What it does not mean: anything about real players, any in-game trophy count, or any ladder rank. There is no bridge in this repository between these numbers and a human opponent. A run that climbs 400 Elo has climbed 400 Elo against its own history. That is genuine progress and it is also the only claim the number supports.
When even the within-pool reading breaks: ladder/transitivity_residual is the share of
well-measured pairs that a single number per player fails to explain. It is near zero on a
population where strength is a straight line. It gets large when A beats B beats C beats A,
which one number per player cannot describe at all. The transitivity alarm fires above
0.10 (AlarmConfig.transitivity_residual). Above that, treat the whole rating column as
unreliable rather than as a slightly noisy truth.
Reading the ladder rows in your metrics file¶
Every iteration writes one flat row to runs/<your-run>/metrics.jsonl. To see just the
ladder part of the last row:
python -c "import json,sys; rows=[json.loads(l) for l in open(sys.argv[1],encoding='utf-8') if l.strip()]; print(json.dumps({k:v for k,v in rows[-1].items() if k.startswith('ladder/')},indent=1))" runs/<your-run>/metrics.jsonl
The rows that are worth your attention, built in royalelearn/metrics/records.py,
ladder_fields:
| Key | What it tells you |
|---|---|
ladder/champion_id |
Which snapshot the next candidate must beat |
ladder/pool_size |
Everything in the archive, including the 2 anchors |
ladder/sampler_size |
Snapshots the matchmaker can actually draw. Anchors are not counted here |
ladder/gate_attempts, ladder/gate_passes |
How many gates ran, how many admitted |
ladder/consecutive_gate_failures |
The plateau signal. 5 trips the gate_starved alarm |
ladder/gate_failed_condition |
Which of the three conditions failed, or none |
ladder/gate_observed_rate |
The candidate's raw score rate against the champion. Appears once a gate decision carries its champion condition |
ladder/gate_lower_bound |
The bottom of its interval, which is what the gate compares to 0.52. Same condition as the row above |
ladder/rating/<member> |
One pool member's fitted rating, only after the first refit |
ladder/rating_ci95_lo/<member> and ..._hi/<member> |
That rating's interval. A wide one means you do not know yet |
ladder/transitivity_residual |
Above 0.10, stop trusting the rating column. Appears once a rating fit exists |
ladder/paired_rho |
How much the starting position decides, rather than the players, in the last gate's champion comparison. Appears once a gate has played the champion and the correlation is defined. Runs before 2026-09-24's fix wrote 0.0 in every row, so ignore it there |
ladder/draw_rate_eval |
Draw rate in evaluation battles |
ladder/gate_seconds_frac |
Share of this run's wall clock spent gating rather than training, so far. On every row, and 0 until the first gate. Absent only after resuming a checkpoint written before runs recorded the gate total |
ladder/elo_readout |
The live dashboard Elo. Never a decision. Appears once a training game has been scored this run |
A row that is missing a key is not a bug. The harness uses absence to say "nothing to report"
rather than publishing a zero that reads like a measurement. So do not read the first rows of a
run and conclude the ladder is broken: most of the table above only starts once a gate has run or
a rating has been fitted, which takes a while. royalelearn/metrics/schema.py lists every key
that may be absent in CONDITIONAL, each with the condition that makes it appear. Print them for
yourself:
python -c "from royalelearn.metrics import schema; [print(k, '->', v) for k, v in schema.CONDITIONAL.items()]"
Rows you should not trust today¶
Three of the ladder numbers do not mean what their names say. All three come from the same
root cause: the live bot is never measured under its own name. Every evaluation battle is
a snapshot against something else, because the gate hands snapshot ids to the evaluator. So
the id learner never appears in the evaluation results, and anything looked up under that
id comes back empty.
ladder/score_vs_noop and ladder/score_vs_random_legal appear on probe iterations, and
nowhere else. Until commit 72c1215 on 2026-09-22 they were published as 0.5 in every row,
which was indistinguishable from a genuine even contest; then they were omitted, because the
pair they asked about had no games and never could. Measured 2026-09-22 across the metrics rows
in runs/: 127 rows written before that fix carry a flat 0.5. If you are reading an older run's
file, ignore those two columns entirely.
The decision that was open is now taken. Set ladder.probe_every_iterations and every
ladder.probe_every_iterations iterations the run plays the LIVE bot against each rung in
ladder.probe_opponents and publishes, for each of them:
| key | what it is |
|---|---|
ladder/score_vs/{rung} |
the live bot's score against that rung, 1.0 for a win |
ladder/score_vs_n/{rung} |
how many SEEDS it was measured over, each played from both sides |
ladder/score_vs_ci95_lo/{rung}, ..._hi/{rung} |
the interval around it |
ladder/score_vs_noop and ladder/score_vs_random_legal are the same numbers under their old
names, for the two rungs that have always been the anchors. It is off by default, because it
costs battles nobody was paying for: at the shipped 40 battles a rung it is under 4% of one
gate. Read the interval before the score. At 20 seeds the instrument's own repeat spread against
scripted:random_legal is about 18 points with the policy unchanged, so a smaller move than that
is the instrument. scripted:noop does not act, so it does not have that noise.
ladder/rating_above_v0 was minus the first snapshot's rating, and is now absent instead.
It is meant to be the live bot's rating above the run's first snapshot. The live bot has no
fitted rating, because every evaluation game is a snapshot against something, so the lookup
returned a default of 0.0 and the row published 0 - rating[v0]. On one run that read -93.9,
-146.2, -191.7 and -129.6, which looks exactly like a bot falling behind its own opening
snapshot and is nothing of the kind. Since 2026-09-22 the key is published only when the fit
holds both sides (royalelearn/metrics/records.py, the guard above
fields["ladder/rating_above_v0"]), so it is simply absent until the live bot is evaluated
under its own id. If you are reading a run's file from before that, ignore the column. To see
the old shape for yourself:
python -c "import json,glob; [print(r.get('ladder/rating_above_v0'), r['ladder/rating/snap:v0']) for p in glob.glob('runs/*/metrics.jsonl') for r in map(json.loads, open(p,encoding='utf-8')) if 'ladder/rating/snap:v0' in r]"
Read ladder/rating/<member> directly instead. Those columns are correct.
ladder/eval_games_total was structurally zero and has just been fixed. It used to read a
counter that nothing ever incremented, because the only caller of the counting method hard
codes the training label. The same commit changed it to count the evaluation log directly. It
is still 0 in every metrics row on disk, including runs made after the fix, but that is now
the truthful answer: those runs never played an evaluation battle. Count the log yourself:
python -c "import json,sys; print(sum(1 for l in open(sys.argv[1],encoding='utf-8') if l.strip() and json.loads(l)['kind']=='eval'))" runs/<your-run>/ladder/games.jsonl
What has actually been run, and what has not¶
This matters more than any of the above.
The shipped gate has only just been run. Measured 2026-09-22: of the 50 run folders under
runs/ with a metrics file, 24 ran at least one gate, and every one used champion_games: 2, which
is the smoke setting. After that, a run used the same 1,000-battle gate that laptop.json and
workstation.json ship. Reading its first real gate showed that conditions 2 and 3 could not fail
there: the champion had no record against the anchors, and the pool held nothing to sample. Both
were fixed on 2026-09-24. No numbers from a real gate are written up on this page yet.
The ladder has never had more than a handful of members. On 2026-09-22 the largest pool in any run on disk was 3, which is 2 anchors plus 1 snapshot. Eviction, the stratified sample in condition 3, and the anti-forgetting argument for keeping old snapshots are all untested at the pool sizes they were designed for.
What a gate costs is measured per battle, not per gate. On 2026-09-23 one evaluation battle
took 8.02 s network against network and 0.195 s scripted against scripted on this laptop, the
median of three full matches (EVAL_BATTLE_SECONDS in ladder/gate.py). So a full 2,200-battle
gate is hours, and preflight prints a gate's battles and hours at that rate before a run starts
(rollout/preflight.py). ladder/gate_seconds_frac tells you what gating is costing a run, and
the schema flags it above 0.1.
What is measured, on a 4-core laptop with an RTX 3050 4 GB and 7.8 GB of RAM shared with six other jobs, on 2026-09-22: in a diagnostic run at 2 workers by 24 battles and 8,192 timesteps per iteration with minibatch 256, an iteration took 30 to 35 seconds. About 25 seconds of that was the learning update and 5 to 8 seconds was collecting battles, with the update taking 71 to 83% of every iteration across 7 iterations. That geometry is not the laptop profile, which uses 3 workers by 32 battles and 32,768 timesteps. It does say that on this machine training time is GPU time rather than simulator time.
Where to go next¶
- The dense version of all of this:
docs/harness-spec.md, section 11. - Why the harness is shaped this way:
docs/design.md. - A single head-to-head between two members, outside the training loop:
python -m royalelearn eval --run runs/<your-run> --a snap:v3 --b snap:v0 --seeds 50 - The raw evidence behind every rating:
runs/<your-run>/ladder/games.jsonl, one JSON line per battle.