The harness, specified¶
This page is for the person writing or reading the harness code. It says what is built, field by
field, so that the code can be written from it without further design decisions. Its companion is
docs/design.md, which says what the harness is for and which commitments it is held to.
Every number here is either [M] measured with its source named, or [A] derived by arithmetic from measured numbers, with the arithmetic shown. Nothing here is a menu: where two defensible choices existed, one was taken and the reason is given.
Where this page says the references, it means existing open-source PPO learners for other games. Their decisions are copied where they are right, and named where they are not.
Terms, fixed once and used throughout:
| term | meaning |
|---|---|
| tick | 50 ms of game time (calibration.json time.TICK_MS, status measured) |
| decision / game-step | one ClashParallelEnv.step, decision_ms of game time rounded up to whole ticks; 500 ms = 10 ticks by default |
| timestep | one stored learner transition: one seat's decision |
| slot | one seat of one battle, held for the whole run; R slots in total |
| cycle | one advance of every slot by one decision; the first index of the experience buffer |
| iteration | T cycles, then one PPO update |
| env step | a game-step, counted for cadences (snapshots, checkpoints) |
1. The decisions everything else follows from¶
D1. The iteration is a rectangle: T cycles × R slots. Slot → (worker, shard, battle, seat)
is fixed for the run. The buffer is a plain two-dimensional array with no ring arithmetic, and GAE is
one vectorised backward scan. The reason that matters most comes third: the composition and the row
order of every inference batch are functions of the slot index rather than of which worker answered
first. Determinism does not depend on arrival order.
D2. A worker owns a vec env of many battles; the boundary carries one packed slab per worker per
shard-round. K worker processes each holding ClashSelfPlayVecEnv(num_games=M), not K×M
processes holding one battle each. RoyaleGym already batches battles into one step and the engine
under it is Rust, so there is no reason to pay K×M interpreters to use three cores.
D3. The process boundary is a byte layout, not a Python protocol. rollout/layout.py defines
offsets, dtypes and a LAYOUT_VERSION; both sides compute the same offsets from the same EnvSpec.
docs/design.md commits to rollout workers moving to Rust when they become the bottleneck. A Rust
worker that writes these bytes and flips these events is a drop-in; the Python worker becomes the
reference implementation and the differential test against it is the acceptance criterion.
D4. Observations are quantised in the worker and written straight into their final resting place in the experience buffer. 14 163 B per row against 55 397 B as float32 (95 cards, 2026-09-22) [A], section 7.2. The learner never copies an observation on the CPU except into a small pinned staging ring.
D5. The policy acts on decode(encode(obs)), never on the raw float32. The bytes that are stored
are exactly the input that produced the stored log-prob, which makes ratio == 1 at the first
minibatch of the first epoch an invariant the update can assert (section 9.6). It is the cheapest detector
there is for a mask, codec or weight-version drift, and every one of those failures is otherwise
silent.
D6. The mask travels with the transition, bit-packed, and the same mask is applied at update.
289 bytes of the 14,163-byte row computed in section 7.2, so 2.0%. Both numbers are on this line
because the percentage is derived from the other one: it read 2.2% until 2026-09-23, which is
289/13,136 and a row this codec has not produced since the mask planes were folded in. A rollout/update mask disagreement pins the clip fraction at 1.0 in one
direction and produces log π = -inf in the other; storing the mask makes both impossible.
D7. Inference runs in the parent, one model, one CUDA context, one batched forward per distinct
policy per round. A 4 GB GPU holds one set of weights comfortably and K copies badly, and a
worker that never imports torch costs ~300 MB less resident memory and three seconds less start-up.
A test asserts "torch" not in sys.modules inside a worker.
D8. Every random stream is name-addressed. derive_seedseq(master_seed, "ppo/minibatch/epoch/1"),
never positional spawning. Adding a new consumer of randomness never perturbs an existing stream, so
a config change that ought to be irrelevant is irrelevant. Action sampling takes caller-supplied
uniforms (section 8.3), so a sampled action is a pure function of (master_seed, iteration, cycle, slot)
and is independent of batch composition.
D9. Learner determinism is on by default. determinism.tier = "run_exact". It costs 10-20% of
throughput [A], and a result that cannot be reproduced costs more than that. The tier is one
config field, it is part of the run identity, and "throughput" is the other value.
D10. Evaluation is structurally separate from training. Its own env spec, its own worker farm built at gate time and closed afterwards, its own fixed seed set, its own result table, no transitions recorded. It is not a percentage of the training mixture.
D11. The authoritative rating is a batch fit over an append-only result log, not an online filter. Bradley-Terry with a Davidson draw term, MAP with a Gaussian prior, standard errors from the inverse observed Fisher information. Same games in, same numbers out, in any order, on any machine. Online Elo stays as the dashboard readout and is never the gate.
D12. Every user-facing behaviour is an ABC in royalelearn/api/ that imports numpy and msgspec and
never torch. Concrete implementations live outside api/ and may import torch. import
royalelearn and royalelearn --help work in an environment without torch, which is the property the
package has today and keeps.
2. Geometry and the throughput budget¶
2.1 Unit costs¶
| quantity | value | source |
|---|---|---|
| engine tick | 50 ms game time | [M] calibration.json time.TICK_MS |
| engine throughput | ~32 000 ticks/s, 20 ticks per call | [M] RoyaleGym docs/architecture.md, 2026-09-21 |
ClashParallelEnv.step through the full Python stack |
826 game-steps/s = 1.21 ms | [M] RoyaleGym tests/test_rust_engine.py throughput report, 2026-09-21 |
| split of that 1.21 ms | 0.31 ms engine (26%), 0.90 ms Python obs + mask for both seats (74%) | [A] from the two rows above |
| one network forward | 378 MFLOP/sample at C=64, N=4, 32×18 board | [A] section 8.1 |
| RTX 3050 Laptop, achieved on convs | ~3 TFLOP/s fp32, ~9 TFLOP/s bf16 | [A] 40% of the 6 / 24 peak |
| observation row, as the environment hands it over | 55 397 B (95 cards, 2026-09-22) | [M] shapes from single_observation_space |
| observation row, this codec | 14 163 B (95 cards, 2026-09-22) | section 7.2 |
Both row figures are arithmetic on the observation space rather than constants, so the next layout
change is a recomputation and not a rewrite. With S spatial planes of H·W cells of which s are
static and f are stored as float16, a vector of width V and an action space of A:
row_bytes = H*W * ( (S - s - f) * 1 + f * 2 ) + 2 * V + ceil(A / 8)
float_row = 4 * ( H*W * S + V ) + A + H*W * hand_size
On the 95-card catalogue of 2026-09-22 (the engine's catalogue has grown since, and the vector with
it; the arithmetic is the example, not a pin), that is S = 20, s = 2 static planes held once
per seat, f = 2 hit-point planes, H·W = 576, V = 1177, A = 2305, hand_size = 4, giving
576*(16 + 2*2) + 2*1177 + 289 = 9 216 + 2 304 + 2 354 + 289 = 14 163 B against
4*(576*20 + 1177) + 2 305 + 2 304 = 55 397 B, 3.91 times smaller [A]. The final term of
float_row is the four mask planes, which the environment emits and the codec never stores: they are
the bit-packed mask reshaped, and the learner reshapes it back at unpack.
2.2 Profiles¶
One config block per machine class; only numbers change.
| field | laptop (default) | workstation | many-core |
|---|---|---|---|
| RAM / VRAM / threads | 8 GB / 4 GB / 8 | 32 GB / 12 GB / 16 | 64 GB+ / 24 GB+ / 64 |
rollout.workers K |
3 | 12 | 24 |
rollout.games_per_worker M |
32 | 64 | 64 |
rollout.shards_per_worker |
2 | 2 | 2 |
| battles G = K·M | 96 | 768 | 1 536 |
| slots R = 2G | 192 | 1 536 | 3 072 |
| learner rows G + mirror battles | 144 | 1 152 | 2 304 |
ppo.timesteps_per_iteration |
32 768 | 131 072 | 262 144 |
| cycles T = ceil(ts / learner rows) | 228 | 114 | 114 |
net.channels C / net.blocks N |
64 / 4 | 96 / 8 | 128 / 12 |
ppo.batch_size |
4 096 | 16 384 | 32 768 |
ppo.minibatch_size |
256 | 2 048 | 8 192 |
| buffer bytes (observations) (95 cards, 2026-09-22) | 623 MB | 2.5 GB | 5.0 GB |
rollout.overlap |
false | false | false |
ladder.candidate_every_env_steps |
4 000 000 | 2 000 000 | 2 000 000 |
shards_per_worker is the intra-worker overlap: each worker holds two independent vec envs of M/2
battles and alternates, so it is stepping one shard while the parent runs inference on the other. A
cycle is shards_per_worker consecutive shard-rounds, so that every slot advances exactly once;
the shard interleaving is a scheduling detail of the parent and never appears in the buffer's index.
What a second shard is worth follows from one measured fact: environment time and inference time
are additive, with no overlap of their own. Standing a fixed 5 ms cost after every vec step adds
5.1 to 5.7 ms to the step at every batch size, on both engines [M]. So the second shard hides
min(env, inference) and nothing else, and its value is set by the ratio of the two per shard-round:
| env ms per shard ÷ inference ms per round | what a second shard buys |
|---|---|
| well under 1 (few battles per worker, inference dominates) | close to double |
| about 1 | most of the inference |
| well over 1 (many battles per worker, the environment dominates) | little; at a ratio near 7, about 14% [M] |
That inverts the obvious intuition: sharding is worth most where battles per worker is small,
because a worker with many battles has already amortised inference by batching. At the laptop
profile a shard is 16 battles against three batched forwards, which sits in the middle of that table.
One thing keeps the number honest rather than assumed: a fixed sleep releases the interpreter cleanly
while a real forward competes for cores, so the table is the optimistic case. What is NOT built, though
this section claimed it was until 2026-09-22: royalelearn bench does not print what a second shard is
worth. Nothing in it separates the shards. The measurement is two runs of bench on a quiet machine
with rollout.shards_per_worker at 1 and at 2, compared on collection seconds; if the second shard
buys under about 10%, set it to 1 and spend the thread on a worker instead.
The buffer row is (T + obs.frame_stack) × R × row_bytes: T+1 cycles for the bootstrap row plus
frame_stack - 1 history cycles carried over from the previous iteration (section 9.1). At the laptop
profile that is 229 × 192 × 14 163 B = 623 MB (95 cards, 2026-09-22) [A]. The two larger profiles'
figures are 115 × 1 536 × 14 163 B and 115 × 3 072 × 14 163 B. They were twice that while they
asked for rollout.overlap, which bought a second rectangle; start-up has refused it since
2026-09-22, because nothing honours it.
2.3 The laptop budget¶
Per iteration at the laptop profile: T=228 cycles × R=192 slots, of which 144 are learner rows,
so 228 × 144 = 32 832 timesteps and 228 × 96 = 21 888 game-steps.
| phase | seconds | basis |
|---|---|---|
| engine ticks + Python obs and mask | 8.8 | 21 888 game-steps ÷ (3 × 826 game-steps/s) [M] |
| rollout inference (actor only) | ~4 | 456 shard-rounds × ~3 forwards × ~1.5 ms launch-bound, plus 16.5 TFLOP [A] |
| critic pass, chunks of 1024 | 1.4 | 32 976 rows × 378 MFLOP ÷ 9 TFLOP/s [A] |
| PPO update: 3 epochs × (fwd+bwd ≈ 3×fwd) × 2 nets | 24.8 | 223 TFLOP ÷ 9 TFLOP/s [A] |
| GAE, advantage standardisation, H2D copies | ~2 | [A] |
| total, serial | ~41 | ≈ 800 timesteps/s |
At determinism.tier = "run_exact" subtract 10-20%: ~680-720 timesteps/s, so 100 M timesteps is
39-41 hours. Overlapped collection would hide the rollout under the update, for
~1 100 timesteps/s and 25 hours, at the cost of a second buffer (623 MB, 95 cards). It is not built:
start-up refuses rollout.overlap = true.
Two consequences, stated so nobody re-derives them:
- The learner is the bottleneck on this machine, by about 2.5×. Rollout capacity is ~3 300
timesteps/s against an update capacity of ~1 330. That inverts the usual situation and it is
why
K = 3and not 32, and why the network is 430 k parameters and not 4 M. - "The harness is never the bottleneck" is an invariant with a number: rollout capacity ≥ 2 ×
update capacity on every shipped profile.
royalelearn benchprints both, the run logsthroughput/rollout_capacity_ratioevery iteration, and start-up warns below 1.5.
RoyaleGym's planned Rust-backed default observation and mask would take the rollout side from
298 µs/transition to about 110 µs [A], i.e. 9 000 timesteps/s. It is not a prerequisite: it is
what later lets K drop to one or two and frees cores.
What the shipped schedules assume about the clock, measured rather than budgeted. The table
above is the design's estimate; the machine gives about 295 timesteps/s at minibatch 256 and about
304 env steps/s [M], so an iteration of 32 832 timesteps takes close to two minutes rather than
41 seconds. Every schedule in laptop.json is written in env steps, and at that rate they are:
| schedule | over | at 304 env steps/s |
|---|---|---|
ppo.ent_coef 0.01 to 0.003 |
30 M env steps | about 27 hours |
ppo.ent_coef_noop 0.02 to 0 |
10 M env steps | about 9 hours |
advantage.gamma 0.997 to 0.999 |
20 M env steps | about 18 hours |
ladder.candidate_every_env_steps |
4 M env steps | about 3.7 hours to the first gate |
So the profile named for this laptop assumes a run of a day or more, and a run stopped after a few hours has not seen its own entropy schedule. That is a statement about the schedules and not about the hardware: anyone judging a short run should read where on these curves it stopped, and anyone shortening a run should shorten the schedules with it rather than leave them and read the result as a policy that would not settle.
2.4 The RAM ledger, printed by royalelearn doctor¶
experience buffer (pinned) 229 x 192 x 14 163 B 623 MB (95 cards, 2026-09-22)
parent torch + CUDA host allocations ~1 400 MB
3 workers x (32 RustEngine battles, no torch) 3 x ~250 = 750 MB
3 workers x shard staging and scalars 3 x ~40 = 120 MB
python interpreters, numpy, msgspec ~400 MB
----------
~3.3 GB of 7.8 GB
at gate time, additionally: 2 eval workers x 24 battles ~380 MB
doctor refuses to start a run whose projection exceeds doctor.ram_budget_mb (default 6 500).
3. The file tree¶
RoyaleLearn/
README.md
LICENSE MIT
pyproject.toml deps: royalegym, numpy, msgspec; extras torch, wandb, dev
docs/
design.md the commitments; unchanged in intent
harness-spec.md this file
throughput.md the budget of section 2 and how to re-measure it
determinism.md the contract of section 5
ladder.md rating, matchmaking, the gate, the statistics
checkpoints.md the format, versioning, what resume proves
metrics.md every metric, its range, its alarm
running.md first run, reading the dashboard, one section per alarm
royalegym-asks.md section 16, with costs
media/ unchanged
examples/
train_1v1.py the file a user runs and edits
configs/laptop.json the shipped default, written out in full
configs/workstation.json the scale-up profile
configs/smoke.json MockEngine, 2 workers x 2 battles, 3 iterations
custom_reward.py how to swap a RewardFunction and keep the identity honest
royalelearn/
__init__.py version and lazy exports; imports nothing that needs torch
version.py __version__, git describe helper
errors.py PreflightError, IdentityMismatch, CheckpointFormatError,
WorkerTimeout, AlarmHalt, StaleEngineBuild
config.py the whole msgspec config tree, hashing, JSON I/O
seeding.py derive_seedseq / derive_generator / derive_int, the namespace
determinism.py apply(tier), the env-var preconditions, the assertions
identity.py RunIdentity, EngineBuild, compute_identity(config)
obs_layout.py vector fields resolved by name from the env's vector_layout()
coordinator.py LearningCoordinator: the loop, the only place phases are ordered
cli.py argument parsing for python -m royalelearn
__main__.py entry point -> cli.main()
rewards.py the potential-based default reward composition
checkpoint.py DirCheckpointStore: manifest, per-component folders, pruning
api/
__init__.py re-exports every ABC and struct
rollout.py EnvSpec, SlotPlan, RolloutRound, WorkerCommand, WorkerFailure,
EpisodeRecord, RolloutSource
policy.py ActionDistribution, Actor, Critic, ActorCritic, NetworkFactory
buffer.py ObsCodec, ExperienceBuffer
advantage.py AdvantageEstimator
update.py Update, UpdateResult
ladder.py Matchmaker, Rater, PromotionGate, SnapshotStore, EvictionPolicy
metrics.py MetricsSink, MetricRow, Alarm, AlarmResult
checkpoint.py Checkpointable, CheckpointStore, Manifest
schedule.py Schedule
rollout/
__init__.py
layout.py the shared-memory byte layout; the cross-language contract
codec.py SpatialObsCodec: pack, unpack_to_device, the quantisation table
envspec.py EnvFactorySpec, ComponentSpec, the class allow-list
preflight.py engine construction, mask_disagreements, digests, the RAM ledger
plan.py SlotPlanner: geometry, slot maps, the per-slot seed paths
scripted.py worker-side scripted opponents: the two anchors and four strategies
worker.py child process main loop; numpy only, torch is a test failure
farm.py ProcessRolloutSource: spawn, handshake, rounds, failure, restart
inline.py InlineRolloutSource: same semantics, one process, no shm
learn/
__init__.py
nets.py ClashTrunk, PointerPolicyHead, ValueHead, DefaultNetworkFactory
distribution.py MaskedCategorical
actor_critic.py SeparateActorCritic, BehaviourSnapshot
buffer.py RectBuffer: the (T+1, R) shared-memory-backed rectangle
gae.py GAE (torch, vectorised) and reference_gae (pure python)
returns.py WelfordReturnScaler
ppo.py PPOUpdate: the losses, the minibatch loop, the diagnostics
inference.py BatchedInference: per-round forwards, streams, staging ring
schedules.py Constant, Linear, Geometric, PiecewiseConstant, the lr backoff
ladder/
__init__.py
snapshots.py DiskSnapshotStore: fp16 actor + spec.json + digest, LRU
matchmaker.py MixMatchmaker: the mixture, PFSP from the fitted ratings
rating.py BradleyTerryDavidsonRater, EloReadout, transitivity_residual
results.py ResultLog: append-only games.jsonl plus the aggregate cache
evaluate.py EvalRunner: fixed seed set, paired sides, no experience
gate.py WilsonGate: the three conditions and the decision record
eviction.py HallOfFameEviction
pool.py LadderPool: royalegym OpponentPool + ResultLog + snapshots
metrics/
__init__.py
schema.py every metric key, unit, description and healthy range
records.py IterationMetrics and the per-source dataclasses
sinks.py JsonlSink (always on), ConsoleSink, CompositeSink
viser_sink.py ViserSink: the learning-status datagram the viewer reads
wandb_sink.py WandbSink: a decorator, run-id persistence, optional import
alarms.py the alarm table of section 13.3 and its evaluator
bundle.py write_bundle: the diagnostic directory written on any halt
tests/ section 15
Deleted in the first commit: continuous_policy.py, discrete_policy.py, multi_discrete_policy.py,
value_estimator.py, experience_buffer.py, ppo_learner.py, the SEED_CLASSES machinery in
__init__.py, NOTICE, LICENSE-APACHE-2.0, the [tool.ruff] extend-exclude list, and the seed
assertions in tests/test_package.py. pyproject.toml's license becomes { text = "MIT" }.
docs/design.md loses its section on the seed modules and gains one sentence recording that
they were a reference and were removed when the harness landed.
4. The ABCs¶
All in royalelearn/api/. They import numpy, typing and msgspec only. torch appears in type
annotations behind if TYPE_CHECKING. Every ABC that owns state also implements
api.checkpoint.Checkpointable.
4.1 api/rollout.py¶
class ObsKeySpec(msgspec.Struct, frozen=True):
"""One key of the environment's observation space, read from the space itself."""
shape: tuple[int, ...]
dtype: str # numpy dtype name
low: tuple[float, ...] # per channel for a 3-D key, length 1 otherwise
high: tuple[float, ...] # per channel for a 3-D key, length 1 otherwise
class EnvSpec(msgspec.Struct, frozen=True):
"""Everything the learner must know about the environment before it builds anything.
Every field here is READ from the running environment at preflight. The harness holds no
layout constant of its own: a width, a plane count and a field offset are all environment
facts, and typing one into this repo would make a RoyaleGym change a silent wrong answer
instead of a loud one."""
num_cards: int # the card catalogue this run is built on
obs_space: dict[str, ObsKeySpec] # every key of single_observation_space, verbatim
vector_layout: tuple[tuple[str, int, int], ...] # (name, offset, size), from ObsBuilder
spatial_layout: tuple[tuple[str, bool], ...] # (name, static), from ObsBuilder
frame_stack: int # obs.frame_stack; how many cycles the trunk sees
vector_size: int # obs_space["vector"].shape[0]
spatial_shape: tuple[int, int, int] # obs_space["spatial"].shape
n_actions: int # single_action_space.n
hand_size: int
tiles: tuple[int, int] # (tiles_y, tiles_x)
decision_ms: int # 500
tick_ms: int # 50
decision_ticks: int # 10
regular_ticks: int # 3600
overtime_ticks: int # 1200
engine_build_digest: str # sha256, section 5.3
obs_digest: str # sha256 over obs_space, vector_layout, spatial_layout,
# frame_stack, n_actions, num_cards and the obs-builder
# ComponentSpec -- so a Reveal change changes it
env_factory: "EnvFactorySpec"
vector_size and spatial_shape are conveniences derived from obs_space and are there so that the
common expressions read well; the space is the authority and a mismatch between the two is an
assertion at construction. The observation carries no velocity, acceleration or difference channel:
motion reaches the policy by frame stacking (obs.frame_stack, section 9.1), which is the harness's
own affair and costs the environment nothing.
Superhuman information, meaning anything a human player at the same moment could not see, is enabled
per field through the obs builder's frozen Reveal struct, which is part of
ClashParallelEnv.config(). One known gap (2026-09-27): an enemy Royal Ghost still shows in the
default observation while it is invisible. This is not fixed yet.
An enabled field adds vector slots, and enemy_spell_aim adds a spatial plane, so both widths and the
plane count move with it. Every one of those changes arrives through vector_layout,
spatial_layout and the observation space like any other layout fact, and through obs_digest like
any other component setting, so a policy that sees revealed information can never be rated against
one that does not (section 11.6).
class Assignment(msgspec.Struct, frozen=True):
"""What one battle's next episode is. Drawn by the Matchmaker at the battle's own episode
boundary, from match/battle/{b}/ordinal/{k}, and therefore a pure function of the master
seed, the battle index and the reset ordinal."""
battle: int
ordinal: int # reset ordinal this assignment takes effect on
role: int # 0 mirror, 1 pool, 2 scripted
opponent_id: str | None # "snap:<id>", "scripted:<name>", or None for mirror
group: tuple[int, int] # per seat: -1 learner, 0..P-1 resident snapshot,
# -2 scripted (worker-side), -3 dead
learner_seat: int # 0 blue, 1 red; -1 for mirror (both)
class SlotPlan(msgspec.Struct, frozen=True):
"""The iteration's opening assignment table: what every battle is playing at the moment the
iteration begins. Assignments change only at a battle's own episode boundary, where the
parent draws a fresh one; nothing in an iteration's span reassigns a battle mid-episode, so
a partially controlled trajectory cannot occur."""
iteration: int
n_battles: int
n_slots: int # 2 * n_battles
assignment: tuple[Assignment, ...] # one per battle, in battle order
resident_snapshots: tuple[str, ...] # len <= ladder.max_resident_opponents
class RolloutRound:
"""One shard-round of observations. Holds zero-copy views; not a Struct."""
cycle: int # buffer row index
shard: int
slots: np.ndarray # int32[n], the slot indices this round covers, ascending
obs_rows: np.ndarray # int32[n], where in the buffer each slot's row was written
group: np.ndarray # int8[n], as reported by the worker for THIS round
reward: np.ndarray # float32[n]
terminated: np.ndarray # bool[n]
truncated: np.ndarray # bool[n]
valid: np.ndarray # bool[n]; False where that slot's worker is dead
deploy_status: np.ndarray # int8[n]; -1 no command, 0 OK, 1..11 = a MASK BUG
tick: np.ndarray # int32[n]
episode_end: np.ndarray # int8[n]; 0 none, 1 win, 2 loss, 3 draw
episodes: list["EpisodeRecord"] # one per episode that ended on this round
timings: dict[str, float]
class EpisodeRecord(msgspec.Struct, frozen=True):
slot: int
worker: int
shard: int
battle: int
seat: int # 0 blue, 1 red
ordinal: int # reset ordinal of this episode within its battle
episode_seed_path: str
policy_id: str # "learner" | "snap:<id>" | "scripted:<name>"
opponent_id: str
bucket: str # "mirror" | "pool" | "scripted"
# the terminal scalars, copied straight out of the env's final_info; the environment's
# EPISODE_STAT_KEYS is the authority on the set, and a key it adds arrives here unchanged
episode_steps: int
episode_ticks: int
own_crowns: int
enemy_crowns: int
own_tower_hp_frac: float # mean of the three towers' hp/max
enemy_tower_hp_frac: float # mean of the three towers' hp/max
elixir_leak_steps: int # decisions spent at full elixir
elixir_count_exact: bool # False if this seat's count of the opponent's elixir was
# ever caught disagreeing with the bar it models, so the
# policy read an estimate in a slot documented as exact
# counted by the worker as the episode runs
winner: int # -1 none, 0 blue, 1 red, 2 draw
outcome: int # +1 win, 0 draw, -1 loss, from this seat's view
cards_played: int
illegal_commands: int # deploy_status in 1..11; must be 0
undiscounted_return: float
reward_terms: dict[str, float] # per-term episode sums for this seat
class WorkerCommand(msgspec.Struct, tag_field="kind", tag=str.upper):
"""Tagged union; six variants, one per round."""
class Step(WorkerCommand):
actions: np.ndarray # int16[n_slots_this_shard]
gamma: float # the current schedule value, threaded into the reward
group: np.ndarray # int8[n_slots_this_shard]
opponent_ix: np.ndarray # int8[n_slots_this_shard], index into the plan's
# resident_snapshots / scripted table
learner_seat: np.ndarray # int8[n_battles_this_shard]
# the three arrays carry the parent's assignment for
# every battle whose episode started on the previous
# round, and are empty otherwise; the worker needs them
# only to know which slots it fills with a scripted
# action
class Plan(WorkerCommand):
plan: SlotPlan # the iteration's opening table, once per iteration
class SetState(WorkerCommand):
snapshots: tuple[bytes | None, ...] # per battle; start from a recorded position
class Spaces(WorkerCommand):
pass # re-read the spaces without stepping
class Defer(WorkerCommand):
pass # nothing this round; hand back the same data
class Close(WorkerCommand):
pass
class WorkerFailure(msgspec.Struct, frozen=True):
worker: int
shard: int
cycle: int
kind: str # "exception" | "crash" | "timeout" | "protocol"
message: str # the child's traceback, verbatim
class RolloutSource(ABC):
"""Where experience comes from. The one seam between the learner and the world.
Implementers may assume: `begin_iteration` precedes any `next_round`; exactly one `submit`
per `next_round`; `Step.actions` are legal under the mask that was handed out; the caller
does not retain a RolloutRound's views past the next `next_round` for the same shard.
Implementers MUST guarantee: `slots` is ascending; every slot appears exactly once per
cycle; a dead worker's slots arrive with `valid=False` rather than not arriving."""
@abstractmethod
def spec(self) -> EnvSpec: ...
@abstractmethod
def begin_iteration(self, plan: SlotPlan, buffer: "ExperienceBuffer", iteration: int) -> None: ...
@abstractmethod
def next_round(self, timeout_s: float = 30.0) -> RolloutRound: ...
@abstractmethod
def submit(self, command: WorkerCommand) -> None: ...
@abstractmethod
def drain_failures(self) -> list[WorkerFailure]: ...
@abstractmethod
def restart(self, worker: int) -> None: ...
@abstractmethod
def close(self) -> None: ...
def stats(self) -> dict[str, float]:
"""Per-round timing: env_ms, wait_ms, codec_ms, parent_wait_frac, bytes_out."""
return {}
Shipped: rollout.farm.ProcessRolloutSource (default) and rollout.inline.InlineRolloutSource
(identical semantics in one process, no shared memory). The inline source is the reference
implementation and what the fast tests run. A future RustRolloutSource implements the same ABC and
writes the same bytes.
4.2 api/policy.py¶
ObsBatch = tuple # a NamedTuple of device tensors, shaped from EnvSpec.obs_space:
# spatial (B, k*S, H, W) f32 k = frame_stack, S spatial planes
# mask_planes (B, k*A_s, H, W) f32 A_s per-slot action planes, derived
# vector (B, V) f32
# mask (B, n_actions) bool
class ActionDistribution(ABC):
@abstractmethod
def sample(self, uniforms: Tensor) -> Tensor: ...
"""(B,) int64. Driven by CALLER-SUPPLIED uniforms in [0,1) so that the sampled action is
reproducible independently of batch composition and of torch's global RNG."""
@abstractmethod
def mode(self) -> Tensor: ... # (B,) int64, mask-respecting argmax
@abstractmethod
def log_prob(self, actions: Tensor) -> Tensor: ... # (B,) float32
@abstractmethod
def entropy(self) -> Tensor: ... # (B,) float32
@abstractmethod
def noop_entropy(self) -> Tensor: ... # (B,) float32, H2 of p(no-op) against p(play)
@abstractmethod
def n_legal(self) -> Tensor: ... # (B,) int64
class Actor(ABC, nn.Module):
@abstractmethod
def logits(self, obs: ObsBatch) -> Tensor: ...
"""(B, spec.n_actions) float32 ALWAYS, even inside autocast. Raw logits: no softmax, no
clamp, no mask."""
def distribution(self, obs: ObsBatch) -> ActionDistribution:
return MaskedCategorical(self.logits(obs), obs.mask)
class Critic(ABC, nn.Module):
@abstractmethod
def value(self, obs: ObsBatch) -> Tensor: ... # (B,) float32
class ActorCritic(ABC, nn.Module):
"""Implementers may assume obs tensors are on the device and already dequantised, and that
obs.mask[:, 0] is True on every row (RoyaleGym sets mask[NOOP]=1 unconditionally,
action.py:284, including after game over). They MUST assert it. Inside the update's
minibatch loop, outside debug iterations, the harness has already asserted it for every
trainable cell, in the critic pass (section 18.1, "How a row is classed, and what checks
it"), and MaskedCategorical skips its own check there. The rollout's forwards skip it too:
they read the no-op bit back in the same copy as the actions, and a row without it stops
the round there."""
arch: "ArchSpec"
@abstractmethod
def act(self, obs: ObsBatch, uniforms: Tensor) -> "ActResult": ...
@abstractmethod
def value(self, obs: ObsBatch) -> Tensor: ...
@abstractmethod
def backprop(self, obs: ObsBatch, actions: Tensor) -> "BackpropResult": ...
"""Recompute under the current parameters, applying exactly the mask read from the
buffer -- never a recomputed one."""
@abstractmethod
def actor_parameters(self) -> Iterator[nn.Parameter]: ...
@abstractmethod
def critic_parameters(self) -> Iterator[nn.Parameter]: ...
@abstractmethod
def actor_state_dict_fp16(self) -> dict[str, Tensor]: ...
"""What a pool snapshot stores: actor only, fp16, no optimizer."""
class ActResult(msgspec.Struct):
actions: Tensor # (B,) int64
log_probs: Tensor # (B,) float32
entropy: Tensor # (B,) float32, diagnostics only
p_noop: Tensor # (B,) float32, diagnostics
n_legal: Tensor # (B,) int64, diagnostics
class BackpropResult(msgspec.Struct):
log_probs: Tensor # (B,) float32
entropy: Tensor # (B,) float32
noop_entropy: Tensor # (B,) float32
values: Tensor # (B,) float32
n_legal: Tensor # (B,) int64
class NetworkFactory(ABC):
@abstractmethod
def build(self, spec: EnvSpec, arch: "ArchSpec", device, dtype) -> ActorCritic: ...
@abstractmethod
def arch_digest(self, spec: EnvSpec, arch: "ArchSpec") -> str: ...
"""sha256 of the canonical JSON of (arch, obs_space, frame_stack, num_cards, n_actions).
A snapshot or checkpoint built from a different architecture is refused, not
shape-errored."""
Shipped: learn.nets.DefaultNetworkFactory building SeparateActorCritic(ClashTrunk +
PointerPolicyHead, ClashTrunk + ValueHead). SharedTrunkActorCritic ships too, because
docs/design.md requires a swappable alternative and because it is the escape hatch if rollout
inference turns out to cost more than a quarter of the iteration.
4.3 api/buffer.py¶
class CodecTable(msgspec.Struct, frozen=True):
"""How each observation key is stored, decided from EnvSpec.obs_space at preflight rather
than from a list of plane indices. Logged, hashed and written into every snapshot."""
plane: tuple[tuple[str, str, float], ...] # per spatial plane: (name, storage, divisor);
# storage in {"uint8", "float16", "static",
# "derived"}; "derived" is reserved,
# and bind refuses it: nothing rebuilds one yet
vector: str # "float16"
mask: str # "bitpack"
def digest(self) -> str: ... # sha256 of the canonical JSON
class ObsCodec(ABC):
"""Quantisation of one observation row. The worker packs; the learner unpacks on the GPU.
Implementers MUST be exact round-trips for the integer-valued channels and MUST declare
row_bytes as a constant given an EnvSpec and its CodecTable."""
@abstractmethod
def table(self, spec: EnvSpec, sample: Sequence[dict[str, np.ndarray]]) -> CodecTable: ...
"""Decide storage per key from the declared bounds and a sample of real observations."""
@abstractmethod
def row_bytes(self, spec: EnvSpec) -> int: ...
@abstractmethod
def pack(self, obs: dict[str, np.ndarray], out: memoryview, row: int) -> None: ...
@abstractmethod
def static_planes(self, obs: dict[str, np.ndarray]) -> np.ndarray: ...
"""The planes EnvSpec.spatial_layout declares static; stored once per seat, never per
row. Declared, never inferred: see section 7.2."""
@abstractmethod
def unpack_to_device(self, raw: Tensor, statics: Tensor, out: ObsBatch) -> None: ...
"""Dequantise, scatter the static planes in, reshape the stored mask into the mask
planes, and gather the frame-stack history."""
@property
@abstractmethod
def codec_version(self) -> int: ...
class ExperienceBuffer(ABC, Checkpointable):
"""A rectangle of (T + frame_stack) cycles x R slots: T collected cycles, one bootstrap
row, and frame_stack - 1 history rows carried over from the previous iteration. Owns the
shared-memory block the workers write into. Implementers may assume each (cycle, slot) cell
is written exactly once, by the worker that owns that slot; the learner only reads."""
@abstractmethod
def shared_handle(self) -> "BufferHandle": ... # name, size, offsets; picklable
@abstractmethod
def begin_iteration(self, plan: SlotPlan, cycles: int) -> None: ...
@abstractmethod
def record_round(self, r: RolloutRound, actions: np.ndarray,
log_probs: np.ndarray) -> None: ...
"""Scalars only: observations are already in place. O(n), no observation copy."""
@abstractmethod
def set_values(self, values: Tensor) -> None: ... # (T+1, R) float32
@abstractmethod
def set_final_values(self, cells: np.ndarray, values: Tensor) -> None: ...
"""V(final_obs) for truncated cells; `cells` is int64[(k, 2)] of (cycle, slot)."""
@abstractmethod
def set_advantages(self, adv: Tensor, ret: Tensor) -> None: ... # (T, R) float32
@abstractmethod
def trainable_mask(self) -> Tensor: ... # (T, R) bool
@abstractmethod
def batches(self, batch_size: int, minibatch_size: int, epochs: int,
rng_for_epoch: Callable[[int], np.random.Generator]) -> Iterator["Batch"]: ...
"""Yields batches; each Batch knows its true sample count and iterates device-resident
minibatches. A batch never straddles an epoch boundary: the epoch's remainder is its own
smaller batch, correctly weighted. Gathers per MINIBATCH, never per batch."""
Shipped: learn.buffer.RectBuffer.
4.4 api/advantage.py, api/update.py, api/schedule.py¶
class AdvantageEstimator(ABC, Checkpointable):
@abstractmethod
def compute(self, *, rewards, values, final_values, terminated, truncated, trainable,
gamma: float, lam: float) -> tuple[Tensor, Tensor, "AdvantageStats"]: ...
"""rewards/terminated/truncated/trainable are (T, R); values is (T+1, R);
final_values is (T, R) and is read only where truncated.
Returns (advantages (T,R), returns (T,R), stats).
Implementers MUST bootstrap a terminated cell from 0 and a truncated cell from
final_values, and MUST NOT carry the recursion across an episode boundary."""
class AdvantageStats(msgspec.Struct):
raw_return_mean: float; raw_return_std: float
reward_scale: float; clipped_reward_frac: float
class Update(ABC, Checkpointable):
@abstractmethod
def step(self, buffer: ExperienceBuffer, sched: "ScheduleState") -> "UpdateResult": ...
class UpdateResult(msgspec.Struct):
policy_loss: float; value_loss: float; entropy: float; noop_entropy: float
entropy_normalised: float; kl: float; clip_fraction: float; dual_clip_fraction: float
explained_variance: float; ratio_max_abs_dev: float
grad_norm_actor: float; grad_norm_critic: float
update_magnitude_actor: float; update_magnitude_critic: float
kl_by_epoch: list[float]; clip_fraction_by_epoch: list[float]
n_minibatches: int; n_optimizer_steps: int; n_samples: int
samples_unused_frac: float; seconds: float
class Schedule(ABC):
@abstractmethod
def value(self, env_steps: int) -> float: ...
Shipped estimators: learn.gae.GAE (torch, vectorised over slots) and learn.gae.reference_gae
(pure python, per slot, used only by the test that proves the vectorised one). Shipped schedules:
Constant(v), Linear(a, b, over_env_steps), Geometric(a, b, over_env_steps) (used for gamma, so
that 1-gamma decays geometrically), PiecewiseConstant(points).
4.5 api/ladder.py¶
class Matchmaker(ABC, Checkpointable):
@abstractmethod
def plan(self, iteration: int, pool: "LadderPool", ratings: "RatingTable",
geometry: "Geometry") -> SlotPlan: ...
"""The iteration's opening table. One Assignment per battle, each drawn at that
battle's current ordinal, so plan() is assign() applied across the geometry."""
@abstractmethod
def assign(self, battle: int, ordinal: int, pool: "LadderPool",
ratings: "RatingTable") -> Assignment: ...
"""The assignment for one battle's next episode. Draws from
match/battle/{battle}/ordinal/{ordinal}, so it is a pure function of the master seed and
its two arguments: the same episode of the same battle always meets the same opponent,
whichever iteration it happens to fall in and whichever worker holds it."""
@abstractmethod
def on_episode(self, record: EpisodeRecord) -> None: ...
class Rater(ABC, Checkpointable):
@abstractmethod
def fit(self, results: "ResultView") -> "RatingTable": ...
"""id -> (rating in Elo units, standard error). MUST be a pure function of `results`:
same games in, same numbers out, in any order, on any machine."""
@abstractmethod
def predict(self, a: str, b: str) -> float: ... # P(a scores against b), draws as 0.5
@abstractmethod
def transitivity_residual(self, results: "ResultView") -> float: ...
class RatingTable(msgspec.Struct, frozen=True):
rating: dict[str, float]; se: dict[str, float]
anchor: str; draw_nu: float | None
n_games: dict[str, int]
transitivity_residual: float; converged: bool; iterations: int
class PromotionGate(ABC):
@abstractmethod
def evaluate(self, candidate: str, pool: "LadderPool",
runner: "EvalRunner") -> "GateDecision": ...
class GateDecision(msgspec.Struct):
candidate: str; champion: str
admit: bool; promote: bool; cycle: bool
conditions: dict[str, "ConditionResult"]
eval_seed_set_sha: str; wall_seconds: float
class ConditionResult(msgspec.Struct):
passed: bool; n: int; observed: float; bound: float; reference: float
class SnapshotStore(ABC):
@abstractmethod
def put(self, snapshot_id: str, ac: ActorCritic, meta: dict) -> str: ...
@abstractmethod
def get(self, snapshot_id: str, device) -> Actor: ... # LRU-cached
@abstractmethod
def digest(self, snapshot_id: str) -> str: ...
@abstractmethod
def list(self) -> list[str]: ...
class EvictionPolicy(ABC):
@abstractmethod
def select_for_eviction(self, *, pool: "LadderPool", ratings: RatingTable,
max_sampled: int) -> list[str]: ...
"""Ids to remove FROM THE SAMPLER. The archive and the result log are never touched:
eviction is about sampling cost, not about forgetting evidence."""
Shipped: MixMatchmaker, BradleyTerryDavidsonRater (authoritative) and EloReadout (dashboard),
WilsonGate, DiskSnapshotStore, HallOfFameEviction.
4.6 api/metrics.py and api/checkpoint.py¶
class MetricsSink(ABC, Checkpointable):
@abstractmethod
def open(self, *, identity: "RunIdentity", config_json: str, run_dir: Path) -> None: ...
@abstractmethod
def write(self, row: "MetricRow") -> None: ...
"""One flat dict per iteration. Implementers MUST NOT mutate the row and MUST NOT raise
on an unknown key."""
def write_episodes(self, rows: Sequence[EpisodeRecord]) -> None: ... # default: ignore
def write_alarms(self, alarms: Sequence["AlarmResult"]) -> None: ... # default: ignore
def write_artifact(self, name: str, path: Path) -> None: ... # default: ignore
@abstractmethod
def close(self) -> None: ...
class Checkpointable(Protocol):
FORMAT_VERSION: int
def save_checkpoint(self, folder: Path) -> None: ...
def load_checkpoint(self, folder: Path, *, strict: bool) -> None: ...
"""With strict=False, a missing file prints the exact path it wanted and continues with a
default. With strict=True it raises. strict defaults to True on resume."""
class CheckpointStore(ABC):
@abstractmethod
def write(self, components: Mapping[str, Checkpointable], manifest: "Manifest") -> Path: ...
@abstractmethod
def read(self, path: Path, components: Mapping[str, Checkpointable], *,
strict: bool) -> "Manifest": ...
@abstractmethod
def latest(self, run_dir: Path) -> Path | None: ...
@abstractmethod
def prune(self, run_dir: Path, keep: int) -> list[Path]: ...
Shipped sinks: JsonlSink (always installed and never optional, because it is the file the resume
test compares), ConsoleSink, CompositeSink, WandbSink (a decorator over any sink). Shipped
store: DirCheckpointStore.
5. Determinism¶
5.1 Three tiers, named and recorded¶
| tier | guarantee | cost | when |
|---|---|---|---|
| T1 env-exact | given the run identity, every episode's engine state-hash sequence, every observation, every mask and every reward are bit-identical, on any machine, forever | free | always, unconditionally |
| T2 run-exact | T1, plus every gradient, every parameter and every metric row bit-identical on the same device class and torch/CUDA build | 10-20% throughput [A] | default, determinism.tier = "run_exact" |
| T3 throughput | T1 only; the learner may use nondeterministic kernels and curves agree in distribution | fastest | opt-in, tier = "throughput", stamped on every metric row and every checkpoint |
tier is part of the run identity. A T3 run may not be resumed into a T2 run without
--allow-identity-drift, which writes drift.json naming every differing field and marks every
later metric row resumed_with_drift=true.
T2 keeps bf16 autocast. Deterministic kernels are bit-reproducible run to run at any precision, so the tier costs nothing in precision and buys exactly what it says: two runs of the same identity agree row for row. It is not a claim that two differently shaped forwards within one run agree, which is a separate matter handled in sections 8.2 and 9.6.
determinism.apply(tier) for T2:
torch.use_deterministic_algorithms(True, warn_only=False)
torch.backends.cudnn.deterministic = True
torch.backends.cudnn.benchmark = False
torch.backends.cuda.matmul.allow_tf32 = False # tf32 is not bit-reproducible across shapes
torch.backends.cudnn.allow_tf32 = False
torch.set_num_threads(config.determinism.torch_threads) # default 1
The fill of memory nothing has written. With deterministic algorithms on, torch also fills what
it allocates without values -- torch.empty, empty_like, new_empty, resize_ and the ops that
allocate through them -- before handing it over: NaN for floating types, the largest value for
integers, and True for bool (torch.utils.deterministic.fill_uninitialized_memory, on by default).
It is a guard, not part of the arithmetic. For floats it turns a read of memory nothing wrote into a
NaN the same in every run; for bool and uint8 it is not loud at all (an unwritten mask would read as
every action legal). A program that reads only what it wrote computes the same bits with it off. It
is also an extra full write of every such buffer. It was timed once: on 2026-09-26, on an RTX 4070
Ti, on one run shape (48 battles, 114 cycles, 8,208 learner rows an iteration) and on the learner
as of e437ffa, before the update's per-minibatch host synchronisations were cut. Two updates were
timed with the fill on, then two with it off. With every row running the actor: on 9.03 and
9.07 s, off 8.44 and 8.14 s, so the means differ by 0.76 s. With forced rows skipping it
(critic_only): on 6.84 and 7.32 s, off 6.66 and 6.61 s, so 0.44 s. Two timings of one setting
differ by as much as that, so what this supports is a few tenths of a second on that learner,
not a precise share. The update has got faster since, and the fill has not been timed on it.
So T2 keeps it where the harness is checking and turns it off elsewhere. determinism.apply turns it
on, so start-up (preflight, the VRAM probe) runs with it, and so does anything else this process
does outside an iteration, such as royalelearn play. The battles of royalelearn eval and gate
are played by evaluation workers when rollout.eval_workers is above 1 (the default is 2), and
those run without it, as below. At the top of every iteration the coordinator sets it from
learn.ppo.checks_due(config.ppo, iteration) -- on in every iteration whose ratio invariant is
checked, the first ppo.debug_assert_iterations included -- and also on in the first iteration a
process runs, because after a resume that is where the process's leftovers first differ from a run
that never stopped. Everything the coordinator's process does in the iteration runs under that
setting: collection, the update, the checkpoint, and any battle it plays itself. Evaluation workers
are separate processes and do not take the determinism settings, so the setting does not reach the
battles they play. Nothing records the setting; it is recomputed from the config and the iteration.
What still guards the iterations it is off in: every allocation on the harness's own path writes
what it allocates (read, 2026-09-26, and tests/test_fill_policy.py compares the learner's digest
with the fill off and forced on, on the CPU, walking the critic_only update); the ratio-invariant and
non-finite alarms run every iteration; and for torch's own CUDA operations, a two-run digest
comparison past the debug window -- the A/B the training runs are measured with -- is the evidence.
That comparison has been run twice on an RTX 4070 Ti with one training configuration, 14
iterations each: on 2026-09-26 with ppo.forced_rows all, and on 2026-09-27 with
critic_only. Each time one run kept the fill on throughout and the other turned it off where the
harness does not check, which in 14 iterations is the last four, and every learner digest was
equal. The
determinism and resume tests in the suite stay inside the debug window and do not reach a fill-off
iteration.
CUBLAS_WORKSPACE_CONFIG=:4096:8 must be set before torch initialises CUDA, so cli.py sets it from
the config before importing torch and determinism.apply asserts it is already set, raising a
message that names the entry point if not. Every worker sets OMP_NUM_THREADS,
OPENBLAS_NUM_THREADS and MKL_NUM_THREADS to 1 before importing numpy: BLAS thread count changes
float reduction order, and the workers are engine-bound rather than BLAS-bound, so it costs nothing.
5.2 Name-addressed seeding¶
royalelearn/seeding.py:
def derive_seedseq(master_seed: int, path: str) -> np.random.SeedSequence:
digest = hashlib.blake2b(path.encode("utf-8"), digest_size=16).digest()
return np.random.SeedSequence(entropy=master_seed, spawn_key=struct.unpack("<4I", digest))
def derive_generator(master_seed: int, path: str) -> np.random.Generator:
return np.random.Generator(np.random.PCG64(derive_seedseq(master_seed, path)))
def derive_int(master_seed: int, path: str) -> int:
"""A reproducible 63-bit integer, for env seeds."""
w = derive_seedseq(master_seed, path).generate_state(2, dtype=np.uint32).astype(np.uint64)
return int((w[0] | (w[1] << np.uint64(32))) >> np.uint64(1))
The namespace is fixed and documented in docs/determinism.md:
| path | consumer |
|---|---|
env/worker/{w}/shard/{s}/gen/{g} |
the one ClashSelfPlayVecEnv.reset(seed=...) for that shard; g is the respawn generation |
match/battle/{b}/ordinal/{k} |
the matchmaker's draw for episode k of battle b |
act/iteration/{i}/cycle/{t} |
the (R,) uniform vector that drives action sampling at cycle t |
scripted/worker/{w}/slot/{r}/gen/{g} |
a worker-side scripted opponent's generator |
ppo/minibatch/iteration/{i}/epoch/{e} |
the minibatch permutation |
eval/seed_set |
the fixed evaluation seed set, drawn once at run start |
eval/match/{comparison_id}/{seed_index}/{side} |
one evaluation battle |
eval/bootstrap/{comparison_id} |
the bootstrap resampling |
torch/init |
network initialisation |
torch/global, torch/cuda |
torch.manual_seed / torch.cuda.manual_seed_all at start-up |
Two rules follow, both enforced by tests:
- No component may call a global RNG.
tests/test_no_global_rng.pyscans the package source and fails onnp.random.not followed byGenerator|PCG64|SeedSequence|default_rng, on a barerandom., and on unseededtorch.rand*. - The matchmaker draws from
match/..., never from an env's generator. Otherwise changing the opponent mixture changes the battle. Because the path is the battle and its ordinal rather than the iteration and the slot, an episode meets the same opponent however the iteration boundary happens to fall across it.
ClashSelfPlayVecEnv autoresets without a seed (env.py:509-510), so an episode's identity is a
function of its shard seed and its reset ordinal rather than of a seed of its own. That is
deterministic and sufficient: the worker records the shard seed path and the reset ordinal on every
EpisodeRecord, which is what royalegym.replay needs to re-simulate one battle. What it is not is
episode-addressable without replaying the battles before it, which is what the
autoreset_seed_fn(game_index, ordinal) hook gives it: an episode's seed is
derive_int(master_seed, "env/battle/{b}/ordinal/{k}"), so it is addressed by name rather than
by how many episodes came before it.
5.3 Run identity¶
royalelearn/identity.py:
class EngineBuild(msgspec.Struct, frozen=True):
engine_class: str # "RustEngine" | "MockEngine"
calibration_digest: str # config()["calibration_digest"]: 16 hex over the values loaded
build_digest: str # config()["build_digest"]: the same hash over the data
# compiled into the extension; equal to the above on a fresh
# build, which is the whole point of having both
catalogue_sha256: str # over [(card.name, card.card_id, card.elixir, card.kind), ...]
path_search: str | None
stale_build_differences: list[str] # MUST be empty; recorded so a failure is auditable
binary_sha256: str = "not recorded" # config()["engine"]["params"]["engine_binary_sha256"]:
# the extension FILE; "not stated" for an engine with none
class RunIdentity(msgspec.Struct, frozen=True):
format_version: int # of this struct; currently 1
royalelearn_version: str; royalelearn_git: str
royalegym_version: str; royalegym_git: str
engine_build: EngineBuild
env_spec_digest: str # envspec.env_value_digest: each component's kwargs bound
# to its signature, defaults applied, numbers as floats
obs_digest: str; action_digest: str
frame_stack: int # obs.frame_stack; it changes the trunk's input width
arch_digest: str
codec_version: int; codec_table_digest: str # the version and the table it produced
algo_digest: str # sha256 over the PPO, GAE and schedule config
rollout_digest: str # sha256 over {workers, games_per_worker, shards, T, R}
ladder_digest: str # sha256 over {mixture, gate params, eval seed set sha}
master_seed: int
determinism_tier: str
torch_version: str
device_kind: str # "cuda:NVIDIA GeForce RTX 3050 Laptop GPU:sm_86" | "cpu:x86_64"
user_code: dict[str, str] | None = None # {user package: sha256 of its .py source}
extensions: dict[str, ExtensionRecord] | None = None # section 19.3: each section by
# content; the struct encodes with omit_defaults, so a run
# without one keeps the run id it had before the field
run_id = sha256(msgspec.json.encode(identity, order="deterministic")).hexdigest()[:16]
Included even though they look cosmetic, because each changes the byte stream: master_seed,
workers, games_per_worker, shards_per_worker, timesteps_per_iteration, determinism_tier,
device_kind. Excluded and merely recorded, because none of them can change a number: run_name,
runs_dir, wandb settings, console verbosity, checkpoint.keep, alarm thresholds and severities
(an alarm can halt a run or warn, never alter a value). A resume diffs the excluded fields and prints
every difference; a difference in an included field is refused by name.
Both digests come from ClashParallelEnv.config(), which reports calibration_digest over the
calibration values as loaded and build_digest over the copies compiled into the extension.
RustEngine.build_digest() is the same number read directly off the engine. The harness hashes no
data file of its own: a digest computed here from royalegym.protocol.data_dir() would be a second
opinion about what the engine is running on, and a second opinion is exactly what a stale build looks
like. When they differ, stale_build_differences carries the engine's own list verbatim and the run
refuses to start (section 7.7).
The environment is identified by value, not by spelling. env_spec_digest, and the ladder's
result context, hash envspec.env_value_digest: each component's kwargs bound to its real signature
with the defaults applied and numbers taken as floats. Hashing the spec as written made "kwargs":
{} and the same defaults written out two identities, made 1 and 1.0 two, and -- the silent
direction -- let a default changed in the code move the objective of every config relying on it
without moving any digest. A bool stays a bool; a default that is an object goes through
msgspec.to_builtins, or is recorded by its type's name if it will not, which is deterministic but
hashes two such defaults of one type alike. EnvFactorySpec.digest() still hashes the spelling and
is used for display only. Found in review of the 2026-09-24 placeholder commits.
The binary is the third field, and the two digests cannot stand in for it. build_digest hashes
the DATA compiled into the extension, not the Rust it was compiled from. Across a rebuild from an
edited state.rs on 2026-09-23 it read 8952c1c4aa7d7923 on both sides while the extension file's
hash went to 04d42611db5d6e5a [M]. Until 2026-09-24 engine_build dropped that hash, so two
runs either side of the rebuild wrote identical blocks and played different games. It is read by one
rule, envspec.engine_binary, from the env's own config: RustEngine states the hash of the file it
loaded, and an engine with no compiled file is "not stated". What is still not published is which
commit built the file, so an engine built from uncommitted source is identified but not attributable.
Three things follow from recording it.
- An identity written before the field existed decodes with
"not recorded". That is neither a match nor a mismatch, because nobody looked. A resume compares the rest ofengine_buildas usual, leaves the binary out of the comparison, and printsnot checked on resume: engine binary: .... Refusing would strand every earlier checkpoint over a difference nobody can show. Passing in silence would claim a check that never happened. - Every rollout worker states the binary it loaded in its startup report, by the same rule, and the farm refuses a worker whose binary is not the identity's. That covers the window between the parent measuring the engine and the workers loading it.
- The user's own code is recorded by content. Every component is named by its dotted path
and the path is hashed as a string, so for code outside the three packages nothing backed the
name: a different body of
mybot.rewards.MyRewardwas the same run, and the ladder pooled the two objectives.user_codeis{top-level package: sha256}over every.pyfile of each user package a component comes from, with line endings normalised. A package counts when a component's class lives in it, or when a string in a component's kwargs names a module underextra_component_modules, which is how a composition refers to its parts. A module that is only LISTED does not count, which keepsextra_component_modulesout of the identity as the table above says. The whole top-level package is hashed, not the one file, because a reward reads its helpers. Content rather than commit, because a user's module is often in no repository at all. An identity from before the field decodes asNoneand is handled like an unrecorded binary.trainandresumealso refuse uncommitted edits to those packages, the way they refuse one to a sibling repo;--allow-dirtyruns anyway. Until 2026-09-24resumehad neither the check nor the flag, although the refusal's own docstring said it lived on both. - A restarted worker is checked again, which is the case that matters. The parent measured the file at start, and a worker restarted after somebody rebuilt the engine loads the new one. It comes up healthy on another game. It is refused in those words, not as a replacement that failed to start, and it is stopped rather than left running.
5.4 What a resume reproduces, operationally¶
After loading a checkpoint written at the end of iteration k:
- The checkpoint's
RunIdentityequals the identity computed from the current config and environment, field for field, or the load is refused with every differing field named. - Every RNG state is restored rather than re-seeded: torch CPU, torch CUDA on every device when the
run is on a CUDA device (a CPU run captures none, and restores none it finds), the
minibatch generator's
bit_generator.state, pythonrandom, and each shard's seed path, reset ordinal and respawn generation. - Schedule positions are restored: current
gamma,ent_coef,ent_coef_noop, both learning rates, the lr-backoff counter,cumulative_env_steps,cumulative_timesteps,iteration,cumulative_updates. -
The learner is byte-identical to what the checkpoint recorded: the
state_digestreported on load equals the one written at iterationk. That is the weights, both optimizers' moments, the return scaler and the schedule positions, so it is the whole learner rather than a curve that resembles one. -
Iteration
k+1reproduces the original's metric row, field for field, when the checkpoint was written at an episode boundary.
Where that last clause stops, and why it stops there. An episode is addressed by its battle
and its ordinal: ClashSelfPlayVecEnv takes an autoreset_seed_fn, the harness gives it
derive_int(master_seed, "env/battle/{b}/ordinal/{k}"), and a resumed worker is put back on the
ordinal the original was about to play through set_episode_ordinals. Both counters are restored:
the environment's, which decides which battle is played, and the matchmaker's, which decides who
plays it. Restoring one without the other gives the right battles against the wrong opponents, or
the reverse.
What is not restored is an episode that was half finished when the checkpoint was written. Its transitions were already in the original run's buffer, and replaying it would count them twice, so a resumed run starts the next episode instead. The battles it then plays are the right ones and their phase is not: the rows differ in how many episodes completed, never in which were played. Two tests draw that line. One asserts the row-for-row continuation from a boundary. The other measures the phase difference from inside an episode, and it fails if resuming mid-episode is ever offered. The warm-up that spreads first-episode phases apart is skipped on a resume for the same reason: it would move every battle off the episode it is continuing.
state_digest, defined once in royalelearn/checkpoint.py, is a sha256 over, in this fixed order:
the actor state_dict tensors (name-sorted, .cpu().numpy() bytes), the critic state_dict, each
optimizer's numeric state_dict leaves, the return-Welford triple, and the schedule positions. It is
logged every iteration, so when two runs do diverge the iteration at which they first differ is a
lookup rather than an investigation.
6. Configuration¶
royalelearn/config.py. msgspec.Struct throughout, every model forbid_unknown_fields=True so a
typo is an error rather than a silently ignored key, one JSON file, hashed with sha256 over the
canonical encoding. msgspec rather than pydantic because it is already a RoyaleGym dependency, it
gives forbidden-unknown-fields, tagged unions and a canonical JSON round-trip, and the package's
dependency list stays royalegym, numpy, msgspec.
There is one source of truth. The config object is it; there is no keyword-argument facade that duplicates its leaves, because two sources of truth for the same number is the kind of variable this harness exists to remove.
class RunConfig(Struct, forbid_unknown_fields=True):
format_version: int = 1
run_name: str = "royalelearn" # cosmetic, not in the identity
runs_dir: str = "runs" # cosmetic
master_seed: int = 20260921 # IN the identity
profile: str = "laptop" # selects the shipped defaults; all still overridable
timestep_limit: int = 100_000_000 # learner transitions
env: EnvFactorySpec
eval_env: EnvFactorySpec | None = None # None -> env with truncation removed
extra_component_modules: list[str] = []
obs: ObsConfig = ObsConfig()
rollout: RolloutConfig = RolloutConfig()
net: NetConfig = NetConfig()
ppo: PPOConfig = PPOConfig()
advantage: AdvantageConfig = AdvantageConfig()
ladder: LadderConfig = LadderConfig()
checkpoint: CheckpointConfig = CheckpointConfig()
metrics: MetricsConfig = MetricsConfig()
alarms: AlarmConfig = AlarmConfig()
determinism: DeterminismConfig = DeterminismConfig()
doctor: DoctorConfig = DoctorConfig()
| field | default | unit | why |
|---|---|---|---|
obs |
|||
frame_stack k |
1 | cycles | the observation carries positions and no motion, so a policy that needs to tell an advancing unit from a retreating one needs more than one frame. Stacking is a gather of buffer rows (section 9.1), costs no extra storage and adds S + hand_size input channels per extra frame. 1 by default because the extra channels cost stem compute and the first run measures whether they buy anything |
rollout |
|||
source |
"process" |
"process" / "inline" |
inline is the semantic reference and runs every fast test |
workers |
3 | processes | 6-8 threads; the parent needs one for the CUDA driver and one for the loop, and the learner is the bottleneck so more workers buy little |
games_per_worker |
32 | battles | with workers=3, 96 battles and 192 slots |
shards_per_worker |
2 | vec envs | a worker steps one shard while the parent infers on the other. Worth min(env, inference) per round and no more, because the two are additive (section 2.2); bench does not separate the shards: two bench runs on a quiet machine, at 1 and at 2, compared on collection seconds, say what it is worth here, and 1 is the right value when the second shard buys under about 10% |
spin_us |
100 | microseconds | spin, timed with perf_counter, before sleeping on the round's semaphore (section 7.3); a Windows Event round trip measures ~40-80 microseconds [A], a spin about 2 |
round_timeout_s |
30.0 | seconds | every wait has a timeout; a timeout is a typed WorkerFailure, never a block |
restart_failed_workers |
true |
||
max_restarts_per_worker |
3 | three restarts of one worker in a run raises rather than silently degrading throughput | |
launch_delay_s |
0.5 | seconds | three RustEngine constructions at once each decode calibration.json and arena.json |
stagger_first_reset |
true |
desynchronise episode phase across battles at run start, so episode ends spread across cycles | |
overlap |
false |
lag-1 collection under the update, for about 1.37x wall clock at the cost of a second buffer (623 MB at 95 cards). Not built: start-up refuses true (since 2026-09-22), and no profile sets it |
|
eval_workers |
2 | processes | the gate's farm, built at gate time and closed after |
eval_games_per_worker |
24 | battles | |
net (ArchSpec) |
|||
channels C |
64 | about 430 k parameters per network; 223 TFLOP per iteration is about 25 s bf16 on a 3050 | |
blocks N |
4 | residual blocks | |
norm_groups |
8 | GroupNorm, never BatchNorm: BatchNorm computes a different function at rollout than at update, which breaks the stored-log-prob contract | |
vector_embed |
32 | channels | the scalar vector is broadcast into the map so elixir can modulate every tile |
value_hidden |
256 | ||
card_embed |
64 | must equal channels |
the pointer inner product |
coord_conv |
true |
the arena is not translation-invariant: own half, enemy half, the river, the bridges, the tower rects | |
separate_trunks |
true |
makes "the mask and the entropy bonus never touch the critic" structural | |
policy_head |
"pointer" |
"pointer" / "factored" |
factored writes the same flat log-probabilities as three stages: wait or act, then which hand slot or ability button, then which tile. A slot is a candidate when any of its tiles is legal; with no candidate it waits |
factored_act_init |
0.1 | probability | the factored head's P(act) at initialisation, through the gate's bias. Not in arch_digest: it is a starting value, not a shape. A factored actor's spec.json records it and the head: royalelearn.learn.nets.head_meta(net), which a run's artifact_spec() carries and a tool writing its own SnapshotSpec must put in its meta |
logit_scale |
"rsqrt_c" |
one over the square root of C on the pointer inner product | |
init |
"orthogonal" |
gain sqrt(2) hidden, 0.01 policy head, 1.0 value head | |
noop_bias |
0.0 | logits | self-limiting at init; the formula for when it is needed is in section 8.4. Refused with the factored head, which has no no-op logit |
autocast_dtype |
"bfloat16" |
Ampere bf16 tensor cores; bf16 keeps fp32's exponent range so finfo.min masking is safe and no GradScaler is needed |
|
device |
"cuda" |
||
ppo |
|||
n_epochs |
3 | the family disagrees by three orders of magnitude (1, 10, about 1000 gradient steps per 50 k), so none of them is evidence; 3 balances reuse against the update being the wall-clock bottleneck. Watch clip fraction across epochs | |
timesteps_per_iteration |
32 768 | timesteps | the compute budget and the episode arithmetic land on the same number independently |
batch_size |
4 096 | samples per optimizer step | 24 optimizer steps per iteration; about nine episodes' worth of terminal signal per step, so the advantage mean is not dominated by one battle |
minibatch_size |
256 | samples per forward | a pure memory knob: gradients accumulate weighted by n/batch_size with one optimizer.step() per batch, and a test proves the accumulated gradient equals the full-batch one. It was 512, chosen as "the largest that fits 4 GB", and 512 does not fit (see below). The default lives in PPOConfig.minibatch_size itself, because every struct default in config.py is the laptop column; the workstation (2 048) and many-core (8 192) profiles set their own. Until 588dc04 the struct still said 512, so royalelearn config --profile laptop, and doctor or bench without --config, built the value that spills |
clip_range |
0.2 | the PPO paper; all three references | |
dual_clip_c |
3.0 | for a negative advantage the standard minimum does not bound the loss below, and in a 2305-way masked space a rarely-sampled action's ratio can be enormous | |
vf_coef |
1.0 | separate networks and separate optimizers, so it only scales the critic's effective learning rate | |
value_clipping |
false |
||
ent_coef |
Linear(0.01 -> 0.003 over 30e6 env steps) |
the community range across the three references; exploration matters less as the pool strengthens | |
ent_coef_noop |
Linear(0.02 -> 0.0 over 10e6 env steps) |
a warm-start guard against no-op collapse, applied to a binary entropy bounded by 0.693 nats, so while it is on it needs a larger coefficient than the joint term. It anneals to zero because H2(p_noop) is maximised at p_noop = 0.5 while a healthy policy sits near 0.94: a constant coefficient would bias every converged policy toward overplaying. The alarms outlive the schedule |
|
entropy_coef_stages |
null | [gate, candidate, tile] |
null puts ent_coef on the whole entropy. Three numbers put one coefficient on each term of the entropy's chain-rule split -- wait or act; which slot or button given act; which tile given the slot -- in place of ent_coef. The terms sum to the whole entropy on either head, and an update that reads them publishes ppo/entropy_gate, ppo/entropy_candidate, ppo/entropy_tile and ppo/p_act |
max_grad_norm |
0.5 | applied to the actor and the critic parameter sets separately | |
lr_actor, lr_critic |
2e-4 | the community moved down from 3e-4 for long runs | |
adam_eps |
1e-5 | the PPO paper's epsilon, not torch's 1e-8 | |
lr_backoff.kl_threshold |
0.02 | ||
lr_backoff.patience |
3 | iterations | |
lr_backoff.factor |
0.5 | ||
lr_backoff.lr_min |
2e-5 | ||
advantage_standardization |
true |
once per iteration, over all trainable cells, before any split | |
critic_chunk |
1024 | rows | the whole-iteration critic pass is a 2.6 GB spike at this row size if it is not chunked |
discard_opponent_rows |
true |
the frozen seat's log-probs came from other weights; V-trace is the wrong tool against a multi-generation-old snapshot | |
keep_previous_iterations |
0 | iterations | the buffer is cleared every iteration: each timestep is trained on exactly n_epochs times and discarded. A rolling window is available at a linear memory cost |
debug_assert_iterations |
10 | iterations | the two mask asserts and the ratio invariant run for this many iterations from a run's start |
check_ratio_invariant_every |
50 | iterations | thereafter |
ratio_atol |
{"fp32": 1e-4, "bfloat16": 2e-2} |
selected by net.autocast_dtype. 0.0 is not claimed and neither is 1e-4 under autocast: the rollout forward and the update forward are different batch shapes, cuDNN picks a kernel per shape, and bf16 carries three significant digits, so the fp32 logits at the end of the two forwards differ at the 1e-2 level. Section 8.2 says what the check does and does not detect at that tolerance |
|
advantage |
|||
gamma |
Geometric(0.997 -> 0.999 over 20e6 env steps) |
0.99 (both reference learners' default) discounts the win to 0.027 by the end of regulation, wrong by about 25 times in terminal reach. The anneal follows OpenAI Five's 0.998 to 0.9997, with an endpoint scaled to a four-minute match | |
gae_lambda |
0.99 | one over one minus gamma times lambda is 91 steps, which is 45.5 s, matching the deploy-push-tower causal chain. The references' 0.95 gives 9.8 s and cannot connect a deploy to the tower it takes. The single highest-leverage number in the config, which is why credit_horizon_seconds is logged every iteration |
|
standardize_rewards |
true |
gamma near 1 inflates return magnitudes roughly tenfold over gamma = 0.99 | |
reward_clip |
10.0 | standard deviations | independent of the standardisation |
ladder |
see section 11 | ||
mix |
(0.50, 0.35, 0.15) |
mirror / pool / scripted | |
max_resident_opponents |
2 | snapshots | at most four batched forwards per shard-round, and an LRU that never thrashes |
pfsp_weighting |
"hard" |
training wants opponents that still beat the learner; evaluation uses "variance" |
|
pfsp_power |
2.0 | ||
pfsp_uniform_floor |
0.2 | the anti-forgetting guard AlphaStar's forgotten-players slice does | |
weight_floor_scale |
0.01 | every snapshot's weight floored at this over the pool size | |
candidate_every_env_steps |
4 000 000 | game-steps | about two hours on the laptop, so the gate is about 5% of wall clock |
floor_admit_every_env_steps |
50 000 000 | game-steps | a plateau cannot starve the pool |
pool_working_size |
48 | snapshots | sampling is linear in the pool per episode, in python |
eval_seed_count |
500 | seeds | |
release_mode |
"stochastic" |
rating both the sampled and the argmax variant doubles the pool and the cost for no decision. It also takes "argmax" and "plugin:<module>:<function>", as opponent_mode does |
|
opponent_mode |
"stochastic" |
how a frozen pool or seed opponent picks its action in TRAINING battles: "stochastic" samples (every run before this field), "argmax" takes the mode, and "plugin:<module>:<function>" calls a function of your own with the rows' masked log-probabilities, mask and observation vector, which returns one legal action per row (royalelearn/learn/decode.py). The learner's seats always sample. Training battles are not rated, so it is not in the ladder context |
|
refit_every_iterations |
10 | ||
probe_every_iterations |
0 | iterations | 0 is off, and off is the default: a probe plays real battles in the parent. Section 11.9 |
probe_games |
40 | battles per rung | paired, so 20 seeds of the frozen set, each from both sides |
probe_opponents |
("scripted:noop", "scripted:random_legal") |
the rungs, refused at config time if one is not a scripted opponent | |
scripted_opponents |
("scripted:noop", "scripted:random_legal") |
who the scripted share trains against, one drawn uniformly per episode. Refused at config time if one is not a scripted opponent, if one is named twice, or if the list is empty while mix has pool or scripted battles. Training only: the gate's anchors and rater.anchor do not move. Section 11.3 |
|
seed_snapshots |
() |
frozen actors (SeedSnapshot: name, path, sha256 of actor.safetensors, and weight, default 1, a multiplier on its draw weight after PFSP's shape and floor) admitted to the pool as seed:<name> before the first iteration, never evicted and never v0. At start each folder's weights are checked against sha256 and its spec against the run's arch_digest, obs_digest and codec_table_digest, then copied into the run's own snapshots/. Refused at config time for a repeated name, a name holding :, a sha256 that is not 64 hex digits, or a weight that is not a positive number. With mix (0, 1, 0) and the gate's and floor's cadences past the run's end, every battle meets the seeds and nothing else |
|
seat_decks |
None |
one deck trained where the learner sits (SeatDecks: deck, field, shuffle). A mirror battle deals deck to both seats; a pool or scripted battle deals it to the learner's seat, fixed per battle (even battles blue, odd red), and a field deck to the opponent's, whose rows are not trained on. Each battle's deal is a RoyaleGym DeckCurriculumStateMutator the workers install in place of the env's own (royalelearn/ladder/seat_decks.py). Decks are 8 distinct card names, checked at preflight and resolved at the first reset |
|
gate.champion_games |
1000 | battles | 500 seeds times 2 side assignments |
gate.champion_lower_bound |
0.52 | score rate | requires an observed 55.2% or better, about 35 Elo |
gate.anchor_games |
200 | battles per anchor | |
gate.anchor_tolerance_pp |
2.0 | percentage points | |
gate.stratified_snapshots |
8 | ||
gate.stratified_games |
100 | battles each | |
gate.bootstrap_resamples |
10 000 | ||
rater.prior_sd |
400.0 | Elo | weakly informative; keeps the fit finite for a player with no losses |
rater.anchor |
"scripted:noop" |
the gauge exists from the first game, before any snapshot | |
rater.draws |
"davidson" |
"davidson" / "half_win" |
fall back if the measured draw rate is under 2% |
checkpoint |
|||
every_env_steps |
2 000 000 | game-steps | |
keep |
10 | checkpoints | |
include_buffer |
false |
the iteration boundary is the resume point; there is no mid-iteration state to preserve | |
strict_load |
true |
tolerate-everything is right for a research tool and wrong for a harness that promises the curve continues | |
metrics |
|||
sinks |
[Jsonl, Console, Viser, Wandb(enable=false)] |
JsonlSink is always installed: it is the file the resume test compares. ViserSink costs one clock read per iteration while no viewer is attached |
|
image_every |
50 | iterations | per-card tile heatmaps |
keep_episode_log_iterations |
200 | iterations | after which episodes.jsonl is gzipped in place |
determinism |
|||
tier |
"run_exact" |
section 5.1 | |
torch_threads |
1 | thread count changes CPU reduction order | |
doctor |
|||
ram_budget_mb |
6500 | MB | of 7800 |
run_mask_disagreement_gate |
true |
every one of the 2304 non-no-op actions, at one state: the one the start-up sample's play reaches, so only the cards in hand then | |
vram_headroom_mb |
256 | MB | device memory that must stay free after one minibatch's measured peak, or the run does not start (section 7.7, step 12). The same peak plus this margin is health/vram_needed_mb, which is published every iteration so a slowdown can be read against the regime it happened in (section 13.3). 0 turns off the gate |
examples/configs/laptop.json and workstation.json are their profiles written out in full.
tests/test_config.py::test_the_shipped_example_is_its_profile_written_out_in_full holds each to its
profile leaf for leaf, twice: the loaded tree must equal the profile, and the raw document must carry
every leaf the profile dumps. Only the second can see a field the file leaves out, because loading
fills an omitted field in from the profile. It found doctor.vram_headroom_mb missing from all three
example files, and they carry it now (588dc04).
EnvFactorySpec (rollout/envspec.py) is the JSON-able env description, simultaneously the thing
sent to workers, the thing recorded in the checkpoint and the thing the ladder's context string is
derived from:
class ComponentSpec(Struct, frozen=True):
cls: str # "royalegym.obs.SpatialObsBuilder"
kwargs: dict[str, JsonValue] = {}
class EnvFactorySpec(Struct, frozen=True):
engine: ComponentSpec
obs_builder: ComponentSpec
action_parser: ComponentSpec
reward_fn: ComponentSpec
state_mutator: ComponentSpec
termination: list[ComponentSpec]
truncation: list[ComponentSpec] = []
decision_ms: int = 500
def build_vec(self, num_games: int) -> "ClashSelfPlayVecEnv": ...
def digest(self) -> str: ... # sha256 of the canonical JSON
build_vec resolves each ComponentSpec to a class and hands the result to RoyaleGym's picklable
EnvFactory(engine=..., obs_builder=(cls, kwargs), decision_ms=...) recipe rather than constructing a
ClashParallelEnv itself. One place knows how a ClashParallelEnv is assembled, and it is the place
that also has to keep it picklable for spawn; a second assembly here would be a copy that drifts.
EnvFactorySpec stays the JSON layer above it: the thing sent to workers, recorded in the checkpoint
and hashed into the ladder's context. ClashParallelEnv.config() read back off the built env is
what proves the two agree.
Only classes on an allow-list (royalegym.*, royalelearn.*, plus anything named in
config.extra_component_modules) may be instantiated. A spec is data that arrives from a config file
and later from a checkpoint, and importlib on an arbitrary string is a code-execution surface.
Why the laptop profile's minibatch_size is 256 [M]¶
Measured 2026-09-22 on the RTX 3050 Laptop (4294 MB), interleaved 512, 256, 512, 256 in one window so the arms cannot drift apart:
minibatch_size |
time/update over four runs |
vram_reserved_mb |
driver free |
|---|---|---|---|
| 512 | 233, 226, 180, 217 s, spread 29% | 4243 MB (0.99x the card) | 0 MB |
| 256 | 47.7, 48.4, 47.4, 48.5 s, spread 2.3% | 2198 MB (0.51x) | 1176 MB |
Read the decomposition, not the ratio. 256 is stable at 48 s; 512 takes 180–233 s depending on what else is resident on the card. Dividing those gives 4.46x, but the spread is a real distribution rather than measurement error, and the ratio invites a reader to expect 4.46x on a 24 GB card where there is probably no difference at all, because there is no cliff to be on.
The mechanism, confirmed rather than inferred. At 512 the allocator reserves 4243 MB against a
4294 MB card. On Windows the driver does not refuse an oversubscribed allocation, it backs it with
host RAM over PCIe, so the update streams tensors across the bus instead of computing from VRAM.
Three things establish it: reserved reaches 1.24x the physical card when cuDNN benchmarking is
also allowed to allocate; set_per_process_memory_fraction(0.95) turns the slow run into an
immediate OutOfMemoryError, so the memory was coming from beyond the card; and the 29% spread
sits entirely on the 512 arm while 256 holds 2.3% in the same window. Machine noise, thermal
throttling and scheduler jitter all predict both arms vary, and only one does.
Two things this measurement should not be read as saying. It is one card on one platform, and the
cliff is a property of the footprint against the device rather than of 512 as a number: the
workstation profiles' larger minibatches are correct for their larger cards. And num_alloc_retries
reads 0 throughout, which is not evidence the allocator is innocent. It counts the
allocation-failure path, and on this platform the allocation never fails. It went to 2 the moment
the memory fraction made failure possible. An instrument that cannot fire here is indistinguishable
from one with nothing to report.
7. The rollout path¶
7.1 Process model¶
parent (learner: torch, CUDA, buffer, ladder, metrics, checkpoints)
|-- worker 0 (python, numpy, royalegym, royalesim; NO torch)
| |-- shard 0: ClashSelfPlayVecEnv(num_games=M/2, env_fn=EnvFactorySpec.build)
| '-- shard 1: ClashSelfPlayVecEnv(num_games=M/2, ...)
|-- worker 1 ...
'-- worker K-1 ...
- Start method
spawnon every platform. Windows has nofork, CUDA is already initialised in the parent soforkwould be wrong anyway, and forcingspawneverywhere means the worker's import-time environment is identical on Linux. - The child is handed an
EnvFactorySpec, a msgspec Struct, never a pickled closure. That removes the reference learner's undocumented 4096-byte factory limit and gives the checkpoint its env description for free. - The child's preamble, in this order, before numpy is imported:
os.environ["OMP_NUM_THREADS"] = os.environ["OPENBLAS_NUM_THREADS"] = "1"
os.environ["MKL_NUM_THREADS"] = "1"
signal.signal(signal.SIGINT, signal.SIG_IGN) # Ctrl-C is the parent's business
launch_delay_sbetween starts:KRustEngineconstructions at once each decodecalibration.jsonandarena.json.- The child holds no policy.
tests/test_worker_hygiene.pyasserts"torch" not in sys.modulesafter start-up.
7.2 The observation codec¶
rollout/codec.py, SpatialObsCodec, codec_version = 1. The codec version is the rule; the
table it produces is computed at preflight from EnvSpec.obs_space and a sample of real
observations, so a plane added, removed or rescaled upstream reaches the right storage without a line
changing here.
The rule, applied per key:
| key | rule | storage |
|---|---|---|
spatial, per plane |
declared static by ObsBuilder.spatial_layout() |
not stored; held once per seat in statics |
spatial, per plane |
declared high <= 255 and integer-valued on 1000 sampled states |
uint8, scale 1, exact |
spatial, per plane |
anything else | float16 |
mask_planes |
equal to action_mask[1:] reshaped, by construction |
never stored; the learner reshapes the stored mask at unpack |
ability_ready |
only with TileActionParser(ability_buttons=True); equal to the mask's last n_buttons bits |
never stored; the network reads it off the stored mask |
vector |
bounded in [0, 1] by the builder | float16, 2 B per element |
action_mask |
bit-packed uint8, LSB first, ceil(n_actions / 8) B, exact |
|
card_ids |
only with SpatialObsBuilder(card_identity=True); vocabulary high + 1 must fit a byte |
uint8, exact, its own region after the mask, ids_planes * H * W B |
| anything else | refused at bind, naming the key |
card_ids is D2's learner half. It is never a plane of spatial, because a card id stored as a scaled
half and read back as 6.997 would be embedded as card 6. It gets a region of its own, after everything
that was already in the row, and CodecTable.ids records it as "uint8". With the switch off the
table has no ids field in its encoding at all (omit_defaults), so every table decided before D2
hashes exactly as it did, and the row is byte-for-byte the same length. A table and a space that
disagree about card_ids, in either direction, are refused, as is a vocabulary over 256.
Any other key is refused too, and that is new. The codec used to read the three keys it knew and ignore the rest, so the day a builder started sending card identity, the planes would have been dropped from every row before the network saw them.
On the 95-card catalogue of 2026-09-22 with Reveal off, S = 20 and that rule splits them 16 / 2 / 2: sixteen
uint8 planes, the two hit-point planes as float16, and the two static planes. The hit-point planes
are the only ones whose declared high reaches 64, since every other plane is a small integer count
or an indicator in [0, 1]. They are also the only ones that fail the integer test, because they carry
a fraction of full health; it is the second fact and not the first that sends them to float16. The
vector is 1 177 wide. Turning on Reveal.enemy_spell_aim adds a spatial plane, and with it one more
uint8 row; the table below is then recomputed from the rule rather than patched.
| part | stored as | bytes | exact? |
|---|---|---|---|
| 16 count and indicator planes | uint8, scale 1 |
9 216 | exact: the declared bound is under 255 and the sampled values are integers |
| 2 hit-point planes | float16 |
2 304 | exact to fp16; the values are hit points scaled into a small range, so the granularity is below the engine's own |
| 2 static planes | not stored | 0 | exact |
| 4 mask planes | not stored, derived from the mask | 0 | exact |
| vector, V = 1177 | float16 |
2 354 | exact to fp16; the builder clips the whole vector to [0, 1] (obs.py:181) |
| action mask | bit-packed uint8 |
289 | exact |
| row total | 14 163 B (95 cards, 2026-09-22) | 3.91x smaller than what the env hands over |
The same rule on MockEngine's 16-card catalogue gives 9 216 + 2 304 + 2*229 + 289 = 12 267 B, and
the whole test suite runs there precisely because its widths are not the Rust catalogue's.
Notes an implementer needs:
- The table is printed at preflight, written into the metric stream once, hashed into
codec_table_digestand stored in thespec.jsonof every snapshot. Two runs whose tables differ are not comparable and the digest is what says so. mask_planesis never stored. RoyaleGym guaranteesmask_planes == action_mask[1:].reshape(hand_size, tiles_y, tiles_x)exactly, with its own test; the learner doesmask[1:].view(...)at unpack. Storing them would be 2 304 B of a 14 163 B row (95 cards) spent on a reshape.- A static plane is one the layout declares static, and nothing else. On this catalogue those are the water and no-deploy masks. There is no sampling fallback and there must not be: the tower planes are constant across any sample in which no tower falls, so a rule that promoted "unchanged on a thousand states" to "static" would freeze a tower at full health for the rest of the run and the policy would never see one die. Storage is decided from a sample; existence is decided from the declaration.
- The static planes are asserted equal between the blue and the red row at start-up. They are, because the observation is in the acting player's own frame and the arena is symmetric under the 180-degree seat rotation; if they ever differ the codec stores one pair per seat and says so in a warning.
- A
uint8plane whose value exceeds 255 is clipped and counted inhealth/obs_codec_clipped. A non-zero value is a bug report, not a tolerance: the plane was admitted touint8on a declared bound, so a clip means the bound was wrong. unpack_to_deviceruns on the GPU:uint8 -> float32times a per-channel constant divisor,float16 -> float32, bit-unpack the mask with shifts, reshape it into the mask planes, scatter the static planes in, and gather the frame-stack history rows. Its FLOP count is nothing against 378 MFLOP per sample.- The per-channel divisors are fixed constants from the declared bounds, not running statistics. Observations are bounded by construction, Welford standardisation would divide a wide one-hot by a near-zero standard deviation and amplify noise that was not there, it adds cross-worker mutable state that a resume has to restore, and it makes the seat-mirror bit-identity unverifiable. The adaptive part is done by the stem's GroupNorm, which has no cross-worker state.
7.3 The boundary is a byte layout¶
rollout/layout.py is the cross-language contract. It defines, as module constants plus
LAYOUT_VERSION: int = 1:
Segment A "rlb-<run_id[:8]>-<pid>-<n>" the experience buffer: learner-owned, worker-written
header 128 B magic, LAYOUT_VERSION, cycles, n_slots, row_bytes, obs offsets, codec_version
obs (T+1) * R * row_bytes the rectangle; index(t, r) = t * R + r
Segment B "rlc-<run_id[:8]>-<w>" one per worker, covering all its shards
per shard, per parity p in {0, 1}:
control 64 B state u32 | cycle u64 | n_slots u32 | err_code u32 | err_len u32 | t_env_ns u64
error 512 B the child's traceback, utf-8
scalars n_slots_shard * 32 B
reward f32, group i8, flags u8 (terminated, truncated, valid),
deploy_status i8, tick i32, episode_end i8, episode_steps i32,
cards_played i16, elixir_leak i16, pad
finals n_trunc_max * row_bytes final_obs of truncated rows, for the bootstrap
actions n_slots_shard * 2 B int16, parent -> child
assign n_slots_shard * 4 B group i8, opponent_ix i8, learner_seat i8, pad;
parent -> child, written only for the slots whose
episode started on the previous round
plan n_slots * 24 B the iteration's opening table: role i8, group i8,
opponent_ix i8, seed hash u64
Signalling: two semaphores per (worker, shard), obs_ready and actions_ready, each released once
per publication or command, after the control word is written. Both sides spin for rollout.spin_us,
then sleep on the semaphore. The spin is timed with perf_counter: time.monotonic() is the tick
count on Windows before Python 3.13 and moves in 15.6 ms steps, so a spin timed by it lasts until the
next step, which is how idle workers once spun through the whole update (38352bf). A wait that sees
the word during its spin takes the token that announced it, so a semaphore holds one token per
command not yet answered, and a late token only wakes a later wait that finds nothing and sleeps
again. A worker sleeps only on the shard whose command is due next, for at most SLEEP_S (50 ms),
and glances at its other shards once per pass; every command a run sends arrives in that order, and
one sent out of turn is taken at most one sleep late. The control word is the truth, and a wake is
never taken as the answer. There is no UDP, no magic float header, no pickling on the hot path and no
length-prefixed stream to desynchronise.
Error signalling: err_code 0 normal, 1 python exception with the message in error, 2 hard crash.
The parent turns either into a WorkerFailure delivered to the coordinator, which restarts the
worker. Both reference learners treat a dead worker as a permanent silent hang.
Because the contract is bytes, a Rust worker is a drop-in: it writes the same header, the same packed
rows and the same scalars, and releases the same semaphores. tests/test_rollout_farm.py is the
acceptance criterion: the process farm and the inline source must produce byte-identical buffers,
scalars and episode records from the same seed over 30 cycles.
7.4 The worker main loop¶
def worker_main(w: int, cfg: WorkerConfig, handles: Handles) -> None:
_set_thread_env() # before numpy
import numpy as np
shards = [build_shard(cfg, w, s) for s in range(cfg.shards_per_worker)]
_startup_gates(shards) # section 7.7
for sh in shards:
seed = derive_int(cfg.master_seed, f"env/worker/{w}/shard/{sh.index}/gen/{cfg.generation}")
sh.obs = sh.vec.reset(seed=seed)
publish(sh, cycle=-1) # hand the parent the first observations
while True:
for sh in shards: # in turn; sleeps only on the shard due next
if not wait_actions(sh, cfg.spin_us):
continue
cmd = read_command(sh) # STEP | PLAN | SET_STATE | SPACES | DEFER | CLOSE
if cmd is CLOSE:
return
if cmd is PLAN:
apply_plan(sh, read_plan(sh)); continue # the iteration's opening table
apply_assignments(sh, read_assignments(sh)) # the parent's draw for the battles
# that just started a new episode
actions = read_actions(sh) # int16[n_slots_shard]
actions = apply_scripted(sh, actions) # worker-side opponents, numpy only
obs, rew, term, trunc, info = sh.vec.step(actions)
pack_round(sh, obs, rew, term, trunc, info) # codec -> the buffer rectangle
publish(sh)
- The scripted opponents run here.
RandomLegalOpponent(noop_prob=0.9)is about 5 microseconds of numpy per row; routing the scripted share of battles through the GPU would cost a forward pass and a boundary crossing for nothing. That share isladder.mix, which ships as(0.50, 0.35, 0.15)-- mirror, pool, scripted -- so it is a configured proportion rather than a measurement, and a run that changes the mixture changes it here too. Which opponents that share meets isladder.scripted_opponents. They draw from a per-slot generator seeded fromscripted/worker/{w}/slot/{r}/gen/{g}, so they stay reproducible. - A battle's assignment changes only at that battle's episode boundary, and the parent decides
it. The worker holds no pending plan and makes no draw: when a round reports
episode_end, the parent draws the next assignment and hands it back on the same round'sStep, before the first action of the new episode is taken. Partial control makes a trajectory unusable, and one place that knows what a battle is playing removes the whole class of bug rather than detecting it. vec.action_masks()is never called, because it restacks the whole batch and the mask is already inobs["action_mask"].info["action_mask"]is ignored for the same reason: it is a second copy of a key the observation already carries.- The terminal statistics are read, not reconstructed. RoyaleGym puts seven flat scalars in
infoon the step an episode ends:episode_steps,episode_ticks,own_crowns,enemy_crowns,own_tower_hp_frac,enemy_tower_hp_fracandelixir_leak_steps. The vec env nests them underinfos["final_info"]as batched arrays with_-prefixed validity masks, so the worker readsinfos["final_info"]["_own_crowns"]to find which rows have one and copies the seven values for those rows into theirEpisodeRecord. The hit-point fractions are already the mean over a player's three towers. What the worker still counts itself is what no one else can: cards played, illegal commands, the undiscounted return and the per-term reward sums. A per-step counter in numpy costs nothing and one fixed record per finished episode crosses the boundary, rather than a per-step metrics array reassembled in the parent.
7.5 The round protocol¶
Per shard-round the parent does:
round = source.next_round(timeout_s) # spin/Event wait; views into shared memory
assign: # the parent owns assignments
for b in battles_that_ended(round.episode_end): # ascending battle order; ends come in pairs
a = matchmaker.assign(b, ordinal[b] + 1, pool, ratings) # match/battle/{b}/ordinal/{k}
apply(a) # updates group for b's two slots, in place
group = round.group # the worker's report, with this round's
# new assignments already folded in
route:
for gid in sorted(set(group)): # sorted: determinism, not arrival order
rows = np.flatnonzero(group == gid) # ascending slot order within the group
if gid == SCRIPTED: continue # the worker fills these in
obs = codec.unpack_to_device(buffer_rows(round, rows), statics)
logits = policy_for(gid).logits(obs).float()
dist = MaskedCategorical(logits, obs.mask)
u = uniforms[round.cycle][round.slots[rows]] # act/iteration/{i}/cycle/{t}
a = dist.sample(u)
actions[rows] = a
if gid == LEARNER:
log_probs[rows] = dist.log_prob(a)
source.submit(Step(actions=actions, gamma=sched.gamma,
group=new_group, opponent_ix=new_opponent_ix,
learner_seat=new_learner_seat)) # only the battles that just reset
buffer.record_round(round, actions, log_probs) # scalars only; observations already in place
The assignment step comes before routing, in the same round, so the first observation of a new episode is already routed by that episode's own assignment and no transition is ever produced under a stale one. The worker is told the result because it fills the scripted seats itself; it is not asked to decide anything.
Row order inside an inference batch is ascending slot index, never arrival order. Group sizes vary from round to round because assignments change at episode boundaries; that costs a little kernel-shape churn and nothing else, and it is the price of binding a policy to a battle for a whole episode, which is the thing that must not be given up.
uniforms[t] is an (R,) float32 vector drawn once per cycle from act/iteration/{i}/cycle/{t} and
indexed by slot, so a sampled action is a pure function of (master_seed, iteration, cycle, slot) and
the logits. That is what makes the trajectory reproducible whatever the batch composition and whatever
rollout.overlap is set to.
Frozen snapshot policies run under torch.inference_mode() in fp16, from a LoadedPolicyCache of at
most ladder.max_resident_opponents + 2 resident modules.
7.6 Failure, timeouts and restart¶
def next_round(self, timeout_s: float) -> RolloutRound:
deadline = time.monotonic() + timeout_s
while not self._shard_ready(shard):
if time.monotonic() > deadline:
self._reap() # is_alive() on every child; exitcode into the log
raise WorkerTimeout(self._diagnose())
self._spin_or_block(shard)
Four properties, none of which either reference learner has:
- Every wait has a timeout.
round_timeout_s = 30,join_timeout_s = 10. is_alive()is checked on every timeout tick and a dead child's exit code is logged.- A dead worker is restarted deterministically: the new worker gets
env/worker/{w}/shard/{s}/gen/{g+1}, so the run remains reproducible from the master seed and the restart count, both of which are checkpointed. health/worker_restartsis a metric, andmax_restarts_per_workerrestarts of one worker raises rather than silently degrading throughput.
The rectangle handles a mid-iteration death without a special case. If worker w dies at cycle t*,
its slots have no observation at t*, so for those slots valid[t, r] = (t <= t* - 2) and the
transition at t* - 2 is marked truncated and bootstrapped from values[t* - 1], which exists.
health/rows_dropped_dead_worker counts the lost cells. The restarted worker rejoins at the next
iteration boundary.
On shutdown: send Close, join(timeout=10), then terminate(), then kill(). close() is
idempotent and is called from a finally in the coordinator and from an atexit hook.
7.7 Start-up gates¶
rollout/preflight.py, run by royalelearn doctor and by the coordinator before the first cycle.
Their results go into the run identity and the checkpoint.
- Construct one engine and reset it once.
RustEngine()raisesRuntimeErrorlisting every stale calibration key. Re-raise asStaleEngineBuildwith the original text verbatim plus the rebuild command. A run must die here, not at cycle 0. The reset comes before anything is read: a freshly constructed env reportsdecision_ticks = 1until its firstreset(), so reading the configuration earlier would record a number that is about to change. - Read
ClashParallelEnv.config(). It returns one JSON-able dict:env,decision_ms,decision_ticks,reveal,engine,obs_builder,action_parser,reward_fn,termination_cond,truncation_cond,state_mutator,calibration_digest,build_digest, each component as{"class": ..., "params": component.config()}. That dict is the single source for the timing fields ofEnvSpec, theEngineBuilddigests of section 5.3 and the ladder'scontextof section 11.6. All three read that one dict, so a component whose parameters change moves the identity, the context and the printed summary together or not at all. - Read the layout off the environment. Every key of
single_observation_spacewith its shape, dtype and per-channel bounds;ObsBuilder.vector_layout();ObsBuilder.spatial_layout(). BuildEnvSpec.obs_space,vector_layoutandspatial_layout, compute the codec table (section 7.2) andobs_digest, and print the table: one line per spatial plane with its storage and divisor, the vector width and its named fields, the mask width and the row total. Nothing in the harness holds a width, a plane count or a field offset of its own, which is why the whole test suite runs onMockEngine, whose widths are not the Rust catalogue's. - Resolve the vector fields the pointer head needs through
vector_layout. They arehand_card_onehot,hand_costandhand_affordable. A missing name is aPreflightErrornaming it. royalegym.action.mask_disagreements(engine, parser, state, team)for both teams, over all 2304 non-no-op actions at one state: the one environment 0 holds after the table sample's play, which runs first. It sees only the cards in hand at that state, so it passes while the mask is wrong for another card, and it does not see the opening state's rules. Non-empty is aPreflightError, not a warning: a policy trained against a wrong mask is worthless, the check is already written and nobody runs it.- Assert
mask[NOOP]on 1000 sampled states including one withgame_overset. This is the one precondition the whole masking scheme rests on. - Assert the action-layout identity exhaustively: for all 2304 non-no-op actions, a one-hot
(hand_size, tiles_y, tiles_x)tensor reshaped in C order has its argmax atparser.encode(slot, x, y); and, where the observation carries mask planes, thatmask_planes == action_mask[1:].reshape(hand_size, tiles_y, tiles_x)on the sampled states. - Record the
EngineBuildfrom step 2's digests and the card catalogue. - Print the RAM ledger and refuse if the projection exceeds
doctor.ram_budget_mb. - Print
credit_horizon_seconds,timesteps_per_iteration,T,R, the learner-row count and therun_id. - Settle the viewer's state stream. The publisher belongs to the vec env, so exactly one vec env
in the whole run carries it: worker 0's first shard is built with
viser="env"and every other shard withviser=None, and within that shard the index-0 battle is the one published. Where the vec env does not take the argument, the workers unsetROYALEVISERin their own environment instead and preflight prints that the state stream is off, because a per-process publisher binds the same UDP port once per env and raisesOSErrorat construction. The learning-status stream of section 13.1 is a separate socket and is always available.
The replay recorder follows the same rule for a different reason. rollout.recorder is an
optional ComponentSpec -- royalegym.replay.SavingReplayRecorder and its JSON kwargs -- and
it reaches worker 0's first shard only, where build_vec attaches ONE built recorder to the
index-0 battle. It is deliberately NOT a field of EnvFactorySpec: a recorder does not change
the game, and hashing it would make a policy trained with one incomparable with every policy
trained without. It is also not passed through the EnvFactory recipe, although EnvFactory
would accept it, because a component is built once per env and the vec env builds one env per
game: that route gives games_per_worker recorders in one process, and the cost is throughput
before it is disk, since ClashParallelEnv leaves the single multi-tick engine.step path
whenever its recorder wants per-tick frames. keep also bounds each instance separately, and
the saved name {pid}-{completed:06d}-tick{tick} cannot separate instances inside one process
-- two collide when two battles end on the same tick, which is the step cap.
12. Measure one minibatch on the device. The coordinator does this once it has built the network;
doctor does not. On CUDA, and unless doctor.vram_headroom_mb is 0, it runs one forward and
backward at ppo.minibatch_size, reads the peak the allocator
reserved, and raises PreflightError naming the next legal minibatch when that peak plus the
headroom exceeds what the driver reports free. The platform backs an oversubscribed allocation with
host memory rather than refusing it, so without this the run would not fail, it would be several
times slower for its whole life. The peak plus the headroom is kept as health/vram_needed_mb, a
reading for the row: the vram_spilling alarm does not use it (section 13.3). A probe that raises is
let through with a printed line saying the gate did not run, and leaves that key unset.
8. The networks and the masked distribution¶
8.1 Exact shapes¶
Card identity (D2), when the builder carries it. Each tile's card_ids go through
nn.Embedding(vocabulary, 8, padding_idx=0) and are concatenated before the stem, as the hand slots'
cards are embedded for the head. The vocabulary is the space's own high + 1, num_cards + 2, never
a literal (it was 100, then 101, since the spec was written). Index 0 is an empty tile and is fixed at
zero. The stem gains k * 2 * 8 input channels. The width 8 is a module constant, CARD_ID_EMBED,
and not a NetConfig field, because a new field would change every run's arch_digest and so refuse
every checkpoint and snapshot on disk for a feature most never switched on. It enters arch_digest
only when card_ids exists. A renumbered catalogue changes the space's bound, so it changes
arch_digest and a trained table meeting the wrong vocabulary is refused rather than reinterpreted.
That is the network's half of the positional-id trap; the builder's card_names check is the other.
Every shape is written in the environment's own terms, because that is how the code computes them.
B batch; C = net.channels; E = net.vector_embed; S, V and A the spatial plane count,
vector width and action count from EnvSpec.obs_space; P = spec.hand_size hand slots;
K = spec.num_cards + 1 card ids including "empty"; k = spec.frame_stack; board
H x W = spec.tiles.
ObsBatch (device tensors, from ObsCodec.unpack_to_device):
spatial (B, k*S, H, W) float32 (bf16 inside autocast)
mask_planes (B, k*P, H, W) float32 derived from mask, never stored
vector (B, V) float32 the current frame only
mask (B, A) bool
card_ids (B, k*2, H, W) int64 only with card identity: own then enemy, per frame;
0 empty, 1 crown tower, 2 + card id; else None
trunk input assembly:
coords (2, H, W) constant buffer, y/(H-1) and x/(W-1) in [0,1]
vemb Linear(V + NB, E)(cat[vector, ready]) -> (B, E, 1, 1) -> expand -> (B, E, H, W)
ready = mask[:, G : G + NB], the ability buttons' readiness; G = 1 + P*H*W is
the grid's width and NB the buttons (0 without ability_buttons, and then
vemb is Linear(V, E) on the vector alone, as it always was)
x cat[spatial, mask_planes, coords, vemb] -> (B, k*(S+P) + 2 + E, H, W)
stem Conv2d(k*(S+P) + 2 + E, C, 3, padding=1) -> GroupNorm(norm_groups, C) -> ReLU
body N x ResBlock: Conv3x3(C,C) GN ReLU Conv3x3(C,C) GN (+skip) ReLU -> (B, C, H, W)
--- actor head ---------------------------------------------------------------
Fp Conv3x3(C, C) GN ReLU -> (B, C, H, W)
hand onehot = vector[:, f("hand_card_onehot")].view(B, P, K)
card_id = onehot.argmax(-1) (B, P) int64
cost = vector[:, f("hand_cost")] (B, P)
afford = vector[:, f("hand_affordable")] (B, P)
Ecard Embedding(K, C) (B, P, C)
q_in cat([Ecard[card_id], cost[..., None], afford[..., None]]) (B, P, C+2)
q Linear(C+2, C)(q_in) (B, P, C)
qb Linear(C+2, 1)(q_in) (B, P, 1)
tiles einsum('bcyx,bkc->bkyx', Fp, q) * C**-0.5 + qb[..., None] (B, P, H, W)
pooled cat([Fp.mean((2,3)), Fp.amax((2,3))]) (B, 2C)
noop Linear(2C, 1)(pooled) + arch.noop_bias (B, 1)
buttons Linear(2C, NB)(pooled), only when NB > 0 (B, NB)
logits cat([noop, tiles.reshape(B, P*H*W), buttons], dim=-1) (B, A) float32
A = G + NB: action G + k presses ability button k (a hero's or a champion's),
masked by the env like every other action
--- critic (its own trunk, identical shape, separate weights) -----------------
pooled_c cat([body_c.mean((2,3)), body_c.amax((2,3))]) (B, 2C)
vfeat cat([pooled_c, Linear(V, E)(vector), legal_frac]) (B, 2C + E + P)
value Linear(2C+E+P, value_hidden) ReLU Linear(value_hidden, 1) -> squeeze (B,)
Worked example, the shipped defaults on the 95-card catalogue of 2026-09-22: k = 1, S = 20, P = 4, E = 32,
C = 64, H x W = 32 x 18, V = 1177, A = 2305, so spatial is (B, 20, 32, 18), the trunk takes
1*(20 + 4) + 2 + 32 = 58 input channels and the stem is Conv2d(58, 64, 3). Those numbers are an
illustration of the expressions above and never appear in the code: on MockEngine's 16-card
catalogue the same expressions give a 229-wide vector and the same 58 channels, and a plane added
upstream changes the stem without anything here being edited.
f(name) is obs_layout.py's lookup into EnvSpec.vector_layout: it returns the (offset, size)
slice the environment declared for that field. The head asks for hand_card_onehot, hand_cost and
hand_affordable by those names and computes no offset of its own. A layout that does not declare a
name the head needs is a PreflightError naming the missing field, raised before the first engine
step rather than discovered as a wrong slice halfway through a run.
The four mask planes are part of the trunk's input because they are a state feature the policy would otherwise have to infer: each is the deploy legality of one hand slot over the whole board, which is exactly "where could I place this card", and it is already computed for the mask. It costs four input channels and nothing else, because it is a reshape of bytes the row already carries.
The head's shape is not a choice. royalegym/action.py:272 is
encode(slot, x_idx, y_idx) = 1 + slot*576 + y_idx*18 + x_idx, so the non-no-op actions are a
(P, H, W) C-order tensor aligned pixel for pixel with the spatial planes, and index 0 is the no-op.
A flat Linear(512, A) head is 1 180 160 parameters against the pointer head's ~10.5 k, and it is the
structural cause of tile spam, because a shared-weight head cannot memorise one output unit the way a
flat head can. Conditioning on a card embedding rather than the slot index is the other half:
slot 0 holds a different card every cycle, so a slot-indexed head has to learn a card-agnostic
detector, and the embedding also buys deck transfer.
legal_frac is P scalars, the mean of each hand slot's slice of the mask. The mask is a function of
state, so feeding a summary of it to the critic is legitimate and mildly helpful; the critic never
sees the full mask and never receives the entropy gradient, which separate trunks make structural
rather than disciplinary.
Parameters at the worked example's values, convolution and linear weights only:
| component | parameters |
|---|---|
| stem 58 -> 64, 3x3 | 33 408 |
| 4 residual blocks (8 convs 64 -> 64, 3x3) | 294 912 |
vector embedding Linear(1177, 32) |
37 664 |
| policy conv + card table 96x64 + query MLPs | 47 400 |
value head, with its second Linear(1177, 32) |
79 120 |
| actor | ~413 k |
| critic | ~445 k |
| actor + critic | ~858 k |
Weights 3.4 MB fp32, Adam moments 6.9 MB. Activations dominate: about 2.36 MB per sample fp32 for the eight body convolutions, so the laptop's minibatch of 256 is 0.59 GB fp32 and 0.30 GB bf16 [A]. That projection counts the body convolutions and nothing else, and it undercounts. It put 512 at 0.59 GB bf16, comfortably inside a 4 GB card; measured, 512 reserved 4243 MB of the 4294 MB card and spilled into host memory, while 256 reserved 2198 MB [M] (section 6). So the minibatch is 256 by measurement, the precision is bf16, and a run measures one minibatch's real peak at startup rather than trusting this paragraph (section 7.7, step 12).
C = 96, N = 8 roughly triples the compute and is the single change to make on a bigger box.
arch_digest covers C, N, the observation space, frame_stack, num_cards, n_actions, the
head name and the trunk class, and a load with a different digest is refused with a message rather
than a shape error.
8.2 Precision rules¶
These are correctness rules, not performance rules.
- Trunk and heads run under
torch.autocast("cuda", dtype=torch.bfloat16). bf16 and not fp16: bf16 has fp32's exponent range, sofinfo.minmasking andlog_softmaxbehave and noGradScaleris needed. - Logits are cast to float32 before masking and before the distribution, at rollout and at update
alike. That makes the masking and the
log_softmaxexact, which is what keepsfinfo.minfrom leaking probability onto an illegal action and keeps the entropy sum finite. - The two forwards do not agree bit for bit, and the spec does not claim they do. The rollout forward runs one policy group of roughly fifty to a hundred and fifty rows; the update forward runs a minibatch of 256 at the laptop profile. cuDNN selects an algorithm per shape and bf16 carries about three significant digits, so the fp32 logits at the end of the two paths differ at the 1e-2 level and the log-probs with them. The fp32 cast fixes the masking, not the convolutions underneath it. Section 9.6 sets the tolerance accordingly and says what the check still catches.
- GroupNorm, never BatchNorm. BatchNorm computes a different function at rollout (a small batch,
running statistics) than at update (a minibatch of 256,
train()mode). That difference is not a rounding difference. It is a different function of the same weights, so the stored log-prob would stop being the log-prob of the action that was taken.
8.3 MaskedCategorical¶
learn/distribution.py is the only place in the package where a mask meets a logit.
class MaskedCategorical(ActionDistribution):
def __init__(self, logits: Tensor, mask: Tensor) -> None: # (B, A) f32, (B, A) bool
assert logits.dtype == torch.float32
assert mask.dtype == torch.bool and mask.shape == logits.shape
assert bool(mask[:, 0].all()), "mask[NOOP] must be True (royalegym action.py:284)"
self._mask = mask
self._logp = torch.log_softmax(logits.masked_fill(~mask, torch.finfo(logits.dtype).min), -1)
def sample(self, u: Tensor) -> Tensor: # inverse CDF
cdf = self._logp.exp().cumsum(-1)
cdf = cdf / cdf[..., -1:].clamp_min(1e-30)
idx = torch.searchsorted(cdf, u.unsqueeze(-1).clamp(0.0, 1.0 - 1e-7), right=True)
return idx.squeeze(-1).clamp_(max=cdf.shape[-1] - 1)
def mode(self) -> Tensor:
return self._logp.argmax(-1) # post-mask, mask-respecting
def log_prob(self, a: Tensor) -> Tensor:
return self._logp.gather(-1, a.unsqueeze(-1)).squeeze(-1)
def entropy(self) -> Tensor:
p = self._logp.exp()
return -(p * torch.where(self._mask, self._logp, torch.zeros_like(self._logp))).sum(-1)
def noop_entropy(self) -> Tensor:
p = self._logp[:, 0].exp().clamp(1e-7, 1 - 1e-7)
return -(p * p.log() + (1 - p) * (1 - p).log())
def n_legal(self) -> Tensor:
return self._mask.sum(-1)
Every detail has a failure it prevents, and every reference implementation gets at least one wrong:
- Mask before the softmax, never after. The two are gradient-identical: a masked logit receives
exactly zero gradient either way. So the difference is numerical, and it is decisive. Post-softmax
masking combined with the references'
clamp(probs, min=1e-11)resurrects every illegal action at p = 1e-11 with a finite log-prob, andmultinomialwill eventually draw one. torch.finfo(dtype).min, neverfloat("-inf").-inf * 0is NaN in the entropy sum; one reference survives only becausetorch.distributions.Categorical.entropy()clamps internally. A library's clamp is not a place to keep a correctness property.- No row can be fully masked, because
action.py:284setsmask[NOOP] = 1unconditionally, including after game over. NaN is structurally impossible and a test says so over 10 000 states. mode()takes the argmax of the normalised masked log-probs. One reference takesprobs.argmaxon an unmasked softmax, which here would emit illegal actions in every evaluation game.sampletakes uniforms. Withright=True, an illegal action's zero-width CDF interval can never be selected, and the sampled action is independent of batch composition and of torch's global RNG. A test compares it against a brute-force inverse CDF.
The mask travels with the transition (289 bit-packed bytes, 2.0% of a row) and the same mask is
applied in backprop. Two asserts run during ppo.debug_assert_iterations:
assert stored_mask.gather(-1, actions.unsqueeze(-1)).all() # every action taken was legal
assert torch.isfinite(log_probs).all() # no -inf reached the loss
A rollout/update mask mismatch is silent and slow-acting: unmasked at update pins the clip fraction at
1.0, masked at update gives log pi = -inf and NaN. Both are in the alarm table.
8.4 Initialisation and the no-op bias¶
Orthogonal, gain sqrt(2) on every hidden convolution and linear, gain 0.01 on the pointer head's
q and qb linears and on the no-op linear, gain 1.0 on the value head's last linear, zero bias
everywhere, Embedding.normal_(0, 0.02). Every draw comes from the torch generator seeded from
torch/init, so the initial weights are a function of master_seed alone.
With a near-zero head the opening policy is near-uniform over the roughly 691 legal actions, so the
agent spends its elixir immediately at random tiles and the mask then forces it to wait. That
oscillation is a good exploratory start and it is self-limiting, so noop_bias defaults to 0.0. If it
is ever needed, the bias that produces a target p on the no-op with n legal actions is
9. The experience buffer, GAE and the PPO update¶
9.1 Buffer layout¶
One shared-memory block of (T + k) x R rows of row_bytes, cycle-major, with k = obs.frame_stack:
Cycle T holds observations only. It is the bootstrap row, and its scalars are unused. The k - 1
rows below cycle 0 are history: at the end of every iteration the last k - 1 cycles are copied down
into them, so cycle 0 of the next iteration has a full stack and an iteration boundary is not a
discontinuity in what the policy sees. At k = 1 there are none and the block is (T+1) x R. At the
laptop profile that is 229 x 192 x 14 163 B = 623 MB (95 cards, 2026-09-22) [A].
Every row of the rectangle is stored, including the seats a frozen or scripted opponent played. That
costs 25% more memory than storing only learner rows and it buys a buffer index that is the slot
index, with no separate buffer_row map to get wrong and no scratch area for discarded rows. What
decides whether a cell reaches the update is trainable, not where it was written.
Frame stacking is a gather, not a second copy. Because the buffer is a rectangle indexed by cycle
and a slot's rows sit at a fixed stride of R in that index, the stack for cell (t, r) is rows
t, t-1, ..., t-k+1 of the same slot, assembled at unpack time in the same kernel that dequantises.
Nothing extra is stored and nothing is written twice. The worker's episode_end flag at t-1 says
where a previous row belongs to an earlier episode. There the stack is zero-filled from that point
back, so the first frame of an episode has a zero history and the policy is never shown the tail of
the battle before it. Only spatial and mask_planes are stacked; vector is the current frame's,
because elixir, hand and clock are already the present state and a stale copy of them is noise. A test
at k = 2 asserts that the ticks of two stacked frames are consecutive, reading info["tick"] off
the round, and that an episode's first cell stacks a zero.
The scalar columns live in the parent's own numpy arrays. All of GAE runs on them, on the GPU.
| column | dtype | shape | bytes at the laptop profile |
|---|---|---|---|
action |
int16 | (T, R) | 88 K |
log_prob, reward, advantage, ret, final_value |
float32 | (T, R) | 5 x 175 K |
value |
float32 | (T+1, R) | 176 K |
terminated, truncated, valid, trainable |
bool | (T, R) | 4 x 44 K |
group, deploy_status, episode_end |
int8 | (T, R) | 3 x 44 K |
tick |
int32 | (T, R) | 176 K |
| total | under 2 MB |
The buffer is cleared every iteration: each timestep is trained on exactly n_epochs times and then
discarded, which is what makes the KL and clip-fraction diagnostics mean what they say. One reference
learner ships a rolling window that is never cleared, so each timestep is in fact trained on about
three times across three consecutive updates with increasingly stale old_probs. That is a
deliberate mildly-off-policy design in its setting and would be an unexamined one here;
ppo.keep_previous_iterations exposes the behaviour at a linear memory cost for anyone who wants it.
9.2 Minibatching¶
def batches(self, batch_size, minibatch_size, epochs, rng_for_epoch):
flat = np.flatnonzero((self.trainable & self.valid).reshape(-1))
for e in range(epochs):
perm = flat[rng_for_epoch(e).permutation(flat.size)] # epochs never straddle batches
for start in range(0, perm.size, batch_size):
yield Batch(perm[start:start + batch_size], minibatch_size, self._gather)
def _gather(self, idx): # one minibatch
idx = np.sort(idx) # sorted: locality
stage = self._pinned.next() # a ring of 4 pinned slabs
np.take(self._obs_view, idx, axis=0, out=stage.obs) # one gather, shared -> pinned
stage.to_device(non_blocking=True)
return self._codec.unpack_to_device(stage, self._statics)
Three rules, each a fix to something the references do:
- Gather per minibatch, never per batch. Fancy-indexing a whole batch materialises a 2.6 GB temporary at this row size.
- Pinned staging with non-blocking copies, so the transfer of minibatch
i+1overlaps the compute of minibatchi. The ring is four slabs ofminibatch_size x row_bytes, about 14 MB at the laptop profile's minibatch of 256, allocated once. - Nothing is dropped. Both references do
while start + batch_size <= totaland silently discard the remainder. Here the remainder is a smaller final batch, weighted by its true sample count;health/samples_unused_fracis in the metric stream and reads 0.
At the laptop profile: 32 832 trainable samples, batch_size = 4096, n_epochs = 3, so 8 batches
per epoch and 24 optimizer steps per iteration, each from 16 accumulated minibatches of 256.
9.3 GAE¶
learn/gae.py, vectorised over slots, one backward loop of T steps over (R,) vectors, on the GPU:
ended = terminated | truncated # the episode ended AT t
boot = torch.where(terminated, zeros,
torch.where(truncated, final_values, values[1:])) # (T, R)
adv = torch.zeros(R); out = torch.empty(T, R)
for t in reversed(range(T)):
delta = rew_scaled[t] + gamma * boot[t] - values[t]
adv = delta + gamma * lam * (~ended[t]) * adv
out[t] = adv
returns = out + values[:T]
- A terminated cell bootstraps from 0. A truncated cell bootstraps from
final_value. A cell that is neither, including the last cycle of the iteration for a row still mid-episode, bootstraps fromvalues[t+1], which is the value of the true next observation, because SAME_STEP autoreset only replaces a row when its episode ended and an episode that ended is flagged. This is the correct treatment and it is where both older references go wrong: one bootstraps a truncation off an unrelated episode's first state, and one self-bootstraps fromV(s_T). ~ended[t]breaks the recursion at every episode end, terminated or truncated alike. Using the terminated flag alone would leak the next episode's advantage backwards across a truncation.- Truncated cells need
final_obs. The worker packsinfos["final_obs"]for those rows into thefinalsarea of its control segment and the parent runs the critic on them intofinal_value. With the shipped config there is noTruncationCondition, so the path fires only for dead-worker rows; it is exercised and tested regardless. - A dead worker's rows fall out without a special case. For a worker that died at cycle
t*:valid[t, r] = (t <= t* - 2),truncated[t* - 2, r] = True,final_value[t* - 2, r] = values[t* - 1, r], and fort > t* - 2the rewards and values are zeroed andendedis set, so the loop runs over them inertly. learn/gae.reference_gaeis the per-slot python implementation.tests/test_gae.pyasserts the two agree to 1e-6 over random inputs including every boundary case. The reference learner ships the same pair and no test; here the test is the point of the pair.
credit_horizon_seconds = decision_ms / 1000 / (1 - gamma * gae_lambda) is computed, printed at
start-up and logged every iteration. At the defaults it is 45.5 s. It is the single highest-leverage
number in the configuration and it must never be set to ten seconds by accident.
9.4 Reward scaling¶
learn/returns.py, WelfordReturnScaler. Divide rewards by the running standard deviation of
unstandardised, unclipped discounted returns, then clip:
raw = discounted_cumsum_per_episode(rewards, gamma) # before any scaling
self.welford.update(raw[trainable & valid]) # {mean, count, m2}, n-1 variance
scale = self.welford.std if (standardize and self.welford.count > 1) else 1.0
rew_scaled = (rewards / scale).clamp(-reward_clip, reward_clip)
The mean is never subtracted: subtracting it changes the sign structure of a zero-sum reward. The Welford triple is checkpointed. This matters more here than in the references because gamma approaching 0.999 inflates return magnitudes roughly tenfold over gamma = 0.99.
Observations are not normalised. They are bounded by construction, Welford would amplify one-hot noise, and it would add cross-worker mutable state that a resume has to restore exactly. Fixed per-channel divisors plus the stem's GroupNorm do the same job with no state.
9.5 The update¶
learn/ppo.py, PPOUpdate.step:
def step(self, buffer, sched) -> UpdateResult:
values = chunked_critic_pass(buffer, chunk=cfg.critic_chunk) # (T+1, R), bf16, no_grad
buffer.set_values(values)
buffer.set_final_values(*critic_on_final_obs(buffer)) # only where truncated
adv, ret, stats = self.gae.compute(..., gamma=sched.gamma, lam=cfg.gae_lambda)
m = buffer.trainable_mask() & buffer.valid_mask()
if cfg.advantage_standardization: # ONCE per iteration
adv = (adv - adv[m].mean()) / (adv[m].std() + 1e-8)
buffer.set_advantages(adv, ret)
before_a, before_c = params_to_vector(actor), params_to_vector(critic)
for batch in buffer.batches(cfg.batch_size, cfg.minibatch_size, cfg.n_epochs, rng_for_epoch):
for opt in self.optimizers:
opt.zero_grad(set_to_none=True)
for mb in batch:
w = mb.size / batch.size # accumulation weight
bp = actor_critic.backprop(mb.obs, mb.action)
logp = bp.log_probs # fp32
ratio = torch.exp(logp - mb.log_prob)
surr = torch.min(ratio * mb.adv,
ratio.clamp(1 - cfg.clip_range, 1 + cfg.clip_range) * mb.adv)
dual = torch.where(mb.adv < 0,
torch.max(surr, cfg.dual_clip_c * mb.adv), surr)
policy_loss = -dual.mean() * w
value_loss = cfg.vf_coef * F.mse_loss(bp.values, mb.ret) * w
entropy_loss = -(sched.ent_coef * bp.entropy.mean()
+ sched.ent_coef_noop * bp.noop_entropy.mean()) * w
(policy_loss + value_loss + entropy_loss).backward()
self._accumulate_diagnostics(ratio, logp, bp, mb, w) # under no_grad
self._record_grad_norms() # BEFORE clipping
clip_grad_norm_(actor.parameters(), cfg.max_grad_norm) # separately
clip_grad_norm_(critic.parameters(), cfg.max_grad_norm) # separately
for opt in self.optimizers:
opt.step()
torch.cuda.synchronize(device) # ONCE, then every .item()
return UpdateResult(...)
The loss, written out, for one minibatch of n samples in a batch of N:
r_i = exp( log pi_new(a_i | s_i) - log pi_old(a_i | s_i) )
surr_i = min( r_i * A_i , clip(r_i, 1 - eps, 1 + eps) * A_i )
L_i = surr_i if A_i >= 0
max( surr_i , c * A_i ) if A_i < 0 # dual clip, c = 3.0
L_policy = - (1/n) sum_i L_i * (n/N)
L_value = vf_coef * (1/n) sum_i (V(s_i) - G_i)^2 * (n/N)
L_entropy = - ( ent_coef * (1/n) sum_i H_i
+ ent_coef_noop * (1/n) sum_i H2_i ) * (n/N)
L = L_policy + L_value + L_entropy
with H_i the entropy of the masked categorical and H2_i the binary entropy of p(no-op) against
p(play), read off the same normalised log-probs at no cost. The actor and the critic have disjoint
parameters, so summing one loss and calling backward() once is safe and the entropy term's gradient
into critic parameters is exactly zero. That has its own test.
ent_coef_noop is a warm-start guard and not a permanent term of the objective. H2(p_noop) is
maximised at p_noop = 0.5 while a healthy Clash policy plays about 22 cards in a match and sits near
p_noop = 0.94, so a coefficient that never decays pulls every converged policy toward overplaying.
It is a term that pays for a behaviour the objective does not want. It therefore anneals to zero over
its schedule; the schedule is in the run identity and run/ent_coef_noop is logged every iteration so
the two regimes are never confused in a plot. The noop_collapse* and noop_entropy_floor alarms are
untouched by this and keep watching after the coefficient reaches zero, and whether the guard was
needed at all is one of the things the first run measures (section 18).
The n/N weighting with one optimizer.step() per batch is what makes minibatch_size a pure
memory knob: the optimisation is identical whatever it is set to, which tests/test_ppo.py proves by
comparing the accumulated gradient against a full-batch one. On a 4 GB GPU that is not a nicety; it
is the mechanism that makes the run possible.
Diagnostics, all under no_grad:
kl = mean( (exp(log r) - 1) - log r ) # Schulman k3: non-negative, low variance
clip_frac = mean( |r - 1| > eps )
dual_frac = mean( (A < 0) & (c*A > surr) )
ev = 1 - Var(G - V) / max(Var(G), 1e-8)
ent_norm = mean( H / log(max(n_legal, 2)) )
grad_norm_actor, grad_norm_critic # pre-clip
update_magnitude_actor = || theta_a_before - theta_a_after ||
update_magnitude_critic = || theta_c_before - theta_c_after ||
kl and clip_frac are recorded per epoch as well as per iteration, because "epoch three's clip
fraction is more than twice epoch one's" is the concrete signal to lower n_epochs.
The learning-rate backoff. If kl > lr_backoff.kl_threshold for patience consecutive
iterations, multiply both learning rates by factor, floor at lr_min, reset the counter and log an
lr_backoff event. The counter and the current rates are checkpointed. One reference computes the KL
and acts on nothing; four lines prevent the class of blow-up that costs a whole run.
Explicitly not included: value-function clipping (the published studies find it
neutral-to-harmful, and one reference omits it with a comment saying so), observation normalisation
(section 9.4), any learning-rate schedule beyond the backoff, and torch.compile (measure it first,
on a machine where it is not a sixty-second warm-up on every run).
Checkpoint loading uses torch.load(..., map_location=device, weights_only=True). Without
map_location a CPU-only resume fails; without weights_only every checkpoint is an execution
surface. The optimizer factory's keyword arguments are restored over a loaded state dict, so
changing the learning rate on a resume actually takes effect.
9.6 The ratio invariant¶
At epoch 0, minibatch 0, before any optimizer step, ratio should be 1.0 for every sample: the
parameters are unchanged since the rollout and the input bytes are identical, because the policy acted
on decode(encode(obs)) (D5).
if self._check_ratio_now(iteration):
dev = (ratio - 1).abs().max().item()
assert dev <= cfg.ratio_atol[precision], _ratio_message(dev, ratio, mb)
The tolerance is a property of the precision, not a fudge factor. In fp32 the two forwards agree
to 1e-4. Under bf16 autocast they do not: the rollout forward is one policy group of roughly fifty
to a hundred and fifty rows and the update forward is a minibatch of 256 on the laptop, cuDNN picks an
algorithm per shape, and bf16's three significant digits put the resulting fp32 logits about 1e-2
apart (section 8.2). ratio_atol is therefore 1e-4 with autocast off and 2e-2 under bf16, and
royalelearn bench prints the actual ratio_max_abs_dev on the machine it is run on, beside the
tolerance, so a user can see the margin rather than trust it. It is one reading: the last timed
iteration's, taken on the first minibatch whose actor ran, not a maximum over several rounds.
AND THE TOLERANCE IS NOT A PROPERTY OF THE PRECISION ALONE. It is a property of the precision AND
of how peaked the policy is, and a run at net.noop_bias 8.0 died at iteration 6 on 2026-09-23
because nobody had noticed the second half. The arithmetic is one line:
An error in the LARGEST logit is multiplied into every other action's log-probability by the
probability that logit holds. At noop_bias 0 the no-op holds about 1/450 of the mass and bf16's
error is invisible -- 383 iterations across three runs never came near the guard. At noop_bias 8.0
it holds 0.87, and bf16's spacing at magnitude 8 is 0.0625, so the same arithmetic lands on 0.02.
Measured that night, same engine, same bias, six iterations each:
worst ratio_max_abs_dev |
tolerance | ||
|---|---|---|---|
| bfloat16 | 0.0261 | 0.02 | died at iteration 6 |
| float32 | 9.54e-07 | 1e-4 | healthy, and reproduced independently by two sessions |
The deviation is a TAIL quantity and the spread is the signature of quantisation, not of luck. The bf16 iterations ran 6.1e-05, 2.3e-03, 7.0e-03, 1.9e-02, 2.6e-02 -- a 428x spread -- because two forwards whose pre-rounding values differ by a hair land on the same grid point most of the time and on adjacent points occasionally. Mostly float noise, sometimes a full grid step. So the criterion is TOLERANCE OVER WORST CASE, and a typical value tells you nothing.
_ratio_precision_gate in rollout/preflight.py computes p_max * ulp(magnitude) from
net.noop_bias, the action space and net.autocast_dtype, prints it beside ratio_atol, and
REFUSES a run that exceeds it. It is an upper bound: within 15% on float32 and about 2x
conservative on bf16, which is the right direction for a gate.
Two things that follow and are worth knowing before choosing a dtype. Moving the bias out of
the reduced-precision sum -- promoting before the + noop_bias rather than after -- lowers the
worst case about 20x, to a 4.5x margin, and does NOT help for a long run: bf16's spacing doubles at
every power of two, the head's own output grows during training (actor.head.query.weight went
0.22 to 1.82 over 147 iterations), and a 4.5x margin absorbs about two doublings. float32 has
sixteen more mantissa bits and absorbs about 65,000x of growth. And float32 costs 2.3x the wall
clock, 62.2 s an iteration against 26.9, while using LESS memory: vram_peak 1,882 MB against
2,011. Nobody predicted the memory going the way it did.
What 2e-2 still detects is the whole reason the check exists, because every failure it is aimed at
produces a deviation of order one, not of order 1e-2:
- a rollout/update mask mismatch makes a masked-out action's log-prob
finfo.min, so the ratio is either zero or astronomically large; - a codec or codec-table mismatch feeds the update a different observation, and the log-prob of a 2305-way categorical moves by far more than a percent;
- a weight-version mismatch, where the policy that acted is not the policy being updated, moves
log-probs by the size of a PPO update, which is what
ppo/update_magnitude_actormeasures and is orders above the tolerance.
Kernel-level nondeterminism is not what this check is for, and at 2e-2 it will not see it. That is
the correct division of labour: reproducibility is determinism.tier's job, and tier T2 keeps bf16
precisely because deterministic kernels are bit-reproducible run to run at any precision. T2 promises
that two runs agree, never that two differently shaped forwards within one run do.
The check runs for the first ppo.debug_assert_iterations iterations of a run and every
check_ratio_invariant_every thereafter, and ppo/ratio_max_abs_dev is logged always. A violation's
message names the worst ten samples with their slot, cycle and episode ordinal, and lists the three
causes above in the order they are worth checking. No reference implementation has this check, and it
is the cheapest possible detector for the failure mode that is otherwise silent and slow-acting.
10. Rewards¶
royalelearn/rewards.py ships four RewardFunction implementations built on RoyaleGym's ABC and
composed with PotentialCombinedReward, which is RoyaleGym's CombinedReward with the discount
forwarded to every term that takes one. Nothing in RoyaleGym is modified; these are additions that
live here until they are absorbed upstream (section 16, ask 4).
def default_potential_reward(*, crown=0.2, tower_hp=0.1, elixir=0.05) -> CombinedReward:
return PotentialCombinedReward([
(WinLossReward(draw=0.0), 1.0), # royalegym's; the objective
(PotentialCrownReward(), 0.2),
(PotentialTowerHPReward(), 0.1),
(CommittedElixirPotential(scale=10.0), 0.05),
])
Each potential term returns gamma * Phi(s') - Phi(s), with Phi computed from the state:
| term | potential |
|---|---|
PotentialCrownReward |
(own_crowns - foe_crowns) / 3 |
PotentialTowerHPReward |
(sum_s own_hp[s]/own_max[s] - sum_s foe_hp[s]/foe_max[s]) / 3 |
CommittedElixirPotential |
((own_bar + own_board) - (foe_bar + foe_board)) / scale, where bar is elixir_milli / 1000 and board is the sum of Fraction(card.elixir, card.count) over that player's live non-tower entities |
This is the reward that ships. default_env_spec, and so all three profiles, and
examples/configs/laptop.json, smoke.json and workstation.json name
royalelearn.rewards.default_potential_reward since a89570d, and
tests/test_rewards.py::test_every_shipped_config_trains_against_the_potential_reward holds each
profile and each of those files to it. Before a89570d every shipped config named RoyaleGym's
royalegym.reward.default_reward, so every real iteration before it trained a different objective.
The reward is part of the environment spec's digest, which is in the run identity and is the first
input of the ladder's context digest (section 11.6). So runs from before a89570d have a different
identity and context and never pool with later ones. That separation is intended.
gamma comes from the schedule. The coordinator builds one ScheduleState per iteration and hands the
same one to collection and to the update, and every Step command carries its gamma in the control
word. On the first step whose gamma differs from the last, ShardRunner.set_gamma passes it through
the term recorder to PotentialCombinedReward.set_gamma, which walks the composition and gives it to
every PotentialReward. A composition built with RoyaleGym's plain CombinedReward does not forward
it, and its potential terms keep a discount of 1. So the reward the worker applies and the discount
the learner's GAE uses are the same number, and
tests/test_rewards.py::test_the_schedule_s_discount_reaches_the_reward_during_collection checks it
end to end: it drives the real coordinator, inline and through worker processes, with a probe term
that pays exactly gamma - 1 under a schedule that steps between iterations, and every round carries
the discount its iteration's schedule gave the learner.
The reason this composition and not RoyaleGym's shipped default_reward():
- The terminal win/loss term is the only one that is the objective. Everything else is shaping, and
shaping that is not a difference of a potential can change which policy is optimal (Ng, Harada and
Russell, 1999). RoyaleGym's
reward.pydocstring already states this principle. TowerHPRewardcomputesPhi(s') - Phi(s)rather thangamma*Phi(s') - Phi(s). It is onegammaaway from being exactly policy-invariant, and fixing it means the term never needs annealing.ElixirTradeRewardis not a potential and it rewards turtling: an agent that never plays a card never incurs the negative term while enemy units still die to its towers and earn the positive one. At weight 0.02 over about forty trades a match that is 0.8, comparable to the plus-or-minus-one terminal reward. The shaping can outweigh the objective.ElixirLeakPenaltyis not zero-sum: both players can leak at once. That is the signature of a term standing in for a missing potential.- The committed-elixir potential replaces both. Playing a unit card moves elixir from the bar to the
board and is net zero. A spell is charged at the tap. A Mirror play loses exactly its own one
elixir: its copy is one level up, and a unit is matched against its card's row at the level the
engine reports for it (see
CommittedElixirPotentialinroyalelearn/rewards.py). A Tri Wizards play gains 7: two of its three wizards come down as their own cards' units, so the play prices at - Losing a unit costs; killing gains; and sitting at ten elixir is penalised automatically,
because the opponent's potential keeps rising while yours does not. There is no coefficient to
re-tune and no annealing schedule, which is what
royalegym/reward.py's own house rule asks for: weights should settle, not drift.
Unit values use exact Fraction(card.elixir, card.count) arithmetic, as RoyaleGym's own elixir term
does, because the seat-mirror antisymmetry test depends on it: a running float sum of the same values
returns plus and minus 1.1e-19 on a perfect mirror instead of zero, which is harmless to a gradient
and makes the property uncheckable.
Every weighted term's per-episode sum, per seat, is on the EpisodeRecord and in the metric stream as
env/reward_terms/<name>, a signed mean, with env/reward_terms_abs/<name>, the mean of each seat's
magnitude, beside it. <name> is the term's class name, because CombinedReward files each term under
type(term).__name__, with one exception: PotentialCombinedReward says which of its terms is the
objective (terminal=0 by default) and files that one under metrics.records.TERMINAL_REWARD_TERM,
which is terminal, so a class rename cannot move an alarm's arithmetic. The shipped composition
therefore reports terminal, PotentialCrownReward, PotentialTowerHPReward and
CommittedElixirPotential. The acceptance criterion, checked by the shaping_dominates alarm, is that
the sum of the absolute shaping terms stays below the terminal term's magnitude.
What that alarm means under a POTENTIAL reward, which is not what it meant under the other one.
A potential term pays gamma * Phi(s') - Phi(s) every step. DISCOUNTED over an episode that
telescopes to gamma^T * Phi(s_T) - Phi(s_0), and with the terminal potential taken as zero
(fc8b53a) only -Phi(s_0) survives. The shares are UNDISCOUNTED sums, though, because that is what
the recorder adds up, and for those the telescoping leaves one more piece:
sum_t F_t = -Phi(s_0) - (1 - gamma) * sum_{t=1}^{T-1} Phi(s_t)
At a level start Phi(s_0) is zero, so a share is 1 - gamma times how far the potential wandered.
test_shaping_strength.py holds both forms exactly. That makes the shares the right instrument for
what this alarm checks, and the wrong one for how loud the shaping is. They fall as the discount
schedule rises, whatever the weights are, while the share divided by 1 - gamma holds. On identical
random-legal battles, moving gamma from 0.997 to 0.999 cut it from 0.0462 to 0.0185 [M].
Shares like these say every term still telescoped. They do NOT say how large the shaping is against
the objective, and read that way they invite raising weights that are already loud. How loud a term
is lives in env/reward_terms_step_abs/<term>, each seat's sum |F_t| over an episode. At the
shipped weights, on a 100-card RustEngine environment, over ten random-legal battles at gamma 0.999,
per seat per episode [M]: CommittedElixirPotential 0.80, PotentialTowerHPReward 0.15,
PotentialCrownReward 0.15, terminal 0.80. The shaping together is about 1.4 times as loud as the
objective, and the elixir term alone is about as loud. No threshold is claimed for that ratio. It is
published so a weight is changed against the number that describes it.
The three shaping weights are keyword arguments of default_potential_reward (crown, tower_hp,
elixir), finite and not negative, with the shipped values as defaults. Each term's
reward_terms_step_abs is linear in its weight.
So the alarm is not a weight check any more, and reading it as one would make it look vestigial. It now says: a shaping term has stopped telescoping. That is the defect fc8b53a fixed, where a finished battle still had a potential and the shaping paid for the margin of a win, and it is the defect a new term that is not a difference of a potential would have. A run whose shaping share climbs towards its terminal term has a term that is no longer policy-invariant, whatever its weight says.
The two shares that alarm reads, env/reward_shaping_abs and env/reward_terminal_abs, are sums of
those per-seat magnitudes. The signed mean is kept because it is how a term that is not antisymmetric
shows itself. Both shares were wrong until 5724eb5 and dc7f7a1: the absolute value was taken on the
mean over both seats, where every zero-sum term cancels, and the objective arrived under its class name
and was summed with the shaping. On every metric row the development machine recorded before then
(161 rows) both shares read exactly 0.0 [M], so the alarm compared two structural zeros. It has
not yet been seen on a real run since.
Closed 2026-09-22 (fc8b53a): the potential on the terminating step. PotentialReward takes
Phi(s') = 0 when state.game_over and keeps the position's potential on a truncation, where the
learner bootstraps from the final observation's value. Before that commit the terminating step paid
gamma * Phi(s') more than potential-based shaping would, a bonus that depended on how the game
ended and could move the optimum. test_a_potential_is_zero_once_the_battle_is_over and
test_a_truncated_battle_keeps_the_potential_of_the_position hold the two cases. This entry said
"open" for two days after the fix landed, two paragraphs below a sentence citing the fix.
11. The ladder¶
11.1 What lives where¶
royalegym.selfplay keeps what is bookkeeping over ids and results: Opponent, PolicySnapshot,
OpponentPool, the sampling strategies, pfsp_weights, save/load, and the streaming Elo. It
keeps them because it needs neither torch nor an env. RoyaleLearn owns the weights, the routing, the
result log, the fit, the evaluation runner and the gate. ladder/pool.py's LadderPool wraps an
OpponentPool, a ResultLog and a SnapshotStore.
OpponentPool.record_result and _evict are not called from this repo. The pool's aggregate record
cannot represent draws separately: wins: float with a draw adding 0.5 makes five wins and five
losses indistinguishable from ten draws, and those carry completely different variance. And _evict
drops the oldest snapshot, which is exactly the diversity the anti-forgetting floor exists to
protect. ladder/results.py owns the result log and ladder/eviction.py owns eviction until those
are fixed upstream (section 16, ask 4).
11.2 Rating¶
Layer 1, the readout. EloReadout, RoyaleGym's zero-sum logistic Elo at k_factor = 32. It goes
on the dashboard between refits and nowhere else. It is order-dependent, which under parallel workers
is nondeterministic, and a fixed k keeps moving a frozen player whose strength is constant.
Layer 2, authoritative. BradleyTerryDavidsonRater, a MAP fit over the whole stored evaluation
result matrix:
q_i = 10 ** (r_i / 400)
P(i wins) = q_i / (q_i + q_j + nu * sqrt(q_i * q_j))
P(draw) = nu * sqrt(q_i * q_j) / (q_i + q_j + nu * sqrt(q_i * q_j))
P(j wins) = q_j / (q_i + q_j + nu * sqrt(q_i * q_j))
prior = r_i ~ Normal(0, prior_sd^2), prior_sd = 400 Elo
gauge = r["scripted:noop"] pinned at 0
The log-posterior is concave, so Newton/IRLS converges in fewer than ten steps on a pool of a few
hundred, in milliseconds. Standard errors come from the diagonal of the inverse observed Fisher
information, the Hessian of the negative log-posterior. A difference between two players uses the
corresponding 2x2 block, never the sum of the two marginals. nu is fitted jointly; if the measured
draw rate is under 2% the config falls back to draws = "half_win" and the run's ladder.json
records that it did.
The property that makes this the authoritative layer: the rating is a pure function of the stored
result matrix. Same games in, same numbers out, in any order, on any machine. That is the same
property RoyaleSim gives for battles and it is what docs/design.md's determinism commitment requires
of this layer.
TrueSkill is rejected explicitly. It is online only, so it cannot refit history; its sigma shrinks monotonically and needs an artificial floor to stop the rating freezing; and at the floor the converged interval is about plus or minus 92 Elo, worse than a hundred games of direct head-to-head. Every one of those compromises exists because, in the references' game, an evaluation game is a real match in a real client. Here an evaluation battle costs a fraction of a second and is seed reproducible, so the compromises should not be inherited along with the mechanism.
rating_above_v0 is reported alongside the anchored rating, so the headline curve reads as "Elo above
the initial random policy" while the gauge itself is a permanent scripted anchor that exists before
any snapshot does.
11.3 Matchmaking¶
MixMatchmaker.assign(battle, ordinal, pool, ratings) is the unit of matchmaking: it answers what one
battle's next episode is, and it is called by the parent at that battle's own episode boundary, from
match/battle/{b}/ordinal/{k}. plan() is the same function applied across the geometry to produce
the iteration's opening table. A policy is therefore bound to a battle for exactly one episode, and a
partially controlled trajectory cannot occur.
Two decisions go into an assignment, and they are made differently.
The role is the battle slot's, for the whole run. The mixture is a count of slots, not a
probability each slot is drawn from. Of the laptop profile's 96 battles, exactly 48 are mirror slots,
34 pool and 14 scripted, and those numbers do not move for the life of the run. config.role_counts
turns the mixture into the three counts; the matchmaker lays them over the battle indices with a
seeded permutation from match/battle/-2/ordinal/{n_battles}, so which battle gets which role is a
pure function of the master seed and the battle count, and nothing else: not the iteration, not the
ordinal, not the pool, not the ratings. A permutation rather than "the first 48 indices" because
battles map to workers and shards in contiguous blocks: a role pinned to low indices would put every
double-seat episode on the first worker and none on the last, and a shard-round is as long as its
slowest worker.
The opponent and the learner's seat are still one draw per episode, addressed by the battle and its ordinal. An iteration spans several episodes in every slot, so one draw per slot per iteration would give every episode that starts inside that iteration the same opponent. That is a correlation across the mixture that nothing in the statistics accounts for. And because that draw is addressed by the battle and its ordinal rather than by the iteration, it does not move when the iteration length changes: the same episode of the same battle meets the same opponent however the iteration boundary falls across it.
| bucket | battle slots (of 96) | opponent | trainable seats |
|---|---|---|---|
| mirror | 48 (mix[0] = 0.50) |
the live policy on both seats | 2 |
| pool | 34 (mix[1] = 0.35) |
one of at most max_resident_opponents frozen snapshots, PFSP-weighted |
1 |
| scripted | 14 (mix[2] = 0.15) |
uniform over ladder.scripted_opponents, run in the worker. The default is the two anchors, NoopOpponent and RandomLegalOpponent(0.9) |
1 |
Trainable rows per cycle are therefore n_battles + mirror_battles, which is 96 + 48 = 144 of 192
slots, exactly three quarters, on every cycle of every iteration whatever the master seed.
config.geometry sizes the iteration from that number and ppo.timesteps_per_iteration counts kept
rows, so cycles × learner_rows ≥ timesteps_per_iteration is a fact about the rectangle rather than
about its average.
Rounding, when the shares do not divide the battles: mirror = round(mix[0] × n_battles), then the
remaining battles split between pool and scripted in the ratio mix[1] : mix[2] with pool rounded and
scripted taking the remainder. Halves round up rather than to even, because Python's round(0.5) is 0
and a one-battle rectangle under the shipped mixture would then contain no mirror at all. The three
counts always sum to n_battles, and each lands within one battle of its exact share. A share whose
quota is under half a battle rounds to zero, because three battles cannot hold a tenth of one. So the
mixture a small run actually plays is the count, not the config.
Why counts and not a draw. Under a per-episode draw the mirror count was Binomial(96, 0.5): a
standard deviation of 4.9 rows per cycle, against a laptop iteration whose margin over the
trainable-row floor (0.98 × 32 768 = 32 113 against a planned 32 832) is 0.2%. Measured over master
seeds 0–299, 76 of them sized an iteration below the floor. That is one in four, and
TrainableRowsShort halted the run in its first minutes. The shipped seed happened to draw 48 and
survive, which is why no recorded run showed it.
A battle share and a row share are not the same quantity. The mixture is a share of battles;
throughput/discarded_rows_frac is a share of rows, and a row is one slot on one cycle. A mirror
episode runs about 9% longer than a scripted one, so mirror battles occupy more of the rectangle's
cycles than their share of the battles, and a sample row share sat about 0.022 below the battle share
it was compared against. With the role fixed to the slot the question does not arise: every slot holds
its role for the whole iteration whether its episodes are long or short, so the learner's rows are
cycles × (n_battles + mirror_battles) and the discarded share is (n_battles - mirror_battles) /
(2 × n_battles) exactly. The iteration invariant of section 14.1 still compares the two within
MIXTURE_TOLERANCE; it can become an equality, short only of rows lost to a dead worker.
Landed in d83027a. Before it, the role was drawn per episode from mix and the iteration was sized
from the expectation. The shipped profiles do not change size. They are still 144, 1 152 and 2 304
learner rows a cycle, as the table in section 2.2 always said, but they are now that size on every
seed rather than on average: measured over master seeds 0–299, 76 iterations short of the floor
became 0.
MixMatchmaker.FORMAT_VERSION is 2, which adds n_battles to the checkpoint; a format-1 checkpoint
still loads and recovers its rectangle at the resumed run's first plan. The one property given up is
that a run replayed at a different worker count lays its roles out differently, since exactly half of
96 battles and exactly half of 192 are not the same set of indices.
- PFSP weighting is
hard, power 2:w_kproportional to(1 - p_k)^2, withp_kthe fitted model's predicted score probability rather than the direct head-to-head record. On a pool of 48 the learner has played few pairs directly andwin_rate's Beta(1,1) prior reads exactly 0.5 for the rest, which collapseshardto uniform and discards the information that the learner crushes v40 and v40 crushes v55. - The learner's rating is its newest admitted snapshot's (2026-09-27).
p_kneeds a fitted rating for the learner, and the live policy never plays an evaluation game under its own id: a gate rates snapshots, and the probe's games arekind="probe", which the fit does not read. Until 2026-09-27 the matchmaker looked up the idlearnerin the fitted table, never found it, and readp_k = 0.5for every candidate, sohardwas uniform in every run, for the residency draw and the per-episode draw alike. The unit tests had supplied alearnerrating by hand, so none of them saw it. Now the newest admitted snapshot stands in for the learner: the pool member with the largest env step, which is the last copy of the learner a gate accepted.p_kis the fitted model's prediction for that member against candidate k, and 0.5 against itself. - The draw falls back to uniform, and says so. Before a gate has admitted a snapshot, or
while the fit has not yet rated it, there is no stand-in, and the weights are the uniform ones
the floor mixes in.
ladder/pfsp_effectiveis 1 on an iteration whose plan used a rating and 0 on one that fell back, andladder/pfsp_learner_stepgives the stand-in's env step, absent when there was none. An arm that means to train againsthardchecks the first key. - The stand-in lags the learner by up to
candidate_every_env_steps: it is the learner as of its last admitted snapshot. The same lag is what keeps the draw fixed inside an iteration, since the stand-in changes only when a snapshot is admitted or a refit lands, both at an iteration boundary. - It is a behaviour change, not a fix that leaves numbers alone. Every run whose pool has more
candidates than
max_resident_opponents, or plays any pool episode once a snapshot is rated, now draws different opponents than it did. - Mixed with a uniform floor:
0.8 * hard + 0.2 * uniform, and every weight floored atweight_floor_scale / Mso no snapshot becomes unreachable. Without a floor the population's tail is forgotten, which is the job AlphaStar's forgotten-players slice does. hardfor training,variancefor evaluation. They serve different objectives: training wants opponents that still beat the learner, measurement wants the most even matchup because that is where a game carries the most information. Both weightings are already implemented inroyalegym.selfplay.pfsp_weights.- The learner's seat is drawn uniformly between blue and red for every pool and scripted battle,
redrawn at every episode boundary, so
env/win_rate_by_seatmeasures the seat advantage the learner actually experiences. Redrawing the seat costs nothing in exactness, because a battle contributes one learner row from whichever seat it takes. That is why the seat stayed a draw when the role stopped being one. The shipped engine is not a 180-degree rotation between seats: multi-unit ground deploys land differently on the two sides, so a seat advantage of a few points can be the game's own and is a warning, not a leak. - A pool slot plays scripted while no snapshot is resident. Before the first candidate is admitted
there is nothing to draw from, and a scripted opponent is the honest substitute: it fills the same
one seat, so the iteration is still exactly the size it was planned at. It draws from the same
ladder.scripted_opponentslist as a scripted slot. - The scripted share trains against
ladder.scripted_opponents. The list defaults to the two anchors, which is what every run before the field played:noopnever plays a card andrandom_legalpasses nine decisions in ten and otherwise plays a random legal move.scripted:pushplays a card as soon as one is playable, as far up the board as the rules allow, andscripted:patientdoes the same once three of its cards are playable. Naming either is how the scripted share meets an opponent that places forward on purpose. The list changes training only. The gate's anchors, the rating anchor and the probe's default rungs stay as they were. The ladder section is part of the run identity (ladder_digest), so adding the field moved every identity, and the default keeps every run's behaviour. - Evaluation is not in this mixture at all (D10). The studied alternative folds evaluation into the rollout worker at a small probability and swaps the live match object in place; that makes "win rate" depend on the training curriculum and, in the reference, leaves a worker permanently misconfigured if anything raises between the swap and the restore.
Deck protocol. Strength in Clash is a function of the deck pair, and deck matchups are the main
source of non-transitivity. The v1 ladder is defined on one declared deck protocol, recorded in the
context string on every result, and results from different contexts are never pooled. Deck
generalisation is a later axis, not a v1 confound.
11.4 Snapshot cadence and the gate¶
Cadence is measured in environment steps, not learner iterations: iteration size changes when the
batch size or the worker count changes, and a cadence in iterations silently changes meaning. A
candidate every candidate_every_env_steps (4 000 000 game-steps at the laptop profile, about two
hours).
Every candidate is evaluated; only those that pass enter the pool. A failed candidate is discarded,
not retried, and its results stay in the log. A run of consecutive failures is exactly the plateau
signal worth having, and gate_starved alarms on five in a row. Independently of the gate a
checkpoint is written at every candidate: checkpoints are for resuming, pool membership is for
rating, and conflating the two is what makes a pool unbounded.
WilsonGate admits candidate C if and only if all three conditions hold:
- It beats the champion. Over
n = 1000paired battles (500 seeds x 2 side assignments), the Wilson 95% lower bound onC's score rate is at least 0.52. In practice that requires an observed rate of at least 55.2%, about 35 Elo. A bound at 0.50 would admit anything merely not worse and the pool would fill with lateral moves. - No regression against the anchors. Score rate against
NoopOpponentandRandomLegalOpponent(0.9), 200 battles each, is not more thangate.anchor_tolerance_pp = 2percentage points below the champion's. This catches the specific failure where a policy beats its recent ancestors by exploiting a shared blind spot and loses basic competence. The champion's score is its RECORD against that anchor. A champion that came through a gate has one. One that did not -- the first snapshot of a run, admitted free -- plays the anchor once, at the same count on the same frozen seeds, and every later gate reads that record. Until 2026-09-24 such a champion fell back to the rater's prediction, which for a snapshot nobody had rated is the prior 0.5: the bound was 0.48, any policy cleared it, and a first real gate ran a regression check that could not fail. - No pool-wide collapse. The candidate's observed mean score against a
variance-weighted stratified sample of 8 pool snapshots, 100 battles each, is at least the champion's fitted mean predicted score against the same eight, less one standard error of the candidate's observed mean. The champion's side is the fit and not a record because the champion has not necessarily played those eight on those seeds, and a condition that silently skipped the snapshots it had no games against would be weakest exactly where the pool is most diverse. The rater's prediction is defined for every pair from the whole result matrix, and its use here is the same one the PFSP weights make of it (section 11.3). The cost stays at 800 battles. With no pool member besides the champion and the anchors there is nothing to sample and nothing to have collapsed against: the condition is recorded SKIPPED with that reason, and promotion rests on (1) and (2). It used to be recorded as passed with n=0, which is how a gate of 1,400 battles read as a full one of 2,200.
Outcomes: all three pass, admit and promote to champion. (1) and (2) pass and (3) fails, admit to the
pool but leave the champion unchanged, with meta["cycle"] = true. That is a useful diverse
opponent and a detected cycle, not progress. (1) fails, discard the candidate and keep the results.
A floor admits unconditionally every floor_admit_every_env_steps, so a plateau cannot starve the
pool.
Total 2200 battles per gate, all of which are rating evidence. Cost on the laptop profile: 2200
battles of about 420 decisions is 924 000 game-steps, which at 2 eval workers is about 6 to 7
minutes [A], against a candidate cadence of about two hours. That is roughly 5% of wall clock,
which is exactly why the gate can afford to be statistically honest. ladder/gate_seconds_frac is
logged. If it exceeds 8%, raise candidate_every_env_steps first: a pool does not need more than
twenty members a day. Only if that is not acceptable, lower n to 600 and raise the required
observed rate. Never drop the interval.
A GateDecision record is written to ladder/gates/<candidate>.json and pushed to the metrics sink
as an artifact. It holds the candidate, each condition's n, observed value, bound and verdict, the
eval seed set hash, the champion id and the wall time.
11.5 Evaluation¶
ladder/evaluate.py, EvalRunner. Four separations, each easy to get wrong:
- No experience is recorded. The runner writes results and nothing else. It drives a
RolloutSourcewhose plan marks every row non-trainable. - A different episode definition. A full match under real rules:
GameOverCondition, no truncation, the real match-start mutator. Training may later use randomised mid-game start states; evaluation never does. It is built from a separate env spec (config.eval_env), never by mutating the live env. A step cap does not make evaluation cheaper, it makes it uninformative: a cut match is scored on crowns that have mostly not been taken yet, so nearly every game is a draw and the rating barely moves (measured onMockEngine: eight of twelve games drawn at a 120-step cap [M]). That holds for the smoke configuration too, where the gate'snis reduced but its episode definition is not. - Separate RNG streams and a separate result table. The evaluation seeds come from a fixed set of
eval_seed_countseeds drawn once at run start fromeval/seed_set, written into the run directory, hashed intoladder_digest, and reused forever. Training results are recorded but taggedkind="train"and excluded from the fit by default: they are PFSP-selected and therefore biased toward hard matchups. - Paired, with common random numbers. Each pairing plays each seed twice with the sides
swapped. The unit of analysis is the seed, scored
x_s = (score as blue + score as red) / 2, taking values in{0, 0.25, 0.5, 0.75, 1}. This removes side bias exactly rather than averaging it away, and removes the start-state variance the two policies share. The interval is a bootstrap over seeds (gate.bootstrap_resamples = 10 000, seeded fromeval/bootstrap/{comparison_id}), not a binomial over games: the two games of one seed are correlated and treating them as independent overstatesnby up to a factor of two.
The empirical correlation between the two side assignments on a seed, rho, is measured and logged as
ladder/paired_rho -- for the last gate's CHAMPION comparison, which is the one whose effective sample
size decides the gate, carried on that condition's record. Where it is undefined (every seed decided
one way on a side) it is None and the key is absent; it used to be written as 0.0, which reads as
"uncorrelated". The effective sample size is about n / (1 + rho); at rho = 0 pairing costs
nothing and above it pairing wins. It is also worth knowing on its own: it says how much of a battle's
outcome the start state decides rather than the policies.
Evaluation uses the release sampling mode, one config field (ladder.release_mode, default
"stochastic"). Rating both the sampled and the argmax variant doubles the pool and the cost for no
decision value.
11.6 The result log, the snapshot archive and eviction¶
ladder/results.py. An append-only ladder/games.jsonl, one line per battle:
{"a": "snap:v18", "b": "snap:v17", "score_a": 1.0, "seed_index": 311, "side_a": "blue",
"context": "9f3c1a2b4d5e6f70", "kind": "eval", "run_id": "3f9a1c...", "iteration": 381,
"wall": "2026-09-22T04:11:07Z"}
{"a": "learner@12500000", "b": "scripted:noop", "score_a": 1.0, "seed_index": 7, "side_a": "red",
"context": "9f3c1a2b4d5e6f70", "kind": "probe", "run_id": "3f9a1c...", "iteration": 381,
"wall": "2026-09-22T04:12:55Z"}
The kind is what separates the three cuts of one file. eval is the authoritative one, the
only one the rating is fitted to. train is the battles the mixture played, recorded because
they are free to keep and excluded because they are PFSP-selected. probe is the live policy
against a fixed rung (section 11.9), excluded for a different reason: the player it names exists
for one moment.
context is
sha256(env_factory.digest() + obs_builder_spec.digest() + deck_protocol + engine_build_digest)[:16].
The obs builder's own ComponentSpec digest is in there because it carries Reveal: a policy that
was shown the opponent's hand and a policy that was not are not the same kind of player, and pooling
their results would make the ladder measure the information rather than the policy. For the same
reason EvalRunner refuses a pairing whose two spec.json obs_digests differ, naming both ids and
both digests rather than producing a number nobody can interpret.
The aggregate table is a derived cache and a test asserts rebuild(log) == aggregate. Writes are
append, flush and os.fsync per batch, so a crash truncates at most the last line and the reader
tolerates it.
ladder/snapshots.py, DiskSnapshotStore: content-addressed under snapshots/<sha256[:16]>/ holding
actor.safetensors (fp16, about 0.8 MB), and spec.json carrying arch_digest, obs_digest,
action_digest, codec_version, codec_table_digest, frame_stack, num_cards, vector_size, the
training step, the context and the run id. No pickled nn.Module: unpickling a module is
arbitrary code execution on load and one class rename invalidates the whole pool. A load whose
arch_digest, obs_digest or codec_table_digest disagrees with the current run is refused with a
message naming the field. A LoadedPolicyCache keeps at most
max_resident_opponents + 2 modules resident in VRAM.
Archive everything and never delete: at a megabyte each, five hundred snapshots is under a gigabyte.
HallOfFameEviction removes snapshots from the sampler only, and never a scripted anchor, never
v0, never a member of the champion chain. Among the rest it keeps a stratified sample across the
fitted rating range and prefers keeping snapshots flagged meta["cycle"], because those are the
diverse ones.
Evicted snapshots stay in the archive and in the result log and their ratings stay in the fit:
eviction is about sampling cost, not about forgetting evidence.
11.7 The statistics¶
Unpaired binomial half-widths at the worst case p = 0.5, with the Elo equivalent at
dp/dElo = ln(10)/1600 = 0.0014386:
| n battles | 95% half-width on score rate | equivalent Elo half-width |
|---|---|---|
| 100 | 9.8 pp | 68 |
| 200 | 6.9 pp | 48 |
| 400 | 4.9 pp | 34 |
| 1000 | 3.1 pp | 22 |
| 2000 | 2.2 pp | 15 |
| 10000 | 1.0 pp | 6.8 |
Score rate to Elo, d = 400 * log10(p / (1 - p)):
| score rate | 0.525 | 0.55 | 0.575 | 0.60 | 0.65 | 0.75 |
|---|---|---|---|---|---|---|
| Elo gap | 17 | 35 | 52 | 70 | 108 | 191 |
The Wilson score interval, z = 1.96:
centre = (p_hat + z^2/(2n)) / (1 + z^2/n)
half = z / (1 + z^2/n) * sqrt( p_hat*(1 - p_hat)/n + z^2/(4 n^2) )
lower = centre - half
| observed | n | Wilson lower bound | clears 0.52? |
|---|---|---|---|
| 55.0% | 1000 | 0.519 | just short |
| 55.2% | 1000 | 0.521 | yes |
| 57.5% | 1000 | 0.544 | yes |
| 55.0% | 400 | 0.501 | no |
A candidate must be about 35 Elo stronger to pass reliably: a true 35-Elo improvement passes with
roughly even odds per attempt and a true 70-Elo improvement passes essentially always, which is the
right shape for a gate applied every few million steps. With a non-trivial draw rate d the variance
of the score rate is (p(1-p) - d/4)/n, strictly less than the binomial, so the table is conservative
once draws are counted separately. That is one more reason to store them separately.
11.8 What the ladder logs¶
The fitted rating and its 95% interval for the learner and every pool member, with the anchor named;
rating_above_v0; the live policy's score rate against each scripted anchor, which is the one scale
that never drifts and which the probe of section 11.9 is what measures; every gate outcome with the condition that failed; the draw rate; paired_rho; the pool and
sampler sizes; eviction events; per-snapshot evaluation game counts; and the transitivity
residual, which is the fraction of head-to-head pairs with at least 30 games whose observed score
contradicts the fit by more than two standard errors.
The residual is the number that says whether a scalar rating is meaningful for this population. Above
about 10% the scalar release metric is lying and the Rater ABC is exactly the seam a Nash-averaging
or alpha-rank implementation drops into. The result log stores every game rather than an aggregate,
so it already holds the dense matrix such a method needs. Publishing the number is how
royalegym/selfplay.py's honest docstring caveat about non-transitivity becomes a measurement instead
of a warning, and it is far better discovered by a logged diagnostic in month one than by a confusing
ladder in month six.
11.9 The probe of the live policy¶
Everything above measures a frozen snapshot. The gate names snap:v{n} and hands it to the
evaluation runner, so the id learner never appears in an evaluation result and a run publishes
no score for the policy it is actually changing. ladder/score_vs_noop and
ladder/score_vs_random_legal asked the result log about a pair nothing could write, and read
0.5 in every row until they were made absent instead (72c1215).
The probe is that missing measurement. Every ladder.probe_every_iterations iterations, after
the update, the coordinator plays the live model against each of ladder.probe_opponents:
probe_games battles per rung, paired, through the same EvalRunner machinery as the gate.
Four things define it.
It names the policy and the moment. The id is learner@{cumulative_env_steps}, so the
result log says which weights were measured. A policy at 4 million steps and the same policy at
40 million are different players, and a single learner id would pool their games into one row
of the matrix.
Its games are not the fit. They are written kind="probe" and eval_view reads eval
only, so a probe moves no fitted rating, no gate condition, no anchor reference and not
ladder/eval_games_total. Two reasons, and both matter. The rungs never change and the learner
does, so a probe is a reading against a ruler rather than a game between two rated players; and
each probe is a one-moment id, so admitting them would grow the fit by a column per probe, each
too thinly played to place, and let the rating scale move because the run measured itself.
The positions are fixed. A comparison takes the first probe_games / 2 seeds of the
frozen evaluation set, never a fresh draw, so every probe of a run plays the same start states
and two of them differ by the policy rather than by the sample. The acting uniforms are not
shared: the stream is addressed by the comparison id, which carries the step, so each probe
draws its own action noise. The start state is the part worth controlling, and it is controlled.
It is off by default. probe_every_iterations = 0 is the shipped value, so no existing run
changes and nothing is paid until a run asks for it.
What it does not do: it is not the gate, it admits and promotes nothing, and it gives the
learner no fitted rating. So ladder/rating_above_v0 stays absent, and the PFSP weighting takes
the learner's newest admitted snapshot as its stand-in (section 11.3). Both want a rated player,
which is what a snapshot is; the probe deliberately makes a different trade.
What it publishes, on the iterations it runs and on no others: ladder/score_vs_noop and
ladder/score_vs_random_legal for the anchors, and for every rung
ladder/score_vs/{rung} with ladder/score_vs_n/{rung} and the two ends of its 95% bootstrap
interval over seeds. n is seeds, not battles: the two battles of a seed are correlated.
time/probe is what the iteration spent and ladder/probe_seconds_frac the share of wall clock
since the run started. A rung nobody probed has no key, because 0.5 is what an even contest also
looks like.
The cost. The battles are played one at a time in the parent process, so a probe is
len(probe_opponents) * probe_games full matches of the evaluation environment, and full means
untruncated. Measured on MockEngine, CPU, one seat driven by the live network:
a probe of the two anchors at probe_games = 40 -- 80 untruncated matches -- cost 95.5 s on
its first iteration and 100.0 s on its second, which is 1.25 s a battle [M]. Two
scripted seats on the same environment cost 0.54 s a battle [M], so about half of a probe is
the live network's forward pass, one decision at a time, and the other half is the environment.
Both numbers were taken on the development laptop with another training run on the machine, so
they are an upper bound rather than a quiet-machine figure; the command is the tiny MockEngine
config of tests/test_coordinator.py with probe_every_iterations = 1, reading time/probe
off the metric row (2026-09-22).
The gate is the comparison to set it against. At the shipped settings a gate plays
1000 + 2x200 + 8x100 = 2200 battles, about every two hours, and the table above budgets that at
5% of wall clock. A probe of the two anchors is 80 battles, under 4% of one gate. So one probe
per gate interval is noise and a probe every iteration is not, which is why the field is an
interval in iterations rather than a flag. ladder/probe_seconds_frac is what a run is watched
with, and time/probe is in the attributed sum rather than in time/residual, so the minutes
show up under a name.
12. Checkpoints¶
12.1 Layout¶
<runs_dir>/<run_name>-<run_id>/
identity.json RunIdentity, written once at run start
config.json the full resolved config, canonical JSON, written once
drift.json present only if --allow-identity-drift was used
metrics.jsonl one row per iteration
episodes.jsonl[.gz] one row per episode
alarms.jsonl one row per alarm firing
bundles/<iteration>/ the diagnostic bundle written on any halt
ladder/
games.jsonl the append-only result log; the source of truth for ratings
aggregate.json derived cache
eval_seeds.json the frozen seed set
ratings/<iteration>.json each refit's RatingTable
gates/<candidate>.json each GateDecision
snapshots/<sha>/ actor.safetensors + spec.json; outside checkpoints, they outlive them
checkpoints/
<cumulative_env_steps:012d>/
manifest.json
actor_critic/ actor.safetensors, critic.safetensors, arch.json
optimizers/ actor_adam.pt, critic_adam.pt, misc.json (cumulative_model_updates)
schedules/ state.json gamma, ent_coef, ent_coef_noop, lr_actor, lr_critic,
backoff counter, positions
advantage/ welford.json {"mean": f64, "count": int, "m2": f64}
rollout/ state.json per (worker, shard): seed path, reset ordinal, respawn generation
ladder/ pool.json, champion.json, ratings.json
metrics/ wandb.json {"run_id": ...}, nested per decorated sink
rng/ rng.json section 12.3
buffer/ buffer.npz optional (checkpoint.include_buffer, default false)
Every component owns its folder and its own save_checkpoint / load_checkpoint. Adding a component
to a checkpoint is adding a folder constant and a dict entry; there is no central serialiser to edit.
Atomicity. Write into <name>.partial/, fsync each file, then
os.replace(partial, final). The directory itself is not fsynced: os.fsync on a directory handle is
a POSIX guarantee and raises on Windows, where this harness's default profile runs, so the durability
step is per file and the atomic step is the rename. os.replace is atomic on both platforms, which is
the property the recovery actually needs: a crash mid-write leaves the previous checkpoint intact and
the partial directory obviously named. latest() reads the run index, never
int(x) for x in os.listdir(...). That form crashes the save on any stray file, and the save is
called from the crash handler, which is exactly where the user most needs it not to.
Pruning keeps checkpoint.keep checkpoints by the manifest index and survives a stray file in the
directory.
No pickle anywhere on a path this repo writes. safetensors for weights, torch.save/torch.load
with weights_only=True for optimizer state, msgspec JSON for everything else, .npz with
allow_pickle=False for the optional buffer. tests/test_snapshots.py scans written files for the
pickle protocol marker.
12.2 manifest.json¶
class Manifest(msgspec.Struct):
format_version: int # refuse a higher one by name; migrate a lower one
run_id: str; run_name: str
identity: RunIdentity # verbatim
config: dict # verbatim, so a checkpoint is self-describing
config_hash: str
iteration: int
cumulative_env_steps: int; cumulative_timesteps: int; cumulative_model_updates: int
wall_seconds: float; created_unix_ns: int
state_digest: str
component_versions: dict[str, int]
files: dict[str, str] # every file in this checkpoint -> its sha256
gate_seconds: float | None = None; last_decision: dict | None = None
last_checkpoint_step: int | None = None # where each cadence last fired; a resume
last_candidate_step: int | None = None # restores them, None (an older manifest)
last_floor_step: int | None = None # starts them at this checkpoint's env step
A resume continues each cadence from where it last fired, as the manifest recorded it: the periodic save, the gate's candidate and the floor's admission. Restarting them at the loaded checkpoint, which every resume did until these were recorded, moved the gate: resumed from a save between two gates, the next gate ran a whole cadence after the save instead of where the run had it due, and a gate that admits a snapshot changes the opponents the learner trains against.
A run that reaches its limit saves on the way out, unless the last periodic save already holds
that learner, and that save does not move the periodic cadence, so a run extended from it saves
where one run straight through would have. Until it did, a finished run's learner since its last
periodic save was on no disk. tests/test_resume.py::test_a_run_that_reached_its_limit_continues_from_its_end
stops a run between two periodic saves and two gates, extends it, and compares every metric field
and the state digest with a run straight through.
read() verifies every hash and raises CheckpointFormatError naming the first mismatch. That turns
a truncated write from a crash six weeks ago from a silent wrong resume into an error.
A checkpoint answers to the run's metric file. Its learner has to be one a row reports, so before
anything is loaded, checkpoint.check_described reads the metrics.jsonl of the run the checkpoint
sits in and refuses a checkpoint whose state_digest no row of its iteration carries, with a
CheckpointFormatError that names both digests and says to resume an earlier checkpoint with
--checkpoint. Any row of that iteration will do, because a resumed run writes the iterations after
its checkpoint a second time and each copy describes a learner that existed. What cannot be compared
is let through: no metric file, no row of that iteration, or a line torn by a crash. Past iteration 0
that absence is printed, so a resume the record could not vouch for does not read like one it did
(8db7e0d).
On resume the store diffs the loaded config against the current one and prints every difference
before continuing. A difference in an identity-hashed field is a refusal, not a warning: a policy
trained on MockEngine's 229-wide vector cannot load into the 12n + 37-wide one of an n-card catalogue, and
today nothing announces that mismatch.
load_checkpoint(folder, strict) defaults to strict=True for a resume and is only False for
load-weights-only, which starts a new run from an old policy. Tolerate-everything is right for a
research tool and wrong for a harness that promises the curve continues; with strict=False a missing
file prints the exact path it wanted and continues with a default.
12.3 RNG state¶
class RngState(msgspec.Struct):
master_seed: int
torch_cpu: str # hex of torch.get_rng_state()
torch_cuda: list[str] # one per device; empty for a run not on CUDA
python_random: list
numpy_minibatch: dict # Generator.bit_generator.state
iteration: int # the counter that addresses act/... and match/... streams
shard_streams: list[dict] # per (worker, shard): seed path, reset ordinal, generation
eval_seed_set_sha: str
One reference learner saves no RNG state at all; the other re-seeds all three generators from the run's initial seed on load, so a run resumed at ten million steps draws the same action noise it drew at step zero. Because every stream here is name-addressed, restoring the iteration counter and the shard positions restores the stream, not merely the parameters.
The buffer is not stored by default. The iteration boundary is the resume point and a fresh iteration
starts from a clean rectangle, so there is no mid-iteration state to preserve; include_buffer exists
for anyone who changes the collection rule.
12.4 How a resume is proved¶
tests/test_resume.py::test_resume_reproduces_metric_stream, marked slow, on CPU, with MockEngine,
a tiny network, InlineRolloutSource and determinism.tier = "run_exact":
- Run six iterations, checkpoint after three, record the full metric stream and every
state_digest. - From the checkpoint in a fresh process, run three more.
- Assert iterations four to six produce byte-identical metric rows and
state_digests.
royalelearn verify-resume --config <cfg> runs the same procedure against a real config on the user's
own machine, which is the only way the promise means anything on hardware this project has not seen.
docs/checkpoints.md states plainly what is and is not bit-exact: the environment always (the
engine is integer-only and seedable, and royalegym.replay re-verifies a trace bit for bit); the
learner under tier = "run_exact" on the same device class and torch/CUDA build. A driver change is
a documented identity change, not a broken promise.
13. Metrics and alarms¶
13.1 The sink¶
MetricsSink per section 4.6. Shipped:
JsonlSinkis always installed and never optional. One msgspec-encoded line per iteration inmetrics.jsonl, plusepisodes.jsonlandalarms.jsonl. No service, no account, replayable, and it is the file the resume test compares.ConsoleSinkprints the per-iteration grouped block. It iterates the metric dict rather than indexing a fixed key list, so a new metric cannot crash the console.ViserSinkcarries the learning status the viewer's panel reads, described below.WandbSinkis a decorator over any sink. It flattens nested keys togroup/name, persistsrun_idin the checkpoint and passes it towandb.init(id=..., resume="allow")so a resumed run continues the same wandb run, carries anenableflag that makes it a pure pass-through, populates the wandb config from the resolved config and theRunIdentity, and importswandblazily behind an optional extra.CompositeSinkfans out. The default isComposite([Jsonl, Console, Viser, Wandb(enable=False)]).
The learning panel is a sink, not a hook. RoyaleViser draws a learning panel beside the battle it
is already showing, fed by a second datagram sender on the state stream's port + 1. ViserSink
implements that sender itself, in about eighty lines of socket and msgpack: RoyaleLearn does not
import royaleviser, both because that package pulls in pygame and because the dependency direction
is the viewer watching the learner and never the reverse. The protocol, pinned in RoyaleViser's
docs/internals.md under "The learning status", is:
- the learner binds
host:porttaken from the sameROYALEVISERsetting as the state stream, with the learning port defaulting to the state port + 1; - it sends nothing at all until a
b"royaleviser 1"hello arrives. A viewer sends one every second and counts as attached while a hello has arrived within three seconds, so a detached run costs one monotonic clock read per iteration and no socket traffic; - while attached it sends one msgpack map
{"learning": {...}}per iteration, and re-sends the last status on a fresh hello so a viewer that starts mid-iteration fills immediately. Either a daemon thread or apump()called once a second serves the hellos; - one message is the whole status. The viewer replaces rather than merges, and renders an omitted field as an em dash, so a partial update would blank the panel rather than leave it stale.
The field names are fixed by royaleviser.model.Learning: run, iteration, policy_loss,
value_loss, entropy, kl, clip_frac, explained_var, grad_norm, learning_rate,
env_steps_per_s, engine_ticks_per_s, episode_ticks, crowns_per_episode, towers_per_episode,
illegal_rate, elixir_wasted, elo, win_rate, pool_size, games_vs_pool, plus an
"extra": {name: number | str} map drawn underneath, of which about three rows fit on screen.
ViserSink maps the metric row onto those names and spends the three extras on rating, formatted
as "1183 ± 22" from the fitted rating and its standard error, cards_per_match, and the last gate
verdict. tests/test_viser_sink.py plays the viewer with a plain socket: send the hello, receive the
datagram, decode it, assert the fields. It imports nothing from RoyaleViser, so the test states the
protocol rather than inheriting it.
Environment-side metrics reach the sink on the round scalars and the EpisodeRecords the worker emits;
learner-side metrics on UpdateResult; ladder metrics on the RatingTable and GateDecision. One row
per iteration merges all three. That is docs/design.md's "one metrics sink fed from both sides".
Aggregation happens in the worker. Counters accumulate per step in numpy and one fixed record per finished episode crosses the boundary. A per-step metric that is always reduced to a mean should be reduced where it is produced; the reference learner collects an array on every one of fifty thousand steps and reassembles it in a double python loop in the parent, which costs seconds per iteration hidden inside a timing residual.
metrics/schema.py is the single source of truth: a dict of key to (unit, description, healthy
range). tests/test_metrics.py asserts that every key a run emits exists in the schema and that every
schema key is emitted, so the documentation and the code cannot drift.
13.2 The metric list¶
run/: iteration, cumulative_timesteps, cumulative_env_steps, cumulative_updates,
wall_seconds, gamma, gae_lambda, credit_horizon_seconds, ent_coef, ent_coef_noop,
lr_actor, lr_critic, determinism_tier, resumed_with_drift, state_digest.
throughput/: overall_steps_per_second, collected_steps_per_second,
engine_ticks_per_second, rollout_capacity_ratio (rollout timesteps/s over update timesteps/s; the
invariant of section 2.3, warn under 1.5), boundary_mb_per_second, parent_wait_frac (the parent
blocked on workers that have not published, which rises when the workers cannot keep up; it is not a
number about workers idling), inference_ms_per_round, discarded_rows_frac, gpu_util_frac.
time/: iteration, collection, inference, env, codec, ipc, critic_pass, gae,
update, checkpoint, gate, probe, residual, and overlap_saved once there is an overlap
to save anything (it was published as a hardcoded 0.0, which reads as "overlap saved nothing this
iteration" rather than "there is no overlap", and is now absent). The ratio of collection to iteration is how you see
whether a run is environment-bound or learner-bound, and the residual is broken out because the
reference's residual silently absorbs the four places it is actually slow.
ppo/: policy_loss, value_loss, entropy, entropy_normalised, noop_entropy, kl,
kl_epoch{0,1,2}, clip_fraction, clip_fraction_epoch{0,1,2}, dual_clip_fraction,
explained_variance, grad_norm_actor, grad_norm_critic, update_magnitude_actor,
update_magnitude_critic, ratio_max_abs_dev, advantage_std_pre_norm, return_running_mean,
return_running_std, reward_clip_frac, n_minibatches, n_optimizer_steps,
samples_unused_frac, lr_backoff_events, adam_eps_floor_frac_actor, adam_eps_floor_frac_critic.
| metric | healthy | what it diagnoses |
|---|---|---|
explained_variance |
rising to 0.5-0.9 | the critic's health. No reference logs it, and value loss is uninterpretable while returns are normalised by a moving standard deviation. Still negative after fifty iterations is the most likely cause of a plateau |
entropy_normalised |
see note | raw entropy falling is ambiguous -- a confident policy or a tighter mask -- and the normalised form separates those. It does NOT separate a policy that has concentrated on one action, and on this action space that is most of what matters: over 250 legal actions a policy putting 15.3x uniform mass on the no-op still reads 0.980, reproduced from first principles in tests/test_hold_lift.py. Read policy/rollout_hold_lift for that question. The documented band of 0.3-0.8 was written for an unconditioned mean and could not be reached once the key was conditioned on choice rows; no replacement band has been earned yet |
noop_entropy |
above 0.05 nats | the leading indicator of no-op collapse, before cards_per_match bottoms out. Taken over the rows whose mask offered more than the no-op, because a decision the elixir bar cannot afford has a binary entropy of zero by construction and most decisions on this environment are that one (section 18, item 8). An unconditioned mean measures the elixir curve: it reads near zero on a healthy run, so a floor on it fires permanently, and a gate that had really collapsed would move it by a fraction of what it moves on the rows that had a choice |
adam_eps_floor_frac_actor |
no band | the share of the actor's parameters whose Adam second moment sits under adam_eps, where the step stops being normalised by the gradient and becomes lr*m/eps -- proportional to it again. Read it against the CRITIC's share in the same row; neither number means anything alone. It was added to test whether the floor accounts for an actor gradient orders of magnitude below the critic's with a near-zero kl and a clip fraction of zero, and it need not: a run can end with its actor LESS floored than its critic. Every config in this project runs 1e-08, not the 1e-5 default, and the same checkpoint measured against 1e-5 reads 94.8% against 29.1% -- a true number about a run that was never executed. That asymmetry is unexplained again |
clip_fraction |
0.05-0.20 | pinned near 1.0 is the signature of a rollout/update mask disagreement, or a learning rate far too high |
kl |
0.003-0.02 | below the band, lower batch_size; above it, raise batch_size or let the backoff act |
grad_norm_* |
below max_grad_norm most steps |
pinned at 0.5 every step means the clip is the binding constraint and the effective learning rate is unknown |
credit_horizon_seconds |
40-50 | logged every run so it can never be ten seconds by accident |
policy/: cards_per_match (healthy about 22; this and not the no-op rate is the
no-op-collapse metric, because a healthy policy is about 94% no-op and 99.5% no-op is 1.8 cards a
match and dead), noop_rate, legal_actions_mean, legal_actions_p05/p50/p95, forced_noop_frac
(the share with exactly one legal action, which carries zero policy gradient), tile_entropy,
tile_top1_share, card_tile_top10_share, rollout_hold_rate, rollout_hold_lift, rollout_legal_actions, rollout_choice_frac, per-card play frequency, and a 32x18 play heatmap per card
as an artifact every metrics.image_every iterations.
The four rollout_* keys are measured in the rollout's own forwards rather than in the update,
so they describe the policy that CHOSE the actions rather than the one the optimizer saw after an
epoch had already moved it. rollout_hold_lift is the one to read: each choice row's p(no-op)
divided by the uniform baseline of that row's own width, averaged over the rows that had a choice.
1.0 is a policy that has learnt nothing about when to wait, and the distance from 1.0 is the only
part of a hold rate that is about the policy rather than about the elixir bar. It has no healthy
band because nobody has yet trained a policy far enough to earn one, and for the same reason no
alarm reads it.
The healthy figure for cards_per_match is elixir arithmetic, not a guess: a full regulation match
generates 85.7 elixir (120 s at 0.357/s plus 60 s at 0.714/s) plus 5 at the start, and at an average
four-elixir card that is about 22 cards.
env/: episode_steps_mean and the histogram (a spike at the 480-step cap is the draw and
turtle equilibrium), episode_steps_p05/p50/p95, episodes_completed, ticks_mean, crowns_for,
crowns_against, crown_diff, tower_hp_frac_end_own/enemy, draw_rate, win_rate_by_seat (near
50%; the observation layer guarantees bit-identical mirrored observations, but the shipped engine's
multi-unit ground deploys are not seat-symmetric, so a small persistent deviation can be the game's
own and a large one is a learner-side leak), elixir_leak_frac, elixir_count_exact_frac (the share
of episode-seats whose count of the opponent's elixir stayed exact; below 1.0 the observation's
enemy-elixir field was an estimate on some episodes and the alarm below says so),
mean_elixir_at_decision,
frac_elixir_above_99, illegal_action_rate, and reward_terms/<name> (signed) and
reward_terms_abs/<name> (magnitude) per weighted term.
ladder/: rating, rating_se, rating_ci95_lo/hi per member, rating_above_v0,
elo_readout, champion_id, champion_step, pool_size, sampler_size, gate_attempts,
gate_passes, gate_observed_rate, gate_lower_bound, gate_failed_condition, gate_seconds_frac,
transitivity_residual, paired_rho, draw_rate_eval, eval_games_total, evictions,
pfsp_effective and pfsp_learner_step (absent while no snapshot stands in for the learner,
section 11.3), probe_seconds_frac, and role_counts/{mirror,pool,scripted} with
role_counts/pool_fallback: the training battles that finished this iteration, one per battle
episode, by what they played, and how many of the scripted ones the plan laid out as pool (a pool
battle plays a scripted opponent until the first snapshot is admitted, which the config's split
cannot show); and on a probe iteration ONLY, score_vs/<rung> with score_vs_n/<rung> and
score_vs_ci95_lo/hi/<rung>, of which score_vs_noop and score_vs_random_legal are the two
anchors under their older names.
health/: illegal_action_rate (exactly zero by construction; this is an alert, not a plot),
mask_disagreements (from the start-up gate), worker_restarts, worker_failures_by_kind,
rows_dropped_dead_worker, obs_codec_clipped, samples_unused_frac, nan_guard_trips,
housekeeping_failures (plus housekeeping/{kind} for each kind that has failed),
vram_peak_mb, rss_peak_mb, buffer_fill_frac, and the device-memory regime read at the end of
each iteration when the run's device is a CUDA device: vram_reserved_mb, vram_inactive_split_mb,
vram_driver_free_mb, vram_alloc_retries, vram_available_mb (driver free plus what this process
reserves) and vram_needed_mb (one minibatch's peak measured at startup plus
doctor.vram_headroom_mb). The vram_spilling alarm reads time/update and vram_driver_free_mb
(and ppo/actor_frozen where a run schedules the actor's rate); the other two are kept because they say which memory regime a slowing update is in.
housekeeping_failures is the total of the retries that did not take, across three actions that
each swallow their own error so the run survives them: pruning a checkpoint directory, compacting
the metric log, and stopping an evaluation worker. Each of those counted itself from the day it
was written and nothing read the number, so a failure incremented a variable that was discarded
with the object. The total is unconditional -- a zero is a measurement, and it is what puts the
key under the row test that every non-conditional key is actually emitted -- while the per-kind
breakdown (prune, compaction, eval_shutdown) appears only when that kind has failed.
13.3 Alarms¶
metrics/alarms.py. Each alarm is (name, predicate, severity, patience) and fires when its predicate
holds for patience consecutive iterations. severity="warn" logs and writes an alarms.jsonl row;
severity="halt" additionally writes a checkpoint and a diagnostic bundle, then raises AlarmHalt so
the process exits non-zero. Thresholds live in config.alarms, are recorded, and are excluded from the
identity hash because an alarm can stop a run but never alter a number.
A halt is a clean stop, never a death, and it explains itself. The checkpoint is written before
the exception, so a halted run is always resumed rather than restarted; the bundle carries the last
fifty metric rows beside it. The final console line, and a halt.json in the bundle, name the alarm,
the threshold it crossed, the last five values of every metric the alarm reads, the checkpoint path
and the resume command verbatim. The reason a halt states its own evidence is that the first question
anyone asks on finding a stopped run is whether the stop was real, and a line reading only that a
metric crossed a threshold costs the same morning as no run at all.
The halting iteration's own row is written to metrics.jsonl before the alarms read it, so the
halt's checkpoint holds the learner that row describes and passes the resume check of section 12.2
(8db7e0d). That iteration's alarm rows are still missing: AlarmSet.evaluate raises before it
returns what fired, so neither the halt nor the warnings beside it reach alarms.jsonl. The halting
alarm is in the bundle, and the warnings beside it are only printed.
This is also why no halting alarm has patience = 1 on a learning quantity. The four that do,
illegal_actions, ratio_invariant, nonfinite and buffer_overflow, are correctness assertions
whose first occurrence is already a defect, and continuing past one wastes the compute that follows.
Everything that measures how training is going waits several consecutive iterations, because a
policy that is still near-random moves these quantities around for reasons that are not the failure
being hunted, and a false halt costs everything the run was for.
| alarm | predicate | patience | severity | what it means |
|---|---|---|---|---|
illegal_actions |
env/illegal_action_rate > 0 |
1 | halt | a mask bug, an unmasked policy, or a wrong action encoding. With a correct mask it is exactly zero |
ratio_invariant |
ppo/ratio_max_abs_dev > 5 * ratio_atol |
1 | halt | a mask, codec or weight-version mismatch (section 9.6). The multiple of the configured tolerance is what keeps the alarm meaningful in fp32 and under bf16 alike; all three failures it is aimed at produce a deviation of order one |
nonfinite |
any loss, gradient or logit non-finite | 1 | halt | |
buffer_overflow |
health/buffer_fill_frac > 1.0 |
1 | halt | an invariant is broken. A healthy rectangle reads exactly one: every cell is written once per iteration, so anything under one is a cell nobody filled and anything over it is impossible. The threshold is above one and not below it for that reason: a bound of 0.98 would halt every healthy run on its first iteration |
worker_failures |
health/worker_restarts rose |
1 | warn | |
worker_failures_persistent |
rose on three consecutive iterations | 3 | halt | |
clip_pinned |
ppo/clip_fraction > 0.5 |
3 | halt | mask disagreement, or a learning rate far too high |
kl_high |
ppo/kl > 0.05 |
3 | warn | the backoff should already be acting |
kl_dead |
ppo/kl < 1e-5 |
10 | warn | nothing is moving: dead entropy, learning rate too low, or a frozen head |
ev_negative |
ppo/explained_variance < 0 |
50 | warn | the most likely cause of a plateau |
noop_collapse |
policy/cards_per_match < 8 |
5 | warn | about 22 is healthy |
noop_collapse_severe |
policy/cards_per_match < 3 |
5 | halt | until 92786d4 this threshold could not be reached, and not because of its value. See below |
noop_entropy_floor |
ppo/noop_entropy < 0.02 |
5 | warn | the leading indicator; it watches after ent_coef_noop has annealed to zero, which is when it matters most |
tile_spam |
policy/tile_top1_share > 0.25 |
5 | warn | |
artefact_exploit |
policy/card_tile_top10_share > 0.5 |
5 | warn, and dump five winning traces | a real meta is not that concentrated |
draw_equilibrium |
env/draw_rate > 0.5 and env/episode_steps_at_cap_frac > 0.8 |
5 | warn | the turtle equilibrium |
seat_bias |
env/win_rate_by_seat's 95% interval excludes 0.45-0.55 |
3 | warn | an unseeded reset, a reward asymmetry, or an observation mirror bug; a few points inside that band can be the shipped engine's own seat asymmetry, which is why this warns rather than halts and why the ladder's paired evaluation swaps sides on every seed |
elixir_count_inexact |
env/elixir_count_exact_frac < 0.99 |
3 | warn | the observation's opponent-elixir field is an estimate on some episodes: a repeated card in a deck, or an engine whose elixir law is not the calibration's. The policy is reading a documented-exact slot that is not. A value near zero rather than slightly under one is the second cause and not a broken counter: it says the engine build and the card data disagree about elixir, so read the run's identity.json engine_build before anything else. The row itself carries no engine digest: run/state_digest is the learner's weights, not the engine |
reward_clipped |
ppo/reward_clip_frac > 0 |
1 | warn | a reward reached advantage.reward_clip. Potential shaping leaves the optimum unchanged only while every step's reward reaches the return intact; on a clipped step the terms stop cancelling, and a clipped terminal step makes a win worth less than a win. Zero at every iteration of the first eight runs on the development machine [M], which was an observation that the rewards stayed small and is now a check. The schema's healthy band for the key was 0 to 0.01 until this row; it is exactly 0 |
shaping_dominates |
sum of absolute shaping terms > absolute terminal term |
5 | warn | under a potential reward this says a shaping term has stopped telescoping, which is what a term that is not a difference of a potential does; it is not a check on the weights. Measured quiet in production 2026-09-22 at 2.9% of the objective (section 10). Until 5724eb5 and dc7f7a1 it compared two structural zeros |
vram_spilling |
time/update >= 2 x the best update this run has had in the same actor state AND health/vram_driver_free_mb < 128 |
3 | warn | the update is several times slower than this run has managed, on a card the driver says is full. That is what an allocation backed by host memory over PCIe looks like from inside the process, and the platform gives no other sign: it does not refuse an oversubscribed allocation, it serves it and reports success. The memory reading is there to tell a spill from a busy machine, which slows an update by 1.4 to 1.7 rather than by 4. The bar is a running minimum, so a slow iteration cannot raise the bar it is judged against and the first iteration's warm-up cannot lower it. Iterations with the actor frozen (ppo/actor_frozen, section 19.5) keep a best of their own: they are much cheaper, and with one best a warm start's unfrozen updates all read as twice it. A row without the key counts as unfrozen. It warns rather than halts, because stopping a long run over a neighbour's memory costs more than the slowdown does. It stays silent on a run with no CUDA device, because health/vram_driver_free_mb is then absent |
transitivity |
ladder/transitivity_residual > 0.10 |
3 | warn | the scalar rating is lying |
gate_starved |
five consecutive gate failures | 1 | warn | the plateau signal, stated as an event |
capacity_ratio |
throughput/rollout_capacity_ratio < 1.5 |
3 | warn | the harness is becoming the bottleneck |
An alarm nobody has seen stay silent is unvalidated. Watching one fire proves only that it
can fire; what validates a threshold is a healthy iteration where the metric sits clearly on the
right side of it with room to move. An alarm that holds on every healthy row is worse than no
alarm, because a channel that always speaks teaches its reader to stop listening, and the event
it was built for then arrives invisibly. That is not hypothetical here: noop_entropy_floor held
on both iterations of the first real run and was read twice, by two people, as a policy warming
up.
Those two iterations are the only healthy rows this harness has produced. Measured against them, four alarms held, and two of them are a threshold problem rather than a finding:
| alarm | held | reading |
|---|---|---|
noop_entropy_floor |
both rows | the threshold was wrong, and the metric is now conditioned on the rows that had a choice (section 18, item 8). The number those rows report is the unconditioned one |
kl_dead |
both rows, at 4.8e-06 and 3.6e-07 | the same fault, and the same repair. A forced row's ratio is exactly one: the distribution is a point mass at the same action before and after the update. So it contributes zero KL structurally. ppo/kl, ppo/clip_fraction and ppo/dual_clip_fraction are now means over the rows that had a choice, and the threshold stands. This reached past the dashboard: lr_backoff reads this KL, so a diluted one put the brake that stops a blow-up out of reach by the same factor |
ev_negative |
first row only | healthy: the critic had seen one batch, and explained variance was +0.25 by the second. Patience is 50 |
noop_collapse |
first row only | healthy: cards_per_match was 6.98 on a freshly initialised policy and 22.08 by the second row. Patience is 5 |
So two alarms are fixed and two behaved. The rule that produced that table is worth more than the table: before shipping a threshold, measure the quantity on the population it will really be averaged over, and if that population is mostly structural zeros, condition the metric rather than lowering the number. The share of rows that are structural zeros is itself a property of the game rather than of the learner. Here it is the elixir economy, and it moves as the policy learns to hold elixir, as the deck changes, and in overtime at double rate. A threshold re-tuned against today's share is a number with an expiry date nobody will notice passing.
What a validated threshold looks like, from the two that behaved: held on the first iteration, cleared by the second, with patience long enough to absorb the start. A new alarm can be checked against that shape in a minute.
A threshold can also be unreachable because of the population it is computed over, and no amount
of tuning it would help. policy/cards_per_match was a mean over every seat that finished,
including the opponents'. Before the first snapshot is admitted, about one seat in eight is
scripted:random_legal, which plays a card whenever it can afford one and scores about 31 an
episode [M], so the blended mean could not fall below about 3.9 however completely the learner
stopped playing. The halt at 3.0 was therefore arithmetically unable to fire for the first gate's
worth of iterations, about 3.7 hours at laptop geometry, which is most of the period the collapse it
guards against actually happens in. The repair was to the metric's population and not to its
threshold: since 92786d4 these are the learner's own seats. Measured and worked out on 2026-09-22.
That is the second alarm in this table to have been unreachable by construction, after
shaping_dominates compared two structural zeros, and the third if the spill alarm above is
counted. The shape they share is a threshold on a quantity whose population nobody stated. Section
13.4's rule, that a metric's row population belongs in its identity, is the systemic answer, and it
is still not built.
vram_spilling is the cautionary one, and it is worth reading before anyone writes another alarm
about memory. The first version compared the memory this process could still take -- driver free
plus what its allocator already holds -- against the peak one minibatch measured at startup. It was
tested on a card on 2026-09-22: a second process took 2,048 MB and left 35 MB free for five
iterations, against a patience of three. health/vram_driver_free_mb went 377.2 to 0.0 while
health/vram_available_mb moved only 3371.9 to 2994.7, against 2349.0 needed, so it stayed silent
by 645.7 MB [M]. The reason is arithmetic, not tuning: an outsider can only consume the free
part, so that sum bottoms out at what this process holds, and the startup gate guarantees that what
it holds exceeds what one minibatch needs. The alarm was least sensitive in the case its own text
named.
No in-process memory reading fixes that, because the platform does not refuse an oversubscribed
allocation: it backs it with host memory and reports success, so the allocation is served, the
counters look ordinary, and only the clock changes. The alarm therefore watches the harm and uses
the memory reading as the evidence that this is the cause. Its numbers are measured: the same update
took 47-49 s with room and 180-233 s without [M], while machine contention alone moved it from
435-531 s to 758 s [M]. A factor of two sits between those two populations. It has still never
been seen firing on a real spill; the check is the experiment above, re-run against this version. It
did fire falsely on a warm-started run, with the update steady and the card at 0 MB driver-free
because the caching allocator had filled it: the frozen-actor iterations before the unfreeze had set
the best, so every unfrozen update was twice it (tests/test_alarms.py replays that shape). A
frozen actor's iterations now keep their own best.
The other alarms have been checked against two iterations of one profile, which by the rule above
is not validation. A metric's row population belongs in its identity rather than in its
implementation, written as ppo/kl@choice against @all and declared in the schema. Then a
threshold cannot be set against the wrong population by accident, and a reviewer sees the fault
without running anything. That is the systemic form of both repairs, and it is not built.
13.4 The diagnostic bundle¶
On any halt, metrics/bundle.py::write_bundle produces <run>/bundles/<iteration>/ containing: the
last fifty metric rows, the alarm rows, the resolved config, the RunIdentity, the last twenty
EpisodeRecords, the ten worst ratio outliers with their slot, cycle and episode ordinal, a
royalegym.replay.Trace of one offending episode re-simulated from its shard seed and reset ordinal
(so it is exact and self-verifying with verify_trace), and the current state_digest. One directory,
attachable to an issue. This is the concrete form of docs/design.md's "a learner that is 80% right
produces a bot that loses for reasons nobody can attribute".
14. The loop, the entry point and the CLI¶
14.1 LearningCoordinator¶
royalelearn/coordinator.py. The only place in the package where the phases of an iteration are
ordered.
class LearningCoordinator:
def __init__(self, cfg: RunConfig) -> None: ...
def __enter__(self) -> "LearningCoordinator":
"""Runs preflight (section 7.7) and raises PreflightError with the RustEngine stale-build
text verbatim if the engine build disagrees with the data on disk."""
def __exit__(self, *exc) -> None:
"""Idempotent close: CLOSE to every worker, join with a timeout, terminate the stragglers."""
def learn(self, until_timesteps: int | None = None) -> None: ...
while cumulative_timesteps < limit:
plan = matchmaker.plan(iteration, pool, ratings, geometry) # the opening table only
source.begin_iteration(plan, buffer, iteration)
for t in range(T): # one cycle = shards_per_worker rounds
for _ in range(shards_per_worker):
round = source.next_round(cfg.rollout.round_timeout_s)
assigned = assign_ended_battles(round, matchmaker, pool, ratings) # section 7.5
actions, log_probs = inference.act(round, uniforms[t])
source.submit(Step(actions=actions, gamma=sched.gamma, **assigned))
buffer.record_round(round, actions, log_probs)
episodes.extend(round.episodes)
handle_failures(source.drain_failures()) # restart, mark rows invalid, count
trainable = buffer.trainable()
check_iteration(collection, trainable, episodes) # raises RoundsMissing, TrainableRowsShort,
# NonFiniteLogProb, DeployRefused, EpisodesUnpaired, MixtureDrifted, AssignmentInsideEpisode
probe = policy_probe.measure(buffer, trainable) # raises StoredActionIllegal
result = update.step(buffer, sched) # critic pass, GAE, PPO; section 9.5
ladder.record_training_results(episodes) # kind="train"; excluded from the fit
if crossed(cfg.ladder.candidate_every_env_steps):
candidate = snapshots.put(actor_critic, cumulative_env_steps)
decision = gate.evaluate(candidate, pool, eval_runner) # builds and closes its own farm
ladder.apply(decision)
if iteration % cfg.ladder.refit_every_iterations == 0:
ratings = rater.fit(results.eval_view())
if cfg.ladder.probe_every_iterations and iteration % cfg.ladder.probe_every_iterations == 0:
rungs = {rung: probe_runner.compare(f"learner@{cumulative_env_steps}", rung,
games=cfg.ladder.probe_games)
for rung in cfg.ladder.probe_opponents} # kind="probe"; section 11.9
row = merge(run_fields, throughput, time, result, probe, episode_stats(episodes),
ladder_fields(rungs))
sinks.write(row); sinks.write_episodes(episodes) # before the alarms read it
alarms.evaluate(row) # may raise AlarmHalt
if crossed(cfg.checkpoint.every_env_steps):
store.write(components, manifest())
sched.advance(cumulative_env_steps)
iteration += 1
if not saved(cumulative_env_steps): # the limit: section 12.2
store.write(components, manifest()) # moves no cadence
Every invariant reads only what collection wrote, so the batch is judged before anything learns from
it, and a refused batch is never trained on. Until 8db7e0d the update ran first, and a refused batch
was trained and then refused. The rollout's statistics are drained right after collection as well,
and the drain compares the rows they were taken over with the rows the decoded masks offer a choice
on. It ran when the row was built until the drain moved, after the update and after a gate's
candidate was stored, where a failure could only end the run with its emergency save refused. Now a
failure is raised before the learner moves, so the emergency save keeps the learner the last row
describes. The row is written before the alarms read it because a halt checkpoints the learner that
row describes, and the row has to be in metrics.jsonl by then (section 12.2).
rollout.overlap is REFUSED as of 2026-09-22, and what follows describes what it would do
rather than what it does. Half of it exists: BatchedInference.begin_iteration takes a
BehaviourSnapshot and samples from it. The driver does not -- nothing passes one, there is no
second thread, there is no second buffer, and ppo/behaviour_lag_iterations is in this sentence
and in no schema. It was not a free thing to leave accepted: the preflight sizes TWO rectangles
when it is set (rollout/preflight.py), so the workstation and many_core profiles shipped
reserving twice the buffer memory for a feature that never ran, and that reservation is what the
memory gate is checked against. check_consistency now says so instead. When the driver is
built, delete this paragraph's first sentence and the refusal together.
With rollout.overlap = true the collection of iteration i runs on a second thread and a second
buffer while the update of iteration i-1 runs on the default CUDA stream, against a
BehaviourSnapshot of the actor taken at the iteration boundary. The lag is exactly one iteration by
construction and the stored log-probs are the snapshot's, so the PPO ratio measures the true
off-policyness and the clip bounds it. ppo/behaviour_lag_iterations is logged as 0 or 1 so the
regime is visible in the run's config panel, and because sampling is driven by name-addressed uniforms
the trajectory is identical either way.
Invariants asserted once per iteration, before the update, each raising a named exception carrying the offending slot and cycle:
assert (round_counter == T * shards_per_worker)
assert buffer.trainable_mask().sum() >= cfg.ppo.timesteps_per_iteration * 0.98
assert np.isfinite(buffer.log_prob[buffer.trainable_mask()]).all()
assert mask_bit(buffer.mask_bits, buffer.action)[buffer.trainable_mask()].all()
assert (buffer.deploy_status[buffer.trainable_mask()] <= 0).all()
assert episodes_end_in_pairs(episodes)
assert abs(discarded_rows_frac - expected_from_mix) < 0.05
assert assignments_constant_within_episodes(buffer.group, buffer.episode_end)
The last of those is what makes "one policy for one whole episode" checkable rather than merely
intended: a battle's group entry may change only on the cycle after its episode_end, anywhere in
the rectangle. It is one vectorised comparison over (T, R) and it catches every way an assignment
could reach the middle of a trajectory.
14.2 The file a user runs¶
examples/train_1v1.py is about fifteen lines, all of it the things a bot creator actually changes.
from pathlib import Path
from royalelearn import LearningCoordinator, load_config
def main() -> None:
cfg = load_config(Path(__file__).parent / "configs" / "laptop.json")
cfg.run_name = "my-first-run"
cfg.advantage.gae_lambda = 0.99 # the number to think about; see docs/harness-spec.md
cfg.metrics.sinks[3].enable = True # weights and biases
with LearningCoordinator(cfg) as run:
run.learn(until_timesteps=100_000_000)
if __name__ == "__main__":
main()
examples/custom_reward.py shows the other thing a bot creator changes: a RewardFunction subclass,
registered through config.extra_component_modules so that EnvFactorySpec may instantiate it, and
therefore recorded in the checkpoint and in the ladder's context like any other component.
14.3 The CLI¶
python -m royalelearn <command>, also installed as the royalelearn console script.
| command | what it does |
|---|---|
config [--profile laptop\|workstation\|many-core] [-o run.json] |
write a fully populated default config |
doctor [--config F] |
the start-up gates of section 7.7 on their own: build one env, print the engine build digests, the observation space and the codec table, run mask_disagreements over all 2304 non-no-op actions, check the action-layout identity, print the RAM ledger, the credit horizon, the geometry and the run_id. Seconds, and it catches most first-run failures |
bench [--config F] [--seconds 60] |
measure and print, for this machine: the number of iterations timed, env milliseconds per game-step, codec microseconds per row, boundary megabytes per second, inference milliseconds per round, update timesteps per second, the rollout/update capacity ratio, peak VRAM, and the last iteration's ratio_max_abs_dev (its first minibatch) beside the configured ratio_atol; under run_exact, also checked_iteration, whether the last timed iteration is one the harness checks, and filled_iteration, whether it filled fresh memory: a checked one or the process's first (section 5.1). It PRINTS that block; nothing writes docs/throughput.md, which is a page kept by hand, and a command that overwrote it would lose the prose around the numbers. --iterations caps the loop (default 3) and --seconds is a lower bound checked between iterations, never inside one |
train --config F [--run-name N] [--until-timesteps T] [--inline] [--device cuda\|cpu] |
a new run |
resume --run DIR [--checkpoint PATH] [--until-timesteps T] [--allow-identity-drift] |
continue; refuses on an identity mismatch by default and names every differing field |
verify-resume --config F [--iterations 6] [--split 3] |
section 12.4's proof, on the user's own machine |
eval --run DIR --a ID --b ID [--seeds N] |
a paired evaluation between two members, with its interval, outside the training loop |
gate --run DIR --candidate ID |
re-run a gate decision from stored snapshots and print the verdict |
rate --run DIR [--output F] |
refit ratings from games.jsonl and print the table with intervals and the transitivity residual |
identity --config F |
print the RunIdentity and the run_id: the thing to paste into an issue |
replay --run DIR --episode W/S/B/O [--out trace.msgpack] |
re-simulate one episode from its shard seed and reset ordinal and verify it with royalegym.replay.verify_trace |
play --checkpoint DIR [--opponent random\|noop\|<snapshot>] [--viser] |
watch one battle; --viser builds the single-battle vec env with viser="env", which is the same path worker 0's first shard takes during a run |
cli.py sets CUBLAS_WORKSPACE_CONFIG and the BLAS thread variables before importing torch, and
royalelearn/__init__.py imports nothing that needs torch, so royalelearn --help, royalelearn
config and royalelearn identity work in an environment without it.
--resume latest resolves by reading the run index and printing what it chose. There is no
"latest" string inside the config: one reference resolves it by string-munging the save path, so a
run named myrun and one named myrun-variant collide.
14.4 Interactive control and crash discipline¶
Checked once per round, not once per iteration, via a sentinel file in the run directory and a
SIGINT handler, so neither a tty nor a busy-wait is needed: c checkpoint now, q checkpoint and
quit, p pause (a blocking wait on an Event).
The whole loop is wrapped in try/except (Exception, KeyboardInterrupt), with a nested try around
the emergency checkpoint, then a finally that closes every worker, joins with a timeout and
terminates the stragglers. KeyboardInterrupt is named explicitly because it is not an Exception,
so a bare except Exception skips its own emergency save on Ctrl-C.
An emergency checkpoint holds only the learner the last metric row describes. From the update's
first change until that iteration's row is written, the learner in memory is one no row describes, so
a crash in that window saves nothing and prints the checkpoint to resume from. What the run did since
the last periodic checkpoint is lost, which the cadence (checkpoint.every_env_steps) already allows.
Outside that window the save is skipped when a checkpoint of the same learner is already on disk (the
periodic one, a halt's, or the one the run resumed from), because that one was written at the
iteration boundary and can still replay the iteration that failed. Otherwise the save is written. A
save taken after collection carries the battles' advanced ordinals: the ordinals address battles, so a
resume plays the next ones, none twice, and the refused iteration's episodes are skipped rather than
trained on (8db7e0d).
15. The test plan¶
Every test runs on CPU, uses MockEngine, and never invokes cargo or maturin. MockEngine's
16-card catalogue gives a vector nowhere near the width of the full one, which doubles as the standing
check that no width is written down anywhere: a test that passes on both catalogues cannot contain a
literal from either. Networks in tests are channels=8, blocks=1. Markers: slow (over five
seconds) and engine (needs a fresh royalesim build). addopts runs everything not marked slow
or engine; the repo gate runs pytest -q and then pytest -q -m "slow and not engine".
Neither reference learner has a single test file. That is the clearest place where the reference must not be imitated: the layers below this one are held to full suites and this one carries the release metric.
| file | asserts | speed |
|---|---|---|
test_package.py |
imports without torch; the public names resolve; the six seed modules are gone; asking for a torch-dependent name without torch raises an ImportError that names the missing package |
fast |
test_config.py |
JSON round-trip is canonical and idempotent; a typo is rejected at every nesting level; the hash is stable under key reordering; each shipped profile is internally consistent (T * learner_rows >= timesteps_per_iteration, batch_size % minibatch_size == 0); each shipped example file (laptop.json, workstation.json) equals its profile leaf for leaf, both loaded and as written |
fast |
test_seeding.py |
derive_* is stable across processes and platforms against pinned values; adding a new stream name does not change an existing stream; two names do not collide on their first four draws |
fast |
test_identity.py |
the identity is order-independent JSON; table-driven over every field, each included field changes run_id and each excluded field does not, frame_stack, codec_table_digest and a changed Reveal among the included ones |
fast |
test_obs_layout.py |
hand_card_onehot, hand_cost and hand_affordable are resolved by name from vector_layout() and the resolved slices match the environment's own, on MockEngine and on a synthetic full-catalogue spec whose widths differ from it; a layout missing a required name raises PreflightError naming that name; no offset is computed from V anywhere in the module |
fast |
test_codec.py |
the codec table is decided from the space: a plane whose declared high exceeds 255 lands in float16 and one under it lands in uint8, a plane the layout declares static is not stored and one it does not declare static is stored even when it is constant on the sample; pack then unpack is exact on every uint8 plane and on the mask; the fp16 parts round-trip within 2^-10 relative on 10 000 real observations; mask_planes reconstructed at unpack equals the environment's, bit for bit; row_bytes matches the formula of section 2.1 on both catalogues; the clipping counter fires on a synthetic count above a declared bound; codec_version and codec_table_digest are in the identity |
fast |
test_layout.py |
the shared-memory offsets and sizes are self-consistent and stable against a golden record, so a change to LAYOUT_VERSION is deliberate; a short block raises at construction, not at first write; the error-byte protocol round-trips a traceback |
fast |
test_action_layout.py |
for all 2304 non-no-op actions, a one-hot (hand_size, tiles_y, tiles_x) tensor reshaped in C order has its argmax at parser.encode(slot, x, y); index 0 is the no-op. Exhaustive, because this is the most load-bearing index identity in the harness |
fast |
test_distribution.py |
illegal actions have p == 0 exactly and log_prob == -inf; entropy is finite with all but one action masked; mode() respects the mask; sampling from fixed uniforms matches a brute-force inverse CDF and never returns an illegal action; mask[NOOP] is set on 10 000 random RoyaleGym states including a game_over one |
fast |
test_nets.py |
every documented shape and dtype at B = 1, 7, 512, on both catalogues and at frame_stack 1 and 2; the stem's input width equals k*(S+P) + 2 + E computed from the spec; logits are float32 under autocast; the gradient into a masked logit is exactly zero; the entropy term's gradient into critic parameters is exactly zero for both trunk variants; orthogonal gains are as specified; parameter counts within 5% of the documented figures; arch_digest round-trips, covers frame_stack, and a mismatching digest is refused with a message rather than a shape error |
fast |
test_gae.py |
the vectorised implementation equals reference_gae to 1e-6 over random inputs; terminated bootstraps from 0 and truncated bootstraps from final_value (named as a regression test for the reference's bug); the recursion breaks at a truncation as well as at a termination; a one-cycle episode; a dead-worker prefix; all-terminated and none-terminated; the reported credit horizon matches 1/(1 - gamma*lambda) |
fast |
test_returns.py |
Welford matches numpy's mean and n-1 variance over 10^5 samples; the state round-trips through JSON; the mean is never subtracted |
fast |
test_buffer.py |
record_round writes exactly the right cells; minibatches cover every trainable valid cell exactly n_epochs times with nothing dropped; a batch never straddles an epoch; the pinned staging ring does not alias; the measured footprint matches the row_bytes formula of section 2.1 to the byte |
fast |
test_frame_stack.py |
at k = 2 the two frames of a stacked cell have consecutive info["tick"] values and belong to the same episode; the first cell of an episode stacks a zero history; cycle 0 of an iteration stacks the history rows carried over from the previous one; vector is the current frame's and is not stacked; frame_stack changes arch_digest and obs_digest; at k = 1 the stacked observation is byte-identical to the unstacked one |
fast |
test_ppo.py |
gradient accumulation over k minibatches gives the same gradient as one full batch to 1e-5, which is the property that makes minibatch_size a pure memory knob; clip fraction and KL match hand-computed values on a synthetic batch; dual clip binds only for a negative advantage; ratio == 1 gives exactly -mean(A); the two mask asserts fire when fed a deliberately mismatched mask; the backoff fires after exactly patience consecutive breaches and floors at lr_min |
fast |
test_schedules.py |
the gamma and entropy anneals hit their endpoints exactly at the stated env-step count; the whole schedule state round-trips | fast |
test_rewards.py |
each potential term equals gamma*Phi(s') - Phi(s) on a hand-built transition; the composition is antisymmetric between seats on a mirrored transition; set_gamma reaches every term; the schedule's discount reaches the reward during a real collection, inline and (slow) through worker processes; every profile and shipped example file names default_potential_reward; the objective is filed under the name the shaping alarm reads; the committed-elixir potential is zero for a card played and negative for elixir left in the bar while the opponent's rises |
fast |
test_rollout_inline.py |
an InlineRolloutSource run of 20 cycles: slots map to the right battles, rewards are antisymmetric on mirror battles, episode ends arrive in pairs, deploy_status is never in 1..11, the seven terminal scalars are read out of final_info and match what the worker counted, and a battle's assignment changes only on the cycle after its episode_end |
fast |
test_rollout_farm.py |
the differential test: ProcessRolloutSource and InlineRolloutSource with the same seed produce byte-identical buffers, scalars and episode records over 30 cycles. This is the acceptance criterion for any future Rust worker |
slow |
test_worker_hygiene.py |
the worker has no torch in sys.modules; the thread variables are set before numpy; a deliberately raised exception arrives as WorkerFailure(kind="exception") with the traceback, the farm restarts the worker with the next generation, the run continues, and health/worker_restarts rises; a hung worker produces WorkerTimeout within the timeout rather than hanging; an idle worker uses under 10% of a core; a command sent out of turn is still taken; a wait takes the token that announced its command and sleeps through one an earlier command left |
slow |
test_rollout_invariants.py |
the section 14.1 assertions fire on deliberately corrupted rounds: a dropped cycle, a mismatched slot count, an action illegal under its stored mask, a truncated cell with no final_value |
fast |
test_rating.py |
the Bradley-Terry-Davidson fit recovers known strengths from synthetic results within its own standard errors over 100 seeds and to within 5 Elo at n = 2000; it is invariant to result order and to a permutation of player ids; standard errors shrink as one over the square root of n; the anchor is pinned exactly; the standard error of a difference uses the 2x2 block; the Davidson term recovers a known draw rate; the residual is near zero on transitive data and large on synthetic rock-paper-scissors; wilson_interval matches the published table |
fast |
test_results_log.py |
the aggregate cache equals a rebuild from games.jsonl; a truncated last line is tolerated; contexts are never pooled without the flag |
fast |
test_matchmaker.py |
the role counts land within one battle of the mixture over 30 master seeds and seven geometries, including an odd battle count, a single battle, (1, 0, 0) and shares that do not divide the battles; a role does not move with the ordinal, the pool or the ratings; every cycle collects exactly n_battles + mirror_battles learner rows and an iteration exactly cycles × that; the shipped laptop iteration clears MIN_TRAINABLE_FRACTION on every one of those seeds; the layout is a permutation, so every worker carries mirror battles; one slot still draws both seats and several opponents over its episodes; PFSP weights are floored; assign(battle, ordinal, ...) is a pure function of its arguments and the master seed; plan() equals assign() applied across the geometry; at most max_resident_opponents snapshots appear; an assignment is binding for a whole episode; a checkpoint round-trips the roles and a format-1 checkpoint still loads |
fast |
test_gate.py |
each of the three conditions fails independently and the decision names which; an observed 55.2% at n = 1000 passes and 55.0% fails; condition 3 compares the candidate's observed mean against the champion's fitted predicted mean and passes on a champion with no games against the stratified eight; the cycle case admits without promoting; paired scoring collapses sides correctly and is symmetric under swapping A and B; the bootstrap interval covers a known rate at the nominal level over 200 synthetic replications | fast |
test_eviction.py |
anchors, v0 and the champion chain are never evicted; the stratified sample spans the rating range; cycle-flagged snapshots are preferred; the archive and the log are untouched |
fast |
test_snapshots.py |
save and load round-trip; a mismatching arch_digest, obs_digest or codec_table_digest is refused with the field named; EvalRunner refuses a pairing whose two obs_digests differ and names both; the LRU evicts; a byte scan finds no pickle protocol marker in any written file |
fast |
test_metrics.py |
every key a run emits exists in schema.py and every schema key is emitted; nested keys flatten to a/b; the wandb sink is a pure pass-through when disabled; the decorator's checkpoint nests |
fast |
test_viser_sink.py |
a plain socket plays the viewer: nothing is sent before a hello, a hello produces one msgpack map within a second, every fixed field of the learning status is present with the right type, the three extras carry the formatted rating, cards_per_match and the last gate verdict, a second hello re-sends the last status unchanged, and a detached sink sends nothing over fifty iterations. Imports nothing from RoyaleViser |
fast |
test_alarms.py |
every alarm fires on a synthetic row and does not fire on a healthy one; a built alarm with no row to trip it, or reading a key the healthy row lacks, fails the suite; patience is honoured; a halt alarm raises AlarmHalt and writes a bundle |
fast |
test_coordinator.py |
a refused batch (short, non-finite, illegal) is never trained on; an update that raises leaves no checkpoint of the weights it moved; a crash while collecting keeps the checkpoint already on disk; a keyboard interrupt during collection saves the learner the last row describes; a halt leaves its row on disk beside its checkpoint; a resume refuses a checkpoint no metric row describes | fast |
test_checkpoint.py |
every component round-trips; with strict=False a missing file prints its path and continues, with strict=True it raises; a flipped byte is detected by the manifest; a higher format_version is refused by name; pruning keeps exactly keep and survives a stray file and a .partial directory; the write is atomic under a simulated crash, on a platform where a directory handle cannot be fsynced as well as on one where it can; a checkpoint no metric row of its iteration describes is refused, and one with nothing to compare against is let through |
fast |
test_rng_roundtrip.py |
torch CPU, the numpy generator and python random round-trip and reproduce their next 1000 draws |
fast |
test_no_global_rng.py |
a source scan finds no module-level np.random.<func>, no bare random., and no unseeded torch.rand* in royalelearn/ |
fast |
test_env_contract.py |
every layout fact the harness relies on, each assertion naming the RoyaleGym file it depends on: the action space is Discrete(1 + hand_size*tiles_y*tiles_x) and index 0 is the no-op; encode matches the head's arithmetic; mask_planes == action_mask[1:].reshape(hand_size, tiles_y, tiles_x) wherever the key exists; vector_layout() is contiguous, covers the whole vector and declares the three hand fields; spatial_layout() declares one entry per plane and marks the static ones; observation_space["spatial"].high is per channel; SAME_STEP autoreset with final_obs; the seven terminal scalars under final_info with their validity masks; info["tick"] present and surviving into final_info; mask[NOOP] always set; episodes end in pairs; deploy_status present; config() returns every documented key after a reset. A RoyaleGym change then breaks this file loudly rather than the learner silently |
fast |
test_resume.py |
section 12.4's identical-metric-stream and identical-state_digest proof, in-process and via a subprocess |
slow |
test_resume_identity_guard.py |
changing master_seed, the arch, codec_version, decision_ms or workers each refuses the resume and names that field; --allow-identity-drift writes drift.json and marks the rows |
slow |
test_ratio_invariant.py |
on CPU in fp32, a two-iteration run has ratio_max_abs_dev under 1e-6; perturbing the stored mask, the codec table or the worker's weight version each trips the assertion with the right message and a deviation of order one. There is no GPU test of the bf16 tolerance: what that tolerance is worth is measured by bench on the machine that will run, not asserted here |
slow |
test_replay_episode.py |
an episode replayed from its shard seed and reset ordinal returns an empty divergence list from royalegym.replay.verify_trace |
slow |
test_smoke_train.py |
examples/configs/smoke.json runs three iterations end to end: no NaN, no illegal actions, a checkpoint written and reloaded, ratings refitted, one gate at reduced n, and a metric row with every schema key present and finite |
slow |
test_bench_report.py |
bench runs and produces every field; asserts nothing about the values, which are machine-dependent, but prints them, exactly as RoyaleGym's throughput test does |
slow |
Target: about 52 fast tests under 30 seconds, about 9 slow tests under five minutes. None needs a GPU,
and the torch-dependent tests skip cleanly when torch is absent so that pytest still passes in a
virtual environment that has not finished installing it.
16. What the harness asks of RoyaleGym¶
Filed as issues there, each with the measured cost. None blocks the first run: the harness works against the surface it is given, and each of these removes a workaround or is a pure speed-up.
The surface the harness reads its layout, its identity and its episode statistics from is in place,
which is why sections 4.1, 7.2 and 7.7 describe reading rather than inferring:
ClashParallelEnv.config() and RustEngine.build_digest() for the identity and the ladder context;
ObsBuilder.vector_layout() and ObsBuilder.spatial_layout() for the field offsets and the static
planes; mask_planes and per-channel observation bounds for the codec table; the seven terminal
scalars under final_info for the EpisodeRecord; per-team reward terms and float32 rewards for the
metric stream; and ClashSelfPlayVecEnv(..., viser=...) for the state stream. What remains open:
| # | ask | cost today | size |
|---|---|---|---|
| 1 | Rust-backed default observation and mask (RoyaleGym's own job 7) | 200 of 298 microseconds per transition, 67% of the rollout side; landing it takes rollout capacity from about 3 300 to about 9 000 timesteps/s and frees two cores | large, already planned |
| 2 | ClashSelfPlayVecEnv(..., copy=False), to drop the per-step copy.deepcopy of the whole batch |
14 microseconds per transition, 387 KB per step | one line |
| 3 | PlacementOracle grid reuse between the mask's grids and the observation's |
up to 7 of 14 point_grid calls per step, about 80 microseconds per game-step |
small |
| 4 | selfplay._Record gains separate wins/draws/losses; record_result(..., eval=False, context=...); atomic save; fit_ratings and is_stronger; an EvictionPolicy ABC whose default is not oldest-first; _evict stops leaking records. Plus the potential-based reward terms of section 10 |
draws are lossy, which breaks every binomial interval; _evict deletes exactly the diverse opponents the uniform floor protects; ElixirTradeReward rewards turtling at a magnitude comparable to the terminal reward |
medium; the harness ships its own until then |
| 6 | call/get_attr/set_attr on the vec env; a counter for entities dropped past max_entities |
reaching into vec.envs[i] is undocumented; a silent truncation during training is the class of bug docs/design.md warns about |
small |
Ask 1 lands as a drop-in speed-up, not a rewrite, because nothing in this design touches the
Python builders' internals. It uses only single_observation_space, the Dict key names, the two
layout methods and the mask contract. That is the property to protect in review.
17. Landing order¶
Seven stages. Each ends green and none leaves a half-wired module in the tree. Nothing in any stage
runs cargo or maturin: the whole plan runs against the already-built royalesim and against
MockEngine, which is deliberate, because it is what lets the middle stages proceed in parallel on a
small machine.
Stage 0, clear the seed. Delete the six vendored modules, SEED_CLASSES, NOTICE,
LICENSE-APACHE-2.0 and the ruff exclusion list; update pyproject.toml (licence, dependencies,
extras, markers); rewrite royalelearn/__init__.py and tests/test_package.py. This lands first
because everything else touches pyproject.toml.
Stage 1, the spine. config.py, seeding.py, determinism.py, identity.py, errors.py,
obs_layout.py, version.py, the whole of api/, rollout/layout.py, rollout/envspec.py,
metrics/schema.py, and their tests. Nothing downstream can start until the ABCs and the byte layout
are frozen, because they are the interfaces everything else is written against.
Stage 2, four independent tracks. Each owns disjoint files and depends only on stage 1.
- Networks:
learn/nets.py,learn/distribution.py,learn/actor_critic.pyand their tests. - Data path:
rollout/codec.py,learn/buffer.py,learn/gae.py,learn/returns.py,learn/schedules.pyand their tests. - Rollout:
rollout/plan.py,rollout/scripted.py,rollout/worker.py,rollout/farm.py,rollout/inline.py,rollout/preflight.pyand their tests.InlineRolloutSourcelands first and is the semantic reference; the farm is accepted only when the differential test shows byte-identical buffers against it. - Ladder, metrics and checkpoints:
ladder/**,metrics/**exceptschema.py,checkpoint.pyand their tests. This track is pure numpy and statistics with no torch and no env, so it is the cleanest one to start first if effort is scarce.
Stage 3, the update. learn/ppo.py, learn/inference.py, rewards.py and their tests. Depends
on the networks and the data path; ppo.py and buffer.py are coupled through the minibatch path, so
splitting them costs more in interface churn than it saves.
Stage 4, the run. coordinator.py, cli.py, __main__.py, metrics/alarms.py,
metrics/bundle.py, examples/**, and the slow tests: resume, the identity guard, the ratio
invariant, replay and the smoke run.
Stage 5, verification. The full suite once, plus royalelearn doctor and royalelearn bench,
with the measured numbers pasted into docs/throughput.md and the README. This is the only stage that
runs the whole sweep; earlier stages scope pytest to the files they touch and report what they
changed rather than proving it with a full gate. Then the documentation pages of section 3, written
against the code as it actually landed, because every number in them is measured here.
Stage 6, the first real run. Twenty-four hours on MockEngine first (no rebuild, and it is
faster), then RustEngine. Acceptance: health/illegal_action_rate exactly zero,
ppo/explained_variance positive by iteration 50, policy/cards_per_match above 10,
env/win_rate_by_seat inside 0.45-0.55, throughput/rollout_capacity_ratio above 1.5, and one gate
passed.
18. What the first run measures¶
Recorded here so that they are measurements rather than arguments. Each has a default that is safe to run with, and each is a number the harness already logs.
- Achieved bf16 throughput and peak VRAM on the 3050. Everything in section 2 scales off the
assumed 9 TFLOP/s and 1.18 MB of activations per sample.
benchanswers it before a single gradient step. If the achieved rate is nearer 4 TFLOP/s the run is 2.2 times slower than budgeted, and the three pre-designed escapes, in order of preference, arenet.channels64 to 48 withblocks4 to 3, thenppo.n_epochs3 to 2, thenrollout.overlapon. The decision is made from the logged timing split, and the chosen values go into the identity so the two regimes are never confused in a plot. - Worker resident memory with 32 battles. If it exceeds about 250 MB, lower
games_per_workerbefore loweringworkers: the inference batch matters more than the worker count. - The draw rate in self-play. It decides Davidson against half-a-win in the rater and it changes
the gate's variance. Measurable today with
RandomLegalOpponentself-play, cheaply. - The paired-seed correlation
rho. It sets the real effective sample size for the gate and it says how much of a battle the start state decides. - The transitivity residual. If self-play here produces genuine rock-paper-scissors between
snapshots, the scalar release metric needs replacing, and the
RaterABC is the seam for it. - Whether the no-op guard was needed at all.
ent_coef_noopanneals from a small positive coefficient to zero over the first ten million env steps, andrun/ent_coef_noop,ppo/noop_entropyandpolicy/cards_per_matchare all logged, so the run says plainly whethercards_per_matchheld up on its own as the coefficient fell, whether it needed the guard for longer, or whether it never came near collapsing. If the last, the term's default becomes zero and the alarms carry the job alone. - Whether frame stacking buys anything. The observation carries no motion, so
obs.frame_stackis the knob that gives the policy a direction of travel.k = 2costsS + hand_sizeextra stem channels and no storage; the comparison is two runs to the same env-step count with everything else held, read offpolicy/tile_top1_shareand the win rate against a fixed anchor. n_epochsandbatch_size. Both are marked measure-first and both have a stated rule: if epoch three's clip fraction is more than twice epoch one's, lowern_epochs; if the per-iteration KL sits below 0.003 lowerbatch_sizeand if it sits above 0.02 raise it.decision_ms. 500 ms is the environment's default and the highest-leverage cost knob: 1000 halves the cost and loses the timing granularity that decides Clash fights, 250 costs four times as much. Ablate it in the second run, not the first, because it is the parameter most likely to be blamed for a plateau that is really something else.- How much of a batch can carry a gradient at all. It is measured, and the answer is 6%. The
first two real iterations on the laptop profile reported
policy/forced_noop_fracat 0.937 and 0.924 [M], exactly equal topolicy/noop_ratein both, withpolicy/legal_actions_meanat 29 of 2305. So in 93% of collected decisions the mask leaves exactly one action, the no-op, because the elixir bar cannot afford anything. Those rows are not a policy choosing to wait; they are a policy with nothing to choose, and they contribute exactly zero policy gradient while occupying a full row of the rectangle, a full share of the boundary's bandwidth and a full share of every epoch.
It is visible in everything downstream and explains all of it: ppo/kl at 4.8e-06 then
3.6e-07, ppo/clip_fraction at 1e-04 then exactly 0, ppo/grad_norm_actor at 0.003,
ppo/entropy_normalised at 0.07 against a healthy band of 0.3-0.8. That last one is an
average over rows whose legal set has one member and whose entropy is therefore zero. A
reader who saw only the KL would conclude the update was broken. The update is fine; the
batch is 94% padding.
And the padding is where the wall clock goes. The same iteration measured
time/collection at 12.6 s against time/update at 531.1 s, so 97.7% of the iteration is
the update [M]: three epochs over a batch that is 93% rows the policy cannot learn
from. (The first iteration reads 84% because it pays for cuDNN's first look at each shape:
time/inference 76.6 s then 10.2 s.)
That is the whole of the measurement. The decision it feeds is section 18.1: ppo.forced_rows,
which is about what the UPDATE does with a forced row and not about whether one is collected.
Two other responses stay open beside it. Raising decision_ms is item 7's ablation and is the
only one of them that changes what the agent is rather than what the learner does with it; never
storing a forced row is answered in 18.1 under what was considered and not adopted.
9. The trunk's discrimination, against its input's. Representation collapse is the encoder
producing nearly the same embedding for boards that differ. If the harness ever alarms on it, the
threshold must be relative, never absolute, and it must be built to three rules that a naive
version of it breaks.
Relative, because the observation is mostly constant. Across eight boards differing in one unit's position, a destroyed tower and the elixir, only 51 of the observation's numbers move, which is 0.4%, and the greatest pairwise cosine is 0.9998 with nothing wrong [M]. The rest is static arena, standing towers and an unchanged hand. An embedding cosine of 0.99 on that input would mean the trunk was increasing discrimination, not losing it.
On the same states, not on a sample. Because so few cells move, the input baseline is dominated
by which boards were chosen: pairs differing only in elixir barely move it, a pair with a tower
down moves it a great deal. Two people measuring one encoder against baselines drawn from
different state sets will disagree about that encoder and both will be right about what they
measured. The baseline is computed on the states the encoding was computed on, or the two numbers
are not a pair. royalegym.measure_variability(builder, states, action_masks) returns the input
side: cells, varying, fraction, raw cosine and varying-cell cosine. It excludes the action
mask, which is legality rather than representation.
Report both cosines, raw and varying-cell. Under a planted total collapse, with the board removed from the observation entirely, the raw cosine moves by less than 0.01 while the varying-cell cosine goes to 1.0 [M]. The number a naive detector would watch is the number that does not move.
One thing the measurement is not for: ranking two different observation builders against each other. It compares a representation against itself, and across representations of different sparsity it is not measuring the same property twice.
18.1 Forced rows: what the update does with them¶
Item 8 measured the problem. This subsection holds the decision it feeds, the field that carries it, the test that settles what is left of it, and the part that is designed and not built.
A forced row is a trainable cell whose stored mask has exactly one legal action. A choice row has more than one. Neither is a policy that chose to wait.
What decides it¶
- A forced row gives the actor exactly zero gradient. A masked log-softmax over one legal entry
is exactly 0.0 -- the other entries are filled with
finfo.minand drop out of the normalisation, and the logits are float32 before the masking -- so the stored log-probability of such a row is the literal zero and its ratio is the literal one. Its entropy is 0 and its play/wait entropy is a clamped constant with no gradient. The design panel measured the actor gradient from forced rows alone at exactly 0 on a network built from the real classes, with dropping them moving the gradient by 2.9e-8 on a norm of 0.038 in float32 and by 1.2e-7 on 0.055 under bfloat16 autocast [M]. Every normalisation in the trunk is per sample, so a forward over a subset computes the same numbers for the rows in it. - The critic and GAE still need every row. At
gamma0.997 andgae_lambda0.99 the cells after a choice cell carry gamma(1 - lambda)/(1 - gamma lambda) = 0.77 of its bootstrap weight, and most of those cells are forced [A]. A forced row is not a decision, but it is a state the critic is read at. - The actor trains in Adam's eps-bound regime. In the optimizer state of a laptop-profile
checkpoint after 72 Adam steps, 99.2% of the actor's 415,428 coordinates had sqrt(v_hat) below
adam_eps1e-5, median 1.0e-6, with every trunk tensor entirely below eps; the critic's share was 11.6% [M]. Below eps, Adam is SGD at lr/eps, so the actor's step is proportional to the scale of its loss. - A mean over every row is therefore a learning rate the elixir bar sets. The actor's three terms
are sums over choice rows divided by the batch's row count, so their scale is the choice fraction
c = 1 -
forced_frac, measured between 0.09 and 0.18 over eight iterations [M]. Adam does not remove a rescale that moves, at any eps: the first moment remembers about ten steps and the second about a thousand, so the step follows c now divided by the root mean square of c over the last thousand [A]. Replaying the measured moments, dividing by the choice rows instead makes the actor's step 4.08x larger atforced_frac0.925 and 3.44x at 0.88 under eps 1e-5, and 1.02x under eps 1e-8 [A]. adam_epsis the larger lever and it is not this field. An interleaved same-seed A/B of 1e-5 against 1e-8 moved the actor 11.1x, 12.5x and 11.1x further per iteration at the same update time, withppo/klunder 4e-5 and nothing clipped [M]. That default is decided on its own; this field is built to be right under either value.- Where the update's time goes. The update costs 254-262 ms per 256-row minibatch across a
fourfold range of iteration size, so it is linear in rows [M]. Per 256 rows the actor's forward
and backward is 323 GFLOP, the critic's is 291 GFLOP and the no-gradient critic pass is 97 GFLOP,
all counted with torch's
FlopCounterModeon the laptop network and the RustEngine spec [M]. The actor is 53% of the per-row compute [A].
The field¶
ppo.forced_rows is a string that check_consistency validates against three values. It sits in the
PPO block, which algo_digest hashes whole, so a run's identity names it and a resume across a change
is refused by name unless --allow-identity-drift is passed.
| value | actor forward and backward | actor loss terms | advantages standardised over | critic and GAE |
|---|---|---|---|---|
all |
every row | sums over the batch, divided by its row count | trainable, valid cells | every row, per tick |
critic_only |
choice rows | sums over the batch's choice rows, divided by its row count | trainable, valid cells | every row, per tick |
critic_only_choice_mean |
choice rows | sums over the batch's choice rows, divided by their count | trainable, valid choice cells | every row, per tick |
all is today's update. Its arithmetic and its batch order are unchanged, so every earlier run
reproduces. It is the control, and after phase 1 it is never the default.
critic_only is an exact speed-up.
- Each epoch keeps
all's permutation and batch boundaries, so every batch holds the same cells. Inside a batch the cells are stably partitioned, choice rows first, then cut into minibatches. The actor therefore runs on whole minibatches of choice rows rather than on a tenth of every minibatch. - The critic runs forward and backward on every row at weight 1/n_batch, as today.
- The actor runs on each minibatch's choice rows, and its three terms are sums over those rows divided
by n_batch. The rows it skips contribute exactly zero, so its gradient is
all's to rounding. - A batch with no choice row gives the actor a zero gradient rather than none, so Adam takes the same
momentum step it takes under
all. At a 16-row remainder batch and c = 0.1 that happens about 18% of the time [A], which is often enough to move the two optimizer trajectories apart. - Value function and GAE: unchanged. The recursion, the returns, the potential shaping's telescoping and the 77-91 tick (38-46 s) horizon all see byte-identical inputs.
- Compute: the actor's share times the forced fraction, less the fixed cost of gathering. That is
0.55-0.65 of today's
time/updateatforced_frac0.88-0.93 [A]. - Measured on MockEngine and the CPU, at a planted
forced_fracof 0.899 over 1,024 rows, with a 32-channel two-block network andminibatch_size256: the epochs went from 4.02-4.22 s to 2.08-2.26 s, 0.52-0.55x, and the whole update, the critic's pass and the recursion included, from 4.84-4.99 s to 2.95-3.04 s, 0.60-0.61x [M]. Two repeats ofpytest -m slow tests/test_ppo.py -k faster_update -s, best of three timed runs per arm, on a machine that was also running a training job. It is not the figure phase 1 will report: this network is small enough for a CPU, so a larger share of each minibatch is the gather and the copy, which neither value skips, and the update here is one epoch rather than three.
critic_only_choice_mean has critic_only's compute and changes one thing: the actor's
population becomes the choice rows, for every statistic its loss uses.
- Each actor term is a sum over the batch's choice rows divided by n_choice(batch), so a choice row weighs 1/n_choice(batch) whatever the minibatch partition.
- The advantages are centred and scaled over the trainable, valid choice cells. That is
standardise's own rule, the statistics of the cells that reach the update: under this value a forced cell reaches only the critic, which never reads an advantage. - The critic keeps 1/n_batch over every row. The KL, clip and entropy diagnostics already condition on
choice rows, so
lr_backoffreads the same quantity it read before. - Value function: no direct change. The critic's rows, weights, targets and pass are identical to
critic_only's. The only indirect change is the state distribution a steadier policy produces, which is why explained variance is the first guardrail. - GAE: the recursion, the returns and the horizon are unchanged. What changes is how the actor reads the advantages, with a baseline and a scale taken over choice rows. That is a variance change, and it moves the entropy terms' weight against the policy term by std_all/std_choice, which is logged.
- The point is that the actor's step size stops following the elixir bar. At eps 1e-5 that also makes the step 3.4-4.1x larger; at eps 1e-8 the steady-state step is about the same and only the drift goes [A].
Both skipping values need net.separate_trunks, which is the default, and the config refuses them
otherwise and says why: under one trunk the critic's forward IS the actor's, so a skipped row saves
nothing and its value loss still reaches the parameters the policy gradient uses. Only the policy head
could be left out, and a choice-row mean would then reweight value against policy inside the trunk by
1/c.
How a row is classed, and what checks it¶
The whole-iteration critic pass already unpacks every cell's mask, so it returns the mask's count
beside the value and the buffer keeps it as an int16 (T, R) column, cleared with every other scalar
when an iteration opens. It is filled under all three values, from the same stored bytes the update
applies.
- In debug iterations the update asserts the column equals each minibatch's unpacked mask count. The column is written from the critic's pass and read back through the minibatch gather, which sorts its cells, so a drift between the two would put one row's count beside another row's observation.
- On every iteration, the critic pass checks the no-op bit of every trainable cell, on the device, and
reads the result once after the pass, before the first minibatch and so before any optimizer step. A
cell with the bit clear raises
MaskedCategorical's own message, built from that cell's mask. The minibatch loop then constructs its distributions without the per-construction check, because that check is a device read-back: one a minibatch, each draining the stream, and it was the last one in the loop once the gather stopped synchronising (2026-09-26). What still paces the loop is the staging ring: a slot is handed out again only after the event recorded behind its last copies, so the host runs at moststaging_slotsminibatches ahead of the device. The minibatches decode the same bytes the pass checked, so no row reaches the loss unchecked. In debug iterations the loop keeps the check as well; an extension actor term's loss runs with it on, because a term may run the actor on rows the pass never saw; and every caller outside the loop but the rollout (evaluation, a model used on its own) keeps it at construction. The rollout's forwards build their distributions without it too: they read the no-op bit back in the same copy as the actions, and a row without it stops the round there. A cell that is not trainable never reaches the loss, whatever its bytes hold: the bootstrap row, a cell no worker ever wrote (zero bytes, no-op bit clear), or one a dead worker left holding an earlier iteration's row. So the check covers the trainable cells, and a failure names how many there are and the first sixteen as (cycle, slot). - On every iteration, two numpy comparisons check every trainable forced cell: its stored log-probability must be exactly 0.0 and its action the no-op. Anything else raises and names up to ten cells. This is the evidence for the zero-gradient argument the skip rests on, and it is the only part of that argument that can be checked without a forward.
- The ratio invariant then reads a first minibatch of choice rows, up to
minibatch_sizerows that can move, where underallit sees about a tenth of that.
What stays true¶
- Minibatch invariance. Every denominator is per batch and never per minibatch, so
minibatch_sizestays a pure memory knob.tests/test_ppo.pyholds it under all three values. - Determinism. The partition is a function of the epoch's permutation and the mask bytes.
- Resume. Nothing new crosses an iteration boundary.
- Collection. Untouched. The farm-against-inline byte identity,
TrainableRowsShortandcumulative_timestepsdo not move. - Peak VRAM. The first minibatch of a batch runs both trunks on
minibatch_sizerows, which is what the preflight probe measures.
New metrics¶
ppo/forced_frac, exact, from the column.policy/forced_noop_fracstays; it is a rollout sample of the same quantity, and the two are read against each other.ppo/actor_rowsandppo/actor_forwards, per iteration over all epochs. These are how a reader sees the arm working: underall,actor_rowsis n_samples x n_epochs exactly.ppo/policy_loss_choice, the surrogate over choice rows, comparable across the three values, and the quantitycritic_only_choice_meanactually optimises.ppo/policy_lossis the whole-batch mean under ALL THREE values, with the skipped rows' constant surrogate added back analytically, so it stays comparable withall's. Read the two the right way round: on a 48-row rectangle with 11 choice rows,critic_only_choice_meanreportspolicy_loss-1.4277 andpolicy_loss_choice-3.25e-08, and a reader who takes the first for the objective that value is minimising is out by seven orders of magnitude. This paragraph said the opposite until the verifier measured it.ppo/explained_variance_choice,ppo/advantage_std_choice_pre_normandppo/advantage_mean_choice, computed under all three values so that the arms report the same keys.
Two groups from the design are not built. ppo/actor_adam_eps_bound_frac and
ppo/critic_adam_eps_bound_frac -- the share of coordinates with sqrt(v_hat) below adam_eps, one
reduction over each optimizer's state per iteration -- belong to the adam_eps item and are a phase-2
mechanism readout rather than something phase 1's forced_rows comparison (section 18.1) decides. time/update_actor and time/update_critic need CUDA
events around each half of the update; under all the backward is fused, so they would be absent
rather than zero, and the split is smdp's prerequisite rather than this field's.
The per-minibatch .item() in the epoch counters has gone with this work. It put a device
synchronisation between every backward pass and the next forward, which the diagnostics' own docstring
forbids, and it would have biased every arm's timing by a different amount.
The default, and the rule that moves it¶
- The field lands with
allas its default, so the landing commit changes no number. - Phase 1 flips the default to
critic_only. It needs four things: the pairing holds, iteration 1 agrees withall, the invariants stay silent, andtime/updateis lower in both repeats. The predicted ratio is 0.55-0.65 [A]. A ratio above 0.8 is recorded as a model miss and explained beforesmdpis considered, becausesmdp's saving rests on the same split. - Phase 2 decides between
critic_only(S) andcritic_only_choice_mean(C). C is adopted when its guardrails hold in both pairs, or in two of three after a split. A guardrail breaks when C fails one of these in a pair where S does not: - C's mean
ppo/explained_varianceover iterations 51-100 is more than 0.05 below S's; - C has an
lr_backoffevent or akl_highalarm; - C's
ppo/entropy_normalisedfalls below 0.3, or C raisesnoop_collapse; - C's score against
scripted:random_legalis more than 0.05 below S's.
Adopted this way, C is recorded as decided on mechanism, with the guardrail table and the number of
training seeds beside it. Better play is claimed only if C scores higher in every pair with
non-overlapping 95% intervals.
4. While phase 2 is open, item 8's rule "if the per-iteration KL sits below 0.003 lower batch_size"
is not applied, because this field moves the KL by the same order. A kl_dead warning under S at
eps 1e-5 is expected and is not a verdict.
The A/B¶
Phase 0, CPU. pytest -q tests/test_ppo.py tests/test_buffer.py tests/test_config.py
tests/test_identity.py tests/test_metric_names.py passes.
The configuration, both GPU phases. examples/configs/laptop.json with only these overrides,
identical across the arms of a pair except the field:
ladder.mix[1.0, 0.0, 0.0]. A pure mirror also makes the learner seat count exact, soTrainableRowsShortcannot fire.env.state_mutator.kwargs.decks: one mirror deck on both teams.ppo.timesteps_per_iteration16,896 andppo.batch_size5,632. At 3 workers x 32 mirror battles that is 88 cycles x 192 rows = 3 x 5,632, so no epoch ends in a remainder batch. Once the equal-batch fix lands, use 16,384 and 4,096.ppo.adam_epsat whatever the default is when the phase starts, recorded, and never varied inside a phase.- A unique
run_nameper run, and amaster_seedshared by the arms of a pair.
Phase 1, GPU, about 25 minutes on a quiet machine. Order: all, critic_only,
critic_only_choice_mean, then the same three again; two iterations each, one master seed.
- Pairing: iteration 1 is collected before any update, so its
policy/*andenv/*keys must match to the last digit across all six runs. If they do not, stop -- the pairing is broken and nothing below means anything. - Equivalence on iteration 1,
critic_onlyagainstall:ppo/grad_norm_actor,ppo/grad_norm_critic,ppo/value_lossandppo/explained_variancewithin 2% relative;ppo/kl,ppo/clip_fraction,ppo/entropy,ppo/noop_entropyand bothupdate_magnitudekeys within 10%;ppo/ratio_max_abs_devwithin the bfloat16 tolerance in both arms; the forced-row invariant silent;ppo/actor_rowsexactly n_samples x n_epochs x (1 -ppo/forced_frac);health/vram_peak_mbno more than 5% aboveall's. - Timing:
time/updatefor every run and iteration, withtime/critic_passandtime/gaebeside it. - Mechanism on iteration 1, C against S on identical data [A]: the
update_magnitude_actorratio, predicted 3.4-4.1x at eps 1e-5 and 0.8-1.3x at 1e-8; theppo/klratio, about 12-17x at 1e-5 and about 1x at 1e-8. A ratio more than 1.5x outside its prediction stops the work until it is explained.
Phase 2, GPU, about 4.5 hours plus evaluation [A]. Prerequisites: phase 1 has flipped the default, the transition-alignment fix below has landed, and the remainder batch is gone -- by the equal-batch fix or by the geometry above.
- Arms: S =
critic_only, C =critic_only_choice_mean. Their compute is equal, so a match in env steps is also a match in wall clock. - Pairs: two master seeds, interleaved S1 C1 S2 C2, 100 iterations each. A third pair only on a split verdict. The claim rests on two training seeds, or three, and says which.
- Outcome: each run's score against
scripted:random_legal, with its 95% interval and the paired difference. Use the live-learner probe every 25 iterations if it has landed; otherwise evaluate a snapshot of the final actor, whose seed set is a function ofmaster_seedand the count, so the games pair across the arms. - Readouts at matched iterations: the guardrails above; the slope of log
update_magnitude_actoragainst log(1 -ppo/forced_frac) across iterations 2-100, predicted positive in S and near zero in C;ppo/advantage_mean_choice;ppo/advantage_std_choice_pre_normagainstppo/advantage_std_pre_norm;ppo/clip_fraction; and for healthpolicy/cards_per_match, the no-op rate on choice rows (policy/noop_rate-policy/forced_noop_frac) / (1 -policy/forced_noop_frac), and every alarm.
What the tests pin¶
Forced rows are planted by rewriting chosen cells' action_mask to the no-op alone before packing, so
the rollout samples under that mask, the stored log-probability comes out of it, and the update reads
the same mask back.
- The saving is measured, not argued. A slow test plants the real forced rate on a rectangle large enough to time and reports the epochs' seconds under both values; it asserts only the direction, because the size of a saving is a number to read rather than a threshold to fail somebody else's machine on.
- The gradient is
all's. With the recording optimizer, in float32, every actor and critic gradient undercritic_onlyequalsall's within a small fraction of its own scale, over two minibatch sizes and a partition whose choice rows cross a minibatch boundary. - The actor's rows are counted. A counting wrapper on the actor's forward sees the choice rows
once per epoch under both skipping values and every row under
all; the critic sees every row under all three.ppo/actor_rowsequals the count. - Forced padding does not move the choice-mean gradient. Slots whose every cell is forced leave the
choice cells' recursions untouched, so under
critic_only_choice_meanthe actor's gradient is unchanged by them, while under the other two it scales by the ratio of row counts. - The parameter delta after one real Adam step, which is item 8's own test, with
advantage_standardizationoff to isolate the denominator. Withadam_epsfar above the largest gradient coordinate -- the laptop's regime -- the actor's delta undercritic_only_choice_meanis n_batch/n_choice timescritic_only's. Withadam_epsat 1e-12 the two are equal to rounding, because Adam's first step is lr x sign(g); that is what shows the difference comes from eps. The critic's delta is identical across all three values in both cases. - Minibatch invariance, the accumulated-gradient test, under all three values.
- The forced-row invariant fires, on a planted log-probability and on a planted action, naming the cell that was planted and not a sibling.
- Standardisation. Under
critic_only_choice_meanthe standardised advantages have mean 0 and standard deviation 1 over the choice cells; under the other two, over the trainable cells. The advantages and returns before standardisation are identical across all three values. - A batch with no choice row. No actor forward runs, the actor's gradient is zeros rather than
None, no NaN appears, both optimizers step, and the actor's parameters land where
allputs them. - The critic is untouched. Its gradients agree across the three values, and so do
ppo/value_lossandppo/explained_variance. - Config and identity. An unknown value is refused by name; both skipping values are refused with
net.separate_trunksfalse and the message says why; changing the value changesalgo_digest; the default isalluntil phase 1's commit changes it and that test with it. - The column. It equals the unpacked mask count on every trainable cell, it is cleared by
begin_iteration, and the debug assertion fires on a planted mismatch. - Names. Every new key is in
schema.pyand intests/test_metric_names.py.
Designed, not built: smdp¶
A fourth value, on the second axis: what the critic and GAE do with forced rows.
- Forced cells stay in the rectangle as frame history, the bootstrap row and statistics. The actor
never sees them, as under
critic_only_choice_mean. The value side folds them into the transitions between a slot's consecutive choice cells. - A transition runs from a choice cell t_i to the slot's next choice cell, or to its episode's end, or to row T, k ticks later. It carries R = sum over j < k of gamma^j r(t_i + j), a discount gamma^k and a trace lambda^k.
- A termination inside a transition bootstraps from 0, a truncation from gamma^k V(final observation), and the edge from gamma^(T - t_i) V(row T).
- One backward loop over T with five carries -- discounted reward sum, discount, next value, next advantage, trace -- computes A = R + gamma^k V_next - V + gamma^k lambda^k A_next at the choice cells.
- Both gamma^k and lambda^k are required. A per-decision gamma changes the objective, and since a stretch's length is set by the policy's own spending it would reward dumping elixir to slow the clock; a per-decision lambda stretches the horizon from 77 ticks to about 238 [A].
- The critic trains on the choice cells plus a uniform sample of forced cells of equal count, p = min(1, n_choice/n_forced), from a named stream addressed by iteration. A sampled forced cell's target is its own SMDP lambda-return, which never uses its own value. The critic pass covers those cells, row T and the truncation finals. The sample is part of the design and not an option: V(row T) and V at the truncation finals are read at forced states 82-94% of the time, and V(row T) carries about 0.6 of a cell's advantage weight at 88 cycles [A].
- The reward scaler,
cumulative_timestepsandTrainableRowsShortstay on every learner tick. - Probe-checked [M]: with every cell a choice it is bit-identical to
gae_recursion; at lambda 1 it equals per-tick GAE at the choice cells with the forced values scrambled (3e-15); at lambda 0.99 the choice-cell advantages do not depend on forced values at all; on random data it sits about 5% from per-tick GAE. - Saving: the critic's half shrinks too, to roughly a third of
critic_only's update. That is a model [A] with an assumed gathering share, which phase 1's split replaces with a measurement. The update's cost would then follow choice points per game second, which the elixir economy sets, rather thandecision_ms. - Why it waits: it depends on the transition-alignment fix; it changes the side of the update named as
the risk, so it is tested against a settled control one change at a time; it touches
gae.py, the critic pass, a new random stream and the population of four published keys; and its saving is unmeasured. It also needs the forced flag before the critic pass, from a mask-only pass or a collection-side flag. - Trigger, all three: phase 2 has decided; the alignment fix has landed; and on the adopted value
time/updateis still more than half oftime/iterationon the laptop profile. Then it gets its own item and runs phase 2's protocol against the adopted value, with three more guardrails: explained variance on choice cells within 0.05, explained variance on a held-out forced sample above 0, andtime/updateat most 0.5x.
Considered and not adopted¶
- Dropping forced rows at collection. Never storing them saves host memory, not time, since
workers pack while the parent waits. It breaks fixed-stride frame stacking, the lockstep rectangle
the farm test certifies, the bootstrap row and the per-tick chain, and it removes most of the
critic's states. Its only gain over
critic_onlyis the critic's half, whichsmdpgets without those breaks. - Fast-forwarding forced stretches inside the worker. A battle steps both seats together and the
other seat is often choosing, so the rectangle's time axis and
run_exactwould both break. - Raising
decision_ms. It changes the agent. It stays the second run's ablation, item 7. - The denominator alone, with the actor still forwarding every row. It pays the whole compute for
critic_only_choice_mean's gradient. - A denominator per minibatch. It makes the gradient depend on the partition.
- A weight fixed for the iteration, 1/(c x
batch_size). Same mean effect; the per-batch count mirrors the critic's per-batch mean and needs no extra pass. - Skipping forced rows in the rollout's forward. A collection-side knob, not bit-identical under bfloat16 because the batch shape changes, and worth 0.5-1.5 s an iteration [A].
- A separate actor batch that decouples actor steps from critic steps. A second hyperparameter change inside one A/B.
- Tuning
adam_epshere. Its own item; phase 2 runs at whatever that item sets. - A separate field for the standardisation population. Only if phase 2 rejects C and
ppo/advantage_mean_choiceunder S is far from zero.
Found on the way: the transition alignment [M]¶
Reproduced in this tree on MockEngine at max_steps 7, decision_ticks 10, one worker, two battles:
the row an episode ends on holds tick 0 -- the next episode's first observation -- with the truncation
flag and the episode end set, and the row before it holds tick 60. The episode record's undiscounted
return, -0.0093000003, is the sum of rows start+1 to end (-0.0093000010) and not of rows start to
end-1 (-0.0094000008).
The rollout writes each round's reward, done flags and episode end on the row of the observation the
step arrived at, which its module docstring says and its tests hold. The update reads row t's reward
and flags as the transition out of row t, with the recursion cut where the row is flagged ended. So
every reward is credited one row late; each episode's first cell, usually a choice at starting elixir,
is scored as the previous battle's ending with its own future cut off; an episode's last cell
bootstraps from the next episode's first state; the transition arriving at row T is never recorded, so
one row in T of the rewards is lost; and frame stacking's liveness test reads the end flag one row
early, which is latent at frame_stack 1 and live at 2.
Each side's tests pass against its own convention, which is why neither saw it. It changes no value of this field and does not confound phase 1, whose arms see identical GAE inputs. Phase 2 waits for it: a learning comparison run on a credit assignment the harness is about to drop measures a harness that will not exist. It is its own item -- record the trailing round, write each transition's scalars and its final observation on the row of the action that caused it, and read liveness at the next row.
19. Extensions, and the freeze¶
A run's config is RoyaleLearn's own keys plus, for each extension the run uses, one top-level
section named after it. An extension is an installed package that declares the section under the
royalelearn.extensions entry-point group; load_config asks for the providers of exactly the
keys RoyaleLearn does not own. What an extension can do is royalelearn.extensions.Extension and
nothing else, and what it may import from RoyaleLearn is royalelearn.extensions.__all__.
extensions.md builds one from start to finish.
Learning from demonstrations is one such package, RoyaleImitate, with two sections: warm_start
(starting weights and the freeze schedule) and imitation (reference policies and the
reference-KL regulariser). Its contract is RoyaleImitate's docs/spec.md, which keeps the numbers
this section had: 19.2-19.4, 19.6-19.12, 19.14 and the imitation half of 19.15 are there. What
stays here is what any extension can use: the sections themselves (19.1), the freeze (19.5) and
export (19.13).
When no section is present nothing here runs, and a run's config.json, identity and run_id are
exactly what they were before sections existed. That is tested.
19.1 Extension sections¶
- Discovery. A key RoyaleLearn does not own is looked up among the installed distributions'
royalelearn.extensionsentry points, every copy of every distribution walked. Refused, by key and all together: a key nothing provides (null or not), two providers for one key, one key declared by two distributions, two copies of one distribution that disagree, a declaration whose module cannot be imported or was imported from outside the distribution that declared it, another extension API version, a name that is not its key, and a section type that would accept a misspelt field. A null section is absent. - Installed and unused. An extension the config does not name is never imported. It cannot move a config, an identity, an alarm table, a schema or a state digest.
- The config. A config with sections is a subclass of
RunConfigwith the sections after the core fields, in name order.RunConfigitself has no field for any of them. - The identity.
RunIdentity.extensionshas oneExtensionRecordper section the run uses: the digest of the section's identity value (files by content digest, thresholds left out), and the providing distribution, its__version__and its commit. A package with no commit to name (one installed from a wheel) is namedcontent:<sha256>over its files instead (identity.package_content_digest). Its folder is watched for uncommitted edits like the core packages. - The hooks.
verify(the files a section names, before preflight, every start),prepare(before the rollout buffer exists),loaded(after a resume's checkpoint load),actor_lr_scale(19.5),actor_terms(terms added to the actor's loss, section 9's update),alarmsandmetric_schema(the run's own alarm table and schema, section 13). - One release.
load_configdrops"imitation": nulland the sixalarms.imitation_*keys at their old defaults, which every config.json written between 7e93217 and the extension API carries, with a notice; a changed one is refused with its new home.
19.5 Freezing and ramping the actor¶
actor_lr_scale(t), the schedule a section supplies (Extension.actor_lr_scale; at most one
section may), multiplies the actor's learning rate after the backoff has set it. Where it is
zero the iteration is frozen:
- the actor runs no loss and no backward, and its optimizer does not step, so Adam's moments are empty at unfreeze rather than filled with a frozen policy's gradients;
lr_backoffdoes not observe the KL, because there is none;- the ratio invariant still runs when it is due, on a no-grad forward of the first minibatch that has a choice row, and must read exactly one;
- the actor's weights are asserted bit-unchanged at the end of the update;
- every key computed from the actor's forward in the update is absent from the row rather than
0.0 (
ppo/kl,ppo/clip_fraction, the entropies,ppo/policy_lossand their relatives, the actor's gradient norm and update magnitude), andppo/actor_frozenis 1.
The critic trains normally on the frozen policy's rollouts. The freeze length is fixed by the
schedule, not triggered by the critic's explained variance, so a resume lands in the same place.
ppo/ev_at_unfreeze is the explained variance of the first iteration after a frozen stretch,
published on that row, and the critic_unready alarm warns when it is under the threshold
the scheduling section sets (warm_start.alarms.ev_at_unfreeze, 0.3, in RoyaleImitate).
Between zero and one the scale is only a learning-rate multiplier.
19.13 Export and evaluation (L6)¶
royalelearn export --run <dir> --checkpoint <id> --out <folder>writes a checkpoint's actor as an actor artifact, float32, with the run's spec. It can then be an init, a reference or an opponent.- Opponents by path. Anywhere the ladder names an opponent (
probe_opponents, the evaluate command),path:<folder>loads an actor artifact, checked withcheck_compatible. royalelearn evaluate --config <cfg> --run <dir> --checkpoints <ids|all> --opponents <list> --games <n> --out <file.jsonl>plays each checkpoint against each opponent on the run's evaluation environment and fixed evaluation seed set, seat-balanced, and writes one row per pair: the score with its interval, games, draws and seeds. It takes the same machine lock as training and resumes an interrupted output file where it stopped.