royalelearn¶
The training harness.
New, and barely tested
The training loop closed on 2026-09-22. python -m royalelearn train runs end to end, and
it needs the torch extra. Every run so far has been a short test, ten iterations at most as
of 2026-09-22, so treat anything below that touches the loop itself as young code. See
The learner for what that means in practice.
Everything on this page is generated from the docstrings in the code, so it shows what exists
rather than what is planned. It covers the main modules, not every one. LearningCoordinator,
the object that runs a training run, lives in royalelearn.coordinator, which is not on this
page. Read its docstrings in the code.
Some of these need torch, which is an optional extra: pip install -e "RoyaleLearn[torch]". The
configuration tree, the run identity and the rollout workers are deliberately torch-free, and
that is a property the package is tested for rather than an accident.
Configuration and identity¶
You write a config, and the run's identity is computed from it so that a resume can be checked against the run it is resuming.
royalelearn.config
¶
The whole configuration of a run, as one tree.
There is one source of truth. The config object is it: no keyword-argument facade duplicates its
leaves, because two sources of truth for one number is the kind of variable this harness exists
to remove. Every struct forbids unknown fields, so a typo is an error at load rather than a
setting that silently does nothing, and every number carries its unit in docs/harness-spec.md
section 6 beside the reason it has the value it has.
msgspec rather than pydantic because it is already a RoyaleGym dependency, it gives forbidden
unknown fields, tagged unions and a canonical JSON round-trip, and the package's dependency list
stays royalegym, numpy, msgspec.
Profiles are functions rather than files: one per machine class, differing only in numbers, and a JSON config names one and overrides what it likes on top.
ConstantSpec
¶
Bases: Struct
A value that does not move.
LinearSpec
¶
Bases: Struct
start to end over over_env_steps game-steps, then end forever.
GeometricSpec
¶
Bases: Struct
start to end geometrically, which is what a discount wants: 1 - gamma decays
by a constant factor rather than gamma moving by a constant amount.
PiecewiseConstantSpec
¶
Bases: Struct
(env_steps, value) breakpoints, held between them.
ObsConfig
¶
Bases: Struct
What the policy sees beyond the current frame.
The observation carries positions and no motion, so a policy that needs to tell an advancing unit from a retreating one needs more than one frame. Stacking is a gather of buffer rows and costs no extra storage; it is 1 by default because the extra channels cost stem compute and the first run measures whether they buy anything.
RolloutConfig
¶
Bases: Struct
The shape and the failure behaviour of the worker farm.
NetConfig
¶
Bases: Struct
The architecture. Hashed into arch_digest, so a snapshot built from another one is
refused rather than shape-errored.
LrBackoffConfig
¶
Bases: Struct
Drop the learning rate when the KL stays above a threshold.
PPOConfig
¶
Bases: Struct
The update.
AdvantageConfig
¶
Bases: Struct
The discount and the credit horizon.
0.99 discounts the win to almost nothing by the end of regulation; the anneal follows
OpenAI Five's, with an endpoint scaled to a four-minute match. gae_lambda at 0.99 gives
a credit horizon of about 45 seconds, which matches the deploy-push-tower causal chain --
the references' 0.95 gives ten and cannot connect a deploy to the tower it takes.
GateConfig
¶
Bases: Struct
The three conditions a candidate must pass to become the champion.
RaterConfig
¶
Bases: Struct
The Bradley-Terry-Davidson fit.
SeedSnapshot
¶
Bases: Struct
A frozen actor the pool holds from the run's first iteration and never evicts.
path is a folder holding actor.safetensors and spec.json: a run's own snapshot
(<run>/snapshots/<digest>) or an actor artifact RoyaleImitate wrote. A relative path is
read from the folder the command runs in. sha256 is of actor.safetensors, and it is
what the run's identity records, so the same path holding other weights is another run; a
run whose folder does not match is refused, and the refusal prints the digest it found.
SeatDecks
¶
Bases: Struct
One deck the learner trains, dealt where the learner sits, against a field of others.
A mirror battle deals deck to both seats. A pool or scripted battle deals deck to
the learner's seat, which is fixed per battle (even battles blue, odd red), and a deck from
field to the other seat, where a frozen or seeded snapshot or a scripted opponent plays
and whose rows are not trained on. field is drawn uniformly (list a deck twice to weight
it), or None for a random deck of eight different cards. Decks are card NAMES, looked up in
the engine's catalogue at the first reset; shuffle is RoyaleGym's ShuffleMode for a
non-mirror deal. See royalelearn/ladder/seat_decks.py.
cls is the curriculum class each battle's deal is: RoyaleGym's by default, or a subclass
that takes its keywords unchanged (one that also deals forms to the seats holding deck,
say), resolved like any component, so its module must be in
config.extra_component_modules. kwargs are that class's own keywords, merged into the
ones the deal decides; they may not name those (deck, p, mirror_p, seat,
pool, shuffle).
LadderConfig
¶
Bases: Struct
Who the learner plays, and when a snapshot joins the pool.
SinkSpec
¶
Bases: Struct
One metrics sink. JsonlSink is always installed: it is the file the resume test
compares.
AlarmConfig
¶
Bases: Struct
Thresholds for the alarm table of section 13.3.
Recorded and excluded from the run identity: an alarm can stop a run and can never alter a value, so two runs that differ only here produce the same numbers for as long as both run. The severities and patiences live with the alarms themselves; the overrides below are for an operator who wants one of them louder or quieter on their own machine.
DeterminismConfig
¶
Bases: Struct
Section 5.1. The tier is part of the run identity; the thread count is part of the tier.
DoctorConfig
¶
Bases: Struct
The start-up gates that can refuse a run.
RunConfig
¶
Bases: Struct
One run, completely: RoyaleLearn's own keys.
A run that uses an extension (royalelearn.extensions) has one more top-level key per
extension, its section. load_config returns such a config as a subclass of this one with
the sections as fields after these, so a config without sections is exactly this type and
encodes as it always did.
Geometry
¶
Bases: Struct
The iteration's rectangle, derived from the config once and passed around whole.
cycles by n_slots, with slot -> (worker, shard, battle, seat) fixed for the run.
Everything downstream -- the buffer's shape, the slot maps, the matchmaker's plan -- reads
these numbers rather than recomputing them, so the arithmetic of section 2.2 lives in one
place (config.geometry).
default_env_spec(engine=RUST_ENGINE, *, max_steps=480)
¶
The environment a run uses unless it says otherwise.
The truncation is a step limit rather than a tick limit because it is decisions that the buffer counts; 480 of them is a full regulation match plus overtime at the default decision granularity, and it is the cap the episode-length histogram spikes at when two policies settle into the turtle equilibrium.
The reward is this package's potential composition rather than RoyaleGym's
default_reward, whose elixir-trade term pays a player who never commits a card and whose
tower term is one discount short of a potential (rewards.py says why each matters). The
reward is part of the environment spec and so of the context digest: runs trained against
the other one are a different objective and do not share a context with these.
role_counts(mix, n_battles)
¶
How many battle slots play mirror, how many pool and how many scripted.
The mixture is a count of slots, not a probability each slot is drawn from, so this is the
one place it is turned into whole battles. MixMatchmaker decides which slots get each
role and geometry sizes the iteration from the same counts: two callers, one function,
because the whole point of the exercise is that the number the iteration is sized for is
the number the matchmaker fills.
It lives here rather than on the matchmaker only because of the import graph -- the
matchmaker reaches config through api.ladder and identity, so config
importing the matchmaker would close a cycle. The dependency runs the way it already did.
Rounding is resolved in two stages, each to the nearest whole battle with a half rounding up, and the last share takes whatever is left:
mirror = round(mix[0] * n_battles),- the remaining battles split between pool and scripted in the ratio
mix[1] : mix[2],poolrounded andscriptedtaking the remainder.
So the three always sum to n_battles, and each lands within one battle of its exact
share (the proof is short: each stage rounds by at most a half, and the second stage's
input is already off by at most a half). Halves round up rather than to even, because
round(0.5) is 0 in Python and a single battle under the shipped mixture would then be a
rectangle with no mirror in it at all.
Two consequences are worth saying out loud. A share small enough that its quota is under a
half rounds to zero: three battles cannot hold a tenth of one, so that role is simply not
in the rectangle, and the mixture the run reports is the count, not the config. And an
all-mirror mixture keeps every slot, which is what makes (1, 0, 0) collect twice the
rows per cycle rather than one and a half.
geometry(config)
¶
The iteration's rectangle, derived from the config.
Learner rows are the slots whose seat the learner plays: both seats of a mirror battle and
one of every other. Because role_counts makes the mirror battles a count rather than a
share, that is n_battles + mirror_battles exactly -- the same number on every cycle of
every iteration of every run, whatever the master seed.
It used to be int(n_slots * (mirror + (1 - mirror) / 2)), the count's expectation. The
two agree at the shipped laptop profile (96 + 48 = int(192 * 0.75) = 144), and what the
expectation could not do was promise it: the realised count was Binomial(96, 0.5) + 96,
with a standard deviation of 4.9 rows per cycle against an iteration whose slack over the
trainable-row floor was 0.2%, so about one master seed in four halted the run.
The cycle count is then what it takes to reach timesteps_per_iteration learner
transitions, rounded up -- an iteration collects at least what was asked for, never less,
and now the "at least" is a fact about the rectangle rather than about its average.
laptop()
¶
8 GB RAM, 4 GB VRAM, 8 threads. The shipped default.
The learner is the bottleneck on this machine by about 2.5 times, which is why there are three workers and not thirty-two, and why the network is four hundred thousand parameters and not four million.
workstation()
¶
32 GB RAM, 12 GB VRAM, 16 threads.
It asked for rollout.overlap until 2026-09-22, which bought a second rectangle and no
overlap, because nothing in the package honours the flag. It is refused now.
many_core()
¶
64 GB RAM, 24 GB VRAM, 64 threads.
profile(name)
¶
The shipped config for one machine class.
check_consistency(config)
¶
Everything wrong with a config, in one list.
All of it at once rather than one exception at a time: on a machine where building an environment costs a minute, finding four mistakes takes four minutes otherwise.
validate(config)
¶
Return the config, or raise PreflightError naming everything wrong with it.
retired_dropped(document)
¶
A stored config with what the one-release shim drops without a word dropped: the six
retired alarm keys at their old defaults, a null imitation and its null init and
actor_lr_scale. For comparing a checkpoint's config with this run's; nothing is refused
or said here.
load_config(source, *, say=print)
¶
Read a config: a JSON string, a path to one, or an already-decoded mapping.
The document is applied ON TOP of the profile it names, so a file that sets three numbers
gets the rest of that machine class rather than the laptop's, and a file that sets none is
exactly the profile. Unknown fields are refused at every level of the tree: a top-level key
RoyaleLearn does not own is a section, and one that no installed extension provides is
refused by name (extensions.providers_for). A section that is null counts as absent.
dump_config(config, *, indent=None)
¶
The canonical JSON of a config: keys in a fixed order, so a reformat is not a new run. A section set to None is absent, as it is on load.
config_hash(config)
¶
sha256 of the canonical JSON. Recorded, printed and diffed on resume.
royalelearn.identity
¶
What a run IS, as sixteen hex characters.
The identity is everything that can change a number the run produces. Two runs with the same
identity produce the same metric stream under determinism.tier = "run_exact"; two runs with
different identities are different experiments and are never plotted as one. A resume compares
the identity field by field and refuses a difference by name.
What is in it, and why:
master_seed, determinism_tier, device_kind, torch_version each changes the byte stream
the engine build and the card catalogue the game itself
env_spec_digest, obs_digest, action_digest what the policy sees and can do
frame_stack the trunk's input width
arch_digest, codec_version, codec_table_digest the network and how a row is stored
algo_digest the PPO, GAE and schedule settings
rollout_digest workers, games, shards, T and R
ladder_digest who the learner plays
extensions each extension section the run
uses, by content, and the code
that provides it
What is recorded and NOT in it, because none of it changes what the run is: run_name,
runs_dir, timestep_limit, the metrics sinks and their settings, checkpoint.keep, the
alarm thresholds and severities, the doctor's budget, and the knobs that set how the collection
is scheduled rather than what it collects. An alarm can stop a run; it cannot alter a number. A
resume diffs those and prints every difference rather than refusing.
EngineBuild
¶
Bases: Struct
Which engine, built from which data and which code, with which cards.
Every field comes from ClashParallelEnv.config(): calibration_digest over the
calibration values as loaded, build_digest over the copies compiled into the extension,
and binary_sha256 over the extension file itself. The harness hashes no data file of its
own -- a digest computed here from the data directory would be a second opinion about what
the engine is running on, and a second opinion is exactly what a stale build looks like.
stale_build_differences MUST be empty; it is recorded so that a failure is auditable
rather than reconstructed.
binary_sha256 is the one that was missing, and the data digests could not stand in for
it. build_digest hashes the DATA compiled in, not the Rust: across a rebuild from an
edited state.rs on 2026-09-23 it read 8952c1c4aa7d7923 on both sides while the file's
hash went to 04d42611db5d6e5a, so two runs either side of the rebuild wrote identical
blocks and played different games. It is NOT_STATED for an engine with no compiled file,
and NOT_RECORDED -- the default, which only a decode ever reaches -- for an identity
written before this field existed.
ExtensionRecord
¶
Bases: Struct
One extension section in a run's identity: what it says, and whose code reads it.
digest is the section's identity value (Extension.identity_value: files by content,
alarms left out). The other three name the providing package -- distribution, __version__
and commit -- and are None for a section RoyaleLearn provides itself, whose code
royalelearn_git already names.
RunIdentity
¶
Bases: Struct
The identity itself. run_id is the first sixteen hex characters of its sha256.
Encoded with omit_defaults so that a field added with a default of None leaves every
identity that does not use it byte-identical, and so every run_id where it was. Every
other field is always set when an identity is computed; the only ones that can sit at their
default are user_code, on an identity decoded from before it existed, and extensions,
on a run that uses none.
Decoding refuses a field this code does not know. An identity written by a build with a field since removed would otherwise decode with that field silently gone, and two identities that differ only in it would compare equal on a resume.
env_spec_digest_of(config)
¶
The environment's identity: what it will be built as, not how its config was spelled.
action_digest_of(env, env_spec)
¶
What the policy can do: the parser, the space it lays out and the decision clock.
One function for the identity and for anything else that must agree with a run about its action space -- a demonstration shard, an artifact -- so the two cannot compute it apart.
extension_records(config)
¶
The identity's record of every extension section the run uses, or None for none.
The digest is of the section's identity value, which names every file by its content digest
rather than its path; that the digest is TRUE of the file is the section's verify, run at
every start. A section provided by an installed package also records the package: which
distribution declared it, its __version__, and its commit (package_provenance). A
package whose commit cannot be named -- installed from a wheel, or not at the top of its own
repository -- is named by its content instead, content:<sha256> over its files
(package_content_digest), so that two runs on two different builds of it stay two runs.
section_digest(value, *, drop=('path',))
¶
The digest of a config section by value, with every key named in drop left out at
every depth.
What the section's identity is when the section names files: each file is named beside its
path by a content digest, so dropping the paths keeps the content in the identity and keeps
a folder moved to another disk from being a new experiment. A section whose files are named
under another key passes that key in drop.
run_id(identity)
¶
The short name of a run: sha256 of the identity's canonical JSON, first sixteen hex.
identity_differences(a, b)
¶
Every field in which two identities differ, so that a refusal can name all of them.
A field one side never recorded is not a difference: nobody can show the two disagree. It is
not a match either, and unverified is where it is said. Only that one field is set aside;
the rest of the block it sits in is compared as usual.
unverified(a, b)
¶
What a comparison of two identities could not check, one sentence each.
catalogue_digest(cards)
¶
sha256 over the card catalogue's identifying fields, in card order.
Name, id, elixir and placement kind: enough that a catalogue with a card renamed, repriced, reordered or added is a different game, and little enough that a change to a statistic the engine already hashes into its calibration digest does not count twice.
engine_build(env_config, cards, *, path_search=None, stale_build_differences=())
¶
The EngineBuild for a built environment, from its own config() and catalogue.
user_packages(config)
¶
{top-level package: where its source is} for the user code this run's env comes from.
A component counts when its class lives outside the three packages, and so does any string in
a component's kwargs that names a module under extra_component_modules -- which is how a
composition refers to its parts, and the only place such a reference can be. A module that is
merely LISTED is not counted: listing cannot move a number, and the identity table excludes
extra_component_modules for exactly that reason.
The whole top-level package, not the one file, because a reward reads its helpers: an edit to
mybot/util.py moves mybot.rewards.MyReward as surely as an edit to the reward itself.
user_code(config)
¶
The identity's record of the user's own code: {package: sha256 of its source}.
What this closes: the identity named every component by its dotted path, so a different
body of mybot.rewards.MyReward was the same run, and the ladder pooled two objectives.
What it does not: data files a component reads, and code reached by anything other than a
component's class path or a string in its kwargs.
dirty_sources(config_path=None, config=None)
¶
What this run would read that is not in any commit, each named with what changed.
A run's identity records each repository by commit and everything else about the environment
by NAME: the reward, the observation builder and the action parser are dotted paths, hashed as
strings. So two different bodies of royalegym.reward.default_reward produce the same
env_spec_digest, the ladder files their games under one context, and a plot draws two
objectives as one line. On 2026-09-22 an uncommitted change to a sibling repo sat under a live
run for twenty minutes and nothing in the identity could have said so.
What this covers, stated rather than implied: the python packages this process resolved, the
engine's data directory, the config file the run was started from, and -- given the config --
the user's own packages the env's components come from. What it does not: any other file in
those checkouts, a package installed from a wheel rather than a checkout, and anything a
component reads at call time from somewhere else. The identity records the user's packages
by content as well (user_code), so an edit this cannot see because it was committed still
makes a different run.
THE ENGINE'S COMPILED CODE is outside this function, and since 2026-09-24 it is inside the
identity instead. The extension is installed rather than imported from a checkout, so there is
no working tree here to ask. What the engine does publish is the hash of the file it loaded,
and EngineBuild.binary_sha256 records it, so two runs on two builds are two runs. What is
still not published is which COMMIT built that file, so an engine built from uncommitted
source is identified but not attributable. Reaching into a sibling checkout to guess would
put back the guessed path this function removed. First measured and put to the engine's
session by the integrator, 2026-09-22; the binary hash was RoyaleGym's answer.
The first version of this checked git describe --dirty on two hardcoded packages. It
missed the engine's data, missed RoyaleViser, and was blind to untracked files -- including,
at the time it was written, the config of the run then in flight. Found by the integrator,
who measured all three rather than arguing them.
package_provenance(package)
¶
<commit sha> or <commit sha>-dirty of the repository package is tracked in, or
"unknown".
Stricter than git describe from the package's folder, which walks UP into any repository
that encloses it: from inside a checkout's virtual environment it names that checkout's own
commit for an installed copy that is in no repository at all. So a commit is recorded only
when the repository's top level IS the folder above the package and the package's
__init__.py is tracked there. The dirty flag is git status over the package's folder,
untracked files included, as dirty_sources reads it. A sha rather than a description,
because a description changes when a tag is added to the same commit.
package_content_digest(package)
¶
sha256 over every file in package's folder, by relative path and bytes, compiled
caches left out; or "unknown" for a package with no folder.
What names a package installed from a wheel, which has no commit: the same files give the same name wherever they are installed, and any change to one gives another.
royalegym_provenance()
cached
¶
(version, git description) of the royalegym this process imported.
Read from the imported module rather than from a pinned number here, because an editable install from a sibling checkout is the normal case in this family and its version string alone does not say which commit.
torch_version()
¶
The torch this process would train with, or "absent".
describe_device(device='cuda')
¶
"cuda:<name>:sm_86" or "cpu:<machine>".
The device is in the identity because a different GPU is a different set of kernels, and run-exactness is promised within a device class rather than across all of them.
compute_identity(config, *, env_spec, build, arch_digest, codec_version, codec_table_digest, torch_version_string=None, device_kind=None)
¶
The identity of the run config describes, against the environment that was built.
The arguments beyond the config are the facts only preflight knows: what the environment turned out to be, what the codec decided to do with it, and what device and torch build the learner will use. They are passed in rather than probed here so that the function is pure and a test can vary one field at a time.
Rewards¶
Where the shipped reward function lives, and the shape your own one takes.
royalelearn.rewards
¶
The reward the harness trains against: one terminal objective and three potentials.
Every shaping term here is a difference of a potential, gamma * Phi(s') - Phi(s). That form
is what keeps shaping from changing which policy is optimal (Ng, Harada and Russell, 1999), and
RoyaleGym's own reward.py states the principle in its header. The terminal win/loss term is
the only term that is the objective; everything else densifies the signal and must leave the
optimum where it was.
The discount in that expression is the run's, not a constant: set_gamma is called once per
iteration with the value the schedule is at, and the same number rides the Step command to
every worker, so the discount the reward is computed with and the discount the learner
bootstraps with are one number rather than two that agree at the start of a run. It is
deliberately not in config(): config() feeds the run identity and the ladder's context,
and a quantity that moves every iteration does not belong in either.
WHY THESE THREE AND NOT ROYALEGYM'S SHIPPED DEFAULTS
TowerHPRewardcomputesPhi(s') - Phi(s). It is onegammaaway from being exactly policy-invariant, and with the gamma in place the term never needs annealing.ElixirTradeRewardis not a potential and it rewards turtling: a player who never plays a card never incurs the negative half, while enemy units still die to its towers and earn the positive half. At a weight that makes it visible it is comparable to the terminal reward.ElixirLeakPenaltyis not zero-sum -- both players can leak at once -- which is the signature of a term standing in for a potential that is missing.
CommittedElixirPotential replaces both. Playing a unit card moves elixir from the bar to the
board and is net zero (a spell is charged at the tap, a Mirror play loses its own one elixir, and a
Tri Wizards play is not priced right yet: see the class); losing a unit costs what the unit cost;
killing one gains it; and sitting at ten elixir is penalised on its own, because the opponent's
side of the potential keeps rising while yours cannot. It replaces the leak penalty's coefficient
with a property of the state, and it needs no annealing schedule, which is what RoyaleGym's house
rule -- weights should settle, not drift -- asks for. Its weight in the composition is configurable
(see default_potential_reward); that is a scale, set once for a run, not a schedule.
EXACT ARITHMETIC
Every potential is a Fraction and the discounting is done on fractions, with one conversion
to float at the end. The seat-mirror property is the reason: Phi(s, red) is exactly
-Phi(s, blue) as a rational, so the two seats' rewards are exact negatives of each other and
"the reward is zero-sum" is a property a test can assert rather than a tolerance it has to
choose. A running float sum of the same unit values returns plus and minus 1.1e-19 on a perfect
mirror -- harmless to a gradient, and enough to make the property uncheckable.
PotentialReward
¶
Bases: RewardFunction
gamma * Phi(s') - Phi(s) for a potential a subclass computes from one state.
The subclass returns a Fraction from the seat's own point of view, antisymmetric between
the two seats, and this class does the discounting and the single conversion to float.
set_gamma(gamma)
¶
Take the discount the schedule is at. Exact: a binary float is a rational.
potential(state, team)
abstractmethod
¶
Phi(s) from team's point of view, exact and antisymmetric in the seat.
get_reward(team, prev, state, results)
¶
gamma * Phi(s') - Phi(s), with Phi of a finished battle taken to be zero.
The zero is taken rather than computed, and it is what makes the shaping harmless. Summed
over an episode the term telescopes to gamma^T * Phi(s_T) - Phi(s_0); Phi(s_0) is
the same whatever the policy does, so if Phi(s_T) is zero as well the shaping adds a
constant and cannot move the optimum (Ng, Harada and Russell, 1999). Read off the final
state instead, Phi(s_T) is the margin -- a 3-0 win has three times the crown potential
of a 1-0 win -- and the shaping would quietly pay for the margin beside the objective,
which is the one thing these terms were chosen not to do.
A truncation is the other case and it is not this one: game_over is the engine saying
the battle was decided, while a step limit cuts a battle that still has a position worth
something, and the estimator bootstraps from that position's value.
PotentialCrownReward
¶
Bases: PotentialReward
Phi = (own crowns - enemy crowns) / 3.
The crown difference is the match's own scoreboard, so the potential is the scoreboard read as a fraction of a whole win. It is the coarsest of the three and the one that most directly anticipates the terminal term.
PotentialTowerHPReward
¶
Bases: PotentialReward
Phi = (sum own_hp/own_max - sum foe_hp/foe_max) / 3.
The same scoreboard as the crowns, read continuously: a tower at a third of its health is a third of the way to the crown it will give up. Dividing by the tower count puts it on the crown potential's scale, so the two weights mean comparable things.
CommittedElixirPotential
¶
Bases: PotentialReward
Phi = ((own bar + own board) - (foe bar + foe board)) / scale.
bar is the elixir a player is holding and board is the elixir value of what that
player has on the field, a unit at Fraction(card.elixir, card.count) of the card that
summoned it. Crown towers are excluded: they were never played, and the tower potential
already owns what happens to them.
What the term says, in the four cases that matter: a unit card played moves elixir from the bar to the board and is worth nothing; a unit lost costs what it cost; a unit killed gains it; and holding a full bar loses ground, because the opponent's side keeps rising while yours cannot. The last is the one that replaces a leak penalty, and unlike a leak penalty it is zero-sum.
WHICH UNITS ARE PRICED, and why this is not every unit on the board. An engine reports each
entity under a card, and the entity need not be that card's own unit. On the RustEngine measured
on 2026-09-24, a Goblin Gang's spear goblins were filed under
the Goblin Hut, at five elixir each; a dying Golem's golemites under the Golem, at eight; a
dying Battle Ram's barbarians under the Battle Ram, at four. Priced by the card they were filed
under, one Goblin Gang play read +15 elixir of potential, a Golem's death +8, a Battle Ram's +4
-- under every run up to then, at every weight. Today's engine files all six of a Gang's units
under the Goblin Gang; its spear goblins still do not match the Gang's row. So a unit is priced
only when it IS the unit its card's catalogue row describes: same hitpoints, same collision
radius, same air or ground. Anything else a card produced scores zero. That is an
understatement and it is deliberate, the rule RoyaleGym's ElixirTradeReward adopted in
2381149: a golemite is worth something, and this term says zero rather than eight. What it
keeps true is the property the term needs. ONE PLAY OF A UNIT CARD PUTS EXACTLY THAT CARD'S
ELIXIR ON THE BOARD: a card's summon count covers exactly the units its row describes, so a
Goblin Gang still totals three. tests/test_elixir_pricing.py taps every card in each
engine's catalogue, on both seats, and names any card no tap landed for, because the rule rests
on that.
One card breaks it, and it is a strict expected failure in that file: the Tri Wizards. From RoyaleSim round 9 a Tri Wizards play puts its Electro Wizard and Ice Wizard down under their own card ids, and each is its own card's unit, so the play prices at 7 + 4 + 3 = 14 for a card of 7: a Tri Wizards play reads as a gain of 7 while those two live. The default catalogue holds it. It is an event-only card, and its fix waits with the other event-only cards.
A MIRROR PLAY LOSES EXACTLY ITS OWN ONE ELIXIR. It pays the copied card's elixir plus its
own one, and puts down a copy one level above the copied card. The copy is priced as the
copied card's own unit, and the extra elixir buys the level, which this term does not price.
The copy's hitpoints are the higher level's (a mirrored Knight has 1938 against the Knight
row's 1766), so the row is read at the level the engine reports for the unit
(EntityState.level), from the engine's own rows (RustEngine.unit_hitpoints), never
from a guessed level curve. A unit at the catalogue's level is matched against the row
itself, which the engine's rows agree with at that level for every card; the tests check it.
An engine that reports no level (MockEngine, and RoyaleSim before round 9) or gives no rows
(RoyaleGym before RustEngine.unit_hitpoints) is matched against the row alone, and there
a Mirror copy scores zero: while it lives, the play reads as its whole cost thrown away.
test_a_mirror_play_puts_the_copied_cards_elixir_on_the_board taps the Mirror on both
seats.
SPELLS ARE CHARGED AT THE TAP, through the bar, by construction: the bar drops by the spell's
cost and nothing it leaves on the board is priced, so the elixir comes back only through what
the spell kills. That is the true trade, and it is what ElixirTradeReward charges too. A
spell that leaves units behind -- a barrel -- is charged the same way, and its units score
zero by the rule above. It telescopes like everything else here, so it moves no optimum.
unit_value(entity)
¶
What one entity on the board is worth: its card's per-unit elixir, or zero.
Zero for anything that is not the unit its card's row describes -- a spear goblin filed under the Goblin Hut, a golemite under the Golem -- because pricing it by the card it is filed under is how a three-elixir play came to read as eighteen. The hitpoints compared are the row's at the unit's own level, so a Mirror's copy, one level up, is its card's.
PotentialCombinedReward
¶
Bases: CombinedReward
RoyaleGym's weighted sum, with the run's discount threaded into the terms that take one.
The worker calls set_gamma on whatever reward function its environment holds, once per
iteration, so the composition has to be the thing that forwards it. A term that does not
take a discount -- the terminal one -- is left alone.
It also says WHICH of its terms is the objective. CombinedReward files each term in the
logged breakdown under its class name, and the metrics group has to tell the objective from
the shaping to publish either one: shaping_dominates is the alarm that watches for the
shaping taking the objective over, and it compares those two sums. Matching on a class name
would put a rename of a class into the arithmetic of an alarm, so the composition names the
objective instead, and the breakdown carries it under TERMINAL_REWARD_TERM.
terminal_class
property
¶
The class name the objective's term would otherwise be filed under.
set_gamma(reward, gamma)
¶
Give gamma to every term of reward that takes one, however it is composed.
A free function as well as a method because a composition is a tree: a CombinedReward
may hold another one, and the discount has to reach the leaves of whatever a bot creator
assembled rather than only the terms this module shipped.
shaping_weight_problems(weights)
¶
Everything wrong with a set of shaping weights, every one named, all at once.
Shared by default_potential_reward and config.check_consistency so that a config is
refused at load for exactly what the reward would refuse at build.
default_potential_reward(*, crown=0.2, tower_hp=0.1, elixir=0.05)
¶
The shipped composition: the objective, and three potentials under it.
The terminal term is the objective at 1.0; the crown potential anticipates it; the tower potential is the same scoreboard read continuously; and the elixir potential is the fastest-moving of the three.
The three shaping weights are keyword arguments, so a config can set them::
"reward_fn": {"cls": "royalelearn.rewards.default_potential_reward",
"kwargs": {"crown": 0.2, "tower_hp": 0.1, "elixir": 0.05}}
Every term is a potential difference, so no weight can change which policy is optimal: a
weight SCALES a term's per-step magnitude and changes nothing about when it arrives. That holds
for any real weight, negative ones included, so it neither justifies raising a weight nor
bounds it. The bound here is a choice, stated below. What a weight does change is how loud
each term is step to step, and that is what to read before changing one:
env/reward_terms_step_abs/<term>, the mean of each seat's sum |F_t| over an episode.
NOT env/reward_terms_abs/<term>, which is the episode's SUM: for a potential it telescopes
to 1 - gamma times how far the potential wandered, so it falls as the discount schedule
rises whatever the weights are.
How loud the shipped weights are, measured 2026-09-24 on a 100-card RustEngine environment, ten random-legal battles at gamma 0.999, per seat per episode [M]: elixir 0.80, tower 0.15, crown 0.15, against a terminal of 0.80. The elixir term alone is already about as loud as the objective. Each magnitude is linear in its weight.
A weight must be a real number, finite and not negative: not a bool, which Python would
quietly take as 1, and not a string, which constructs and then fails inside a worker on the
first step. A negative potential weight pays a seat for losing ground, and a NaN reaches every
return it touches; both would train, silently. config.check_consistency applies the same
rule when a config is loaded, so a bad weight is named before any environment is built.
Collecting experience¶
royalelearn.rollout
¶
The rollout side: the env description, the shared-memory byte layout, and the workers.
Names resolve lazily. api/rollout.py imports EnvFactorySpec from here, and
layout.py imports EnvSpec from there, so an eager re-export in this file would close
that loop at import time; resolving on first use keeps both directions working whichever module
a caller imports first, and keeps the cost of importing the env description down to msgspec.
royalelearn.obs_layout
¶
The vector fields the network needs, resolved by name.
The pointer policy head reads three runs of the flat observation vector: the one-hot of the card in each hand slot, each slot's cost, and whether each slot is affordable now. Their offsets are not written down here and are not computed from the vector's width. They are looked up by name in the layout the observation builder itself declares, so that a field added, removed or reordered upstream moves them, a field the head needs and the builder no longer emits is a refusal at start-up naming it, and the same code runs against a sixteen-card catalogue and a sixty-five-card one without knowing which it has.
HandFields
¶
Bases: NamedTuple
Where the hand lives in the observation vector, and the shape its one-hot unfolds to.
card_onehot is hand_size consecutive one-hot blocks of onehot_width -- the
catalogue plus one for an empty slot. The width is read off the field's own size divided by
the hand size the action space declares, never from the vector's width or the card count, so
a builder that widens the block says so by widening the field.
field_slice(spec, name)
¶
The slice of the observation vector holding name.
A missing name is a PreflightError that lists what the layout does declare, because the
failure it catches -- a renamed field -- is otherwise a silently wrong slice of somebody
else's numbers.
resolve_fields(spec, names=REQUIRED_FIELDS)
¶
Every named field's slice, or a PreflightError naming the first one that is missing.
hand_fields(spec)
¶
The three hand fields, checked against each other and against the action space.
The learner¶
royalelearn.learn
¶
The learner: the networks, the masked distribution, and everything that touches a device.
Every module under here imports torch. Names resolve lazily so that importing the package, the config tree or the CLI's read-only commands costs nothing and works in an environment where torch is not installed; asking for one of these names without torch raises that import's own error, which already says which package is missing.
Rating and checkpoints¶
royalelearn.ladder
¶
The ladder: the rating, who plays whom, what a gate decides, and where the evidence lives.
Names resolve lazily, so that reading a result log -- which is what the rating tests and any offline analysis do -- does not import an environment, a snapshot store or torch.
royalelearn.checkpoint
¶
Checkpoints: per-component folders, a manifest that proves the bytes, and an atomic write.
Every component owns its folder and its own pair of methods, so adding one to a checkpoint is adding a folder name and a dict entry rather than editing a central serialiser. Nothing about a component's state is described in two places.
Atomicity. Everything is written into <name>.partial/, every file is fsynced, and then
the directory is renamed into place. The directory itself is not fsynced: os.fsync on a
directory handle is a POSIX guarantee and raises on Windows, where this harness's default
profile runs, so the durability step is per file and the atomic step is the rename -- which is
atomic on both platforms and is the property recovery actually needs. A crash mid-write leaves
the previous checkpoint intact and the partial directory obviously named.
No pickle on any path this module writes. safetensors for weights, msgspec JSON for
everything else, and torch.load(weights_only=True) wherever a component reads an optimizer
state back. A checkpoint is loaded six weeks later from a directory nobody has looked at since;
it must not be able to run code.
IndexEntry
¶
Bases: Struct
One checkpoint, as the index knows it.
CheckpointIndex
¶
Bases: Struct
The run's checkpoints, newest last.
latest() reads this rather than parsing directory names. int(x) for x in
os.listdir(...) crashes the save on any stray file, and the save is called from the crash
handler, which is exactly where the user most needs it not to.
RngComponent
¶
Every random stream's position, as one folder of a checkpoint.
One reference learner saves none of this; the other re-seeds all three generators from the run's initial seed on load, so a run resumed at ten million steps draws the same action noise it drew at step zero. Because every stream in this harness is name-addressed, restoring the iteration counter and the shard positions restores the stream itself rather than merely the parameters -- the torch and python states below are for the few draws that are not name-addressed, such as a dropout mask.
capture()
¶
Everything that would have to be true again for the next draw to be the same one.
DirCheckpointStore
¶
Bases: CheckpointStore
One directory per checkpoint, named by the cumulative env step that produced it.
verify(path, manifest)
¶
Every file in the manifest, hashed and compared. Names the first that fails.
This is what turns a truncated write from a crash six weeks ago into an error at the moment of the resume rather than into a run that continues from something else.
prune(run_dir=None, keep=None)
¶
Keep the newest keep and remove the rest, along with any abandoned partial.
HOUSEKEEPING MUST NOT BE ABLE TO END A RUN. shutil.rmtree raises PermissionError
on Windows while any file inside the folder is open, and reading a checkpoint while a run
continues is an ordinary thing to do. This is called at every checkpoint -- about 1,990
times in a long run -- and before 2026-09-23 a single open handle anywhere in the oldest
folder would have ended it.
A folder that could not be removed STAYS IN THE INDEX, so the next prune tries it again rather than losing track of it, and it is not reported as removed. The returned list is what actually went.
Found by a platform-portability review, which measured it here rather than arguing it.
sha256_of(path, chunk=1 << 20)
¶
The hash the manifest records, read in chunks so a buffer file costs no memory.
config_differences(before, after)
¶
Every leaf in which two configs differ, keyed by its dotted path.
Over the whole tree rather than the top level, so that a changed learning rate reads as
ppo.lr_actor instead of as the whole ppo block having moved.
check_resume(manifest, identity, config, *, allow_drift=False)
¶
What a resume has to agree about, checked once before anything is loaded.
A difference in an identity field is a refusal rather than a warning: a policy trained on one catalogue's observation vector cannot load into another's, and the failure that follows is a shape error deep inside a forward pass rather than a sentence naming the field. Every other difference is printed and continued past, because a run that resumes with a different checkpoint interval is the same run.
Returns the config differences it printed. allow_drift records the identity differences
instead of refusing them, which is what --allow-identity-drift is for.
check_described(manifest, rows)
¶
Refuse a checkpoint whose learner no metric row of its iteration describes.
A checkpoint's weights are supposed to be ones the run's record reports: the row of the iteration its manifest names carries the same state digest. One that breaks this holds a learner trained past its row -- which is what the emergency save wrote while the update ran before the batch was judged -- and resuming it continues the run from a state its record does not contain, under counters that belong to a different learner.
Any row of that iteration will do: a run resumed from an earlier checkpoint writes the iterations after it a second time, and each copy describes a learner that existed. What cannot be compared is let through rather than refused -- no metric file, no row of that iteration, a line torn by a crash mid-write -- because an absence says nothing about the weights. Past iteration zero, which never has a row, it is printed: a resume the record could not vouch for should not read like one it did.
Seeding and determinism¶
Every random draw in a run descends from one seed tree, so a resumed run lands on the original's curve.
royalelearn.seeding
¶
Name-addressed random streams.
Every generator in the harness is derived from the run's master seed and a PATH -- a string
naming the consumer, such as "ppo/minibatch/iteration/4/epoch/1" -- rather than by spawning
children of a root SeedSequence in the order the code happens to ask for them. Positional
spawning makes every stream downstream of a new consumer move, so adding a diagnostic that draws
one number changes the actions a policy takes. Here a stream is a pure function of its name, so a
config change that ought to be irrelevant is irrelevant, and two runs of the same identity draw
the same numbers however their code paths are ordered.
The namespace below is the whole of it. A new consumer adds a row rather than reusing a neighbour's path, because two consumers on one path advance each other's stream.
Stream
¶
Bases: NamedTuple
One row of the namespace: a path template and what draws from it.
derive_seedseq(master_seed, path)
¶
The SeedSequence for one named stream of one run.
The path is hashed to a 128-bit spawn key rather than appended to the entropy, so that the master seed stays the run's single number and the path cannot collide with a seed value. blake2b because it is in the standard library, is fast on short strings, and -- the property that matters -- is fixed forever, which a string hash with per-process randomisation is not.
derive_generator(master_seed, path)
¶
A PCG64 generator for one named stream. PCG64 is numpy's default and is stable across versions and platforms, which is what makes a recorded run reproducible on another machine.
derive_int(master_seed, path)
¶
A reproducible 63-bit integer, for env seeds.
63 rather than 64 bits because the seed crosses into ClashSelfPlayVecEnv.reset(seed=...)
and gymnasium's seeding refuses a value that does not fit a signed 64-bit integer.
stream_path(template, /, **fields)
¶
Fill one row of STREAMS.
Going through here rather than writing an f-string at the call site means a typo in a path is
a KeyError naming the template at the moment it is drawn, instead of a private stream that
quietly works and silently fails to be the stream anyone meant.
royalelearn.determinism
¶
The determinism tiers, and the process settings each one needs.
Three tiers, of which the first is unconditional (docs/harness-spec.md section 5.1):
T1 env-exact given the run identity, every episode's engine state-hash sequence, every
observation, every mask and every reward are bit-identical, on any machine,
forever. It costs nothing and is a property of the engine and of name-addressed
seeding, so there is no switch for it and apply does not mention it.
T2 run-exact T1, and every gradient, parameter and metric row bit-identical on the same device
class and torch/CUDA build. The default, at 10-20% of throughput.
T3 throughput T1 only; the learner may pick nondeterministic kernels.
Two of the settings T2 needs cannot be applied from here, because they must be in place BEFORE
torch initialises CUDA and before numpy imports its BLAS: CUBLAS_WORKSPACE_CONFIG and the
three BLAS thread counts. They are environment variables, the entry point sets them, and
apply asserts rather than sets, so a run that would have been silently non-reproducible dies
at start-up naming the entry point instead of producing a curve that cannot be repeated.
This module imports torch inside the functions that need it. Importing it at module scope would
put a 300 MB, three-second dependency on the import path of royalelearn.cli, which has to
work in an environment without torch at all.
apply_cublas_workspace_config(env=None)
¶
Set CUBLAS_WORKSPACE_CONFIG if it is unset, and return its value.
Called by the entry point before torch is imported. An existing value is left alone even when
it is not one this harness would have chosen: the operator's own setting wins, and
require_cublas_workspace_config is what decides whether it is good enough.
apply_blas_thread_env(env=None)
¶
Pin the BLAS thread counts to one. Must run before numpy is imported.
require_cublas_workspace_config(entry_point='royalelearn.cli', env=None)
¶
Refuse to run T2 without a deterministic cuBLAS workspace.
cuBLAS reads this variable once, when CUDA initialises, so setting it here would be too late and setting it silently would be worse than not setting it at all: the run would look reproducible and would not be.
set_fill_uninitialized(on)
¶
Turn torch's fill of unwritten memory on or off (docs/harness-spec.md section 5.1).
It has an effect only while deterministic algorithms are on. Read at every allocation, so a change takes effect at the next one; nothing already allocated is touched.
apply(tier, *, torch_threads=1, entry_point='royalelearn.cli')
¶
Put the process into tier and return what was applied, for the record.
T2 keeps bf16 autocast. Deterministic kernels are bit-reproducible run to run at any precision, so the tier costs nothing in precision: it buys exactly what it says, which is two runs of one identity agreeing row for row. tf32 is off because it is not bit-reproducible across the different batch shapes the rollout forward and the update forward use.
Metrics¶
royalelearn.metrics
¶
Where a run's numbers go: the schema they are checked against, the sinks that write them, and the alarms that read them.
Names resolve lazily, so that reading the schema -- which is what the documentation build and the metric tests do -- does not import a sink, a socket or wandb.
The base classes¶
royalelearn.api
¶
Every ABC and struct the harness is written against.
This subpackage imports numpy and msgspec and nothing else that matters: torch appears in type
annotations only, behind TYPE_CHECKING. That is what lets a rollout worker, the CLI's
config and identity commands, and import royalelearn itself run in an environment with no
torch installed -- and what keeps a worker's resident memory three hundred megabytes smaller
than the parent's.
The concrete implementations live outside api/ and may import whatever they need:
RolloutSource rollout/farm.py, rollout/inline.py
ActorCritic learn/actor_critic.py
ObsCodec rollout/codec.py
ExperienceBuffer learn/buffer.py
AdvantageEstimator learn/gae.py
Update learn/ppo.py
Schedule learn/schedules.py
Matchmaker, Rater ladder/matchmaker.py, ladder/rating.py
MetricsSink, Alarm metrics/sinks.py, metrics/alarms.py
CheckpointStore checkpoint.py
AdvantageEstimator
¶
Bases: ABC, Checkpointable
Rewards and values in, advantages and returns out.
compute(*, rewards, values, final_values, terminated, truncated, trainable, gamma, lam)
abstractmethod
¶
rewards/terminated/truncated/trainable are (T, R); values is
(T+1, R); final_values is (T, R) and is read only where truncated.
Returns (advantages (T, R), returns (T, R), stats).
Implementers MUST bootstrap a terminated cell from 0 and a truncated cell from
final_values, and MUST NOT carry the recursion across an episode boundary. A
truncation is an episode that was cut, not one that was decided, and treating the two
alike throws away the value of every position a step limit ended.
AdvantageStats
¶
Bases: Struct
What the estimator saw, for the metric row.
reward_scale is the divisor the return scaler applied and clipped_reward_frac is how
much of the batch the clip bound touched: both are how a scaled reward stays legible after
the scaling.
CodecTable
¶
Bases: Struct
How each observation key is stored, decided from EnvSpec.obs_space at preflight
rather than from a list of plane indices. Logged, hashed and written into every snapshot.
plane is one entry per spatial plane: its name, its storage in
{"uint8", "float16", "static", "derived"}, and the divisor that takes the stored integer
back to the value the environment produced. "derived" is reserved: nothing rebuilds such a
plane yet, so the codec refuses it at bind. Two runs whose tables differ are not comparable,
and digest() is what says so.
digest()
¶
sha256 of the canonical JSON of this table.
ExperienceBuffer
¶
Bases: ABC, Checkpointable
A rectangle of (T + frame_stack) cycles x R slots: T collected cycles, one
bootstrap row, and frame_stack - 1 history rows carried over from the previous
iteration. Owns the shared-memory block the workers write into.
Implementers may assume each (cycle, slot) cell is written exactly once, by the worker
that owns that slot; the learner only reads.
shared_handle()
abstractmethod
¶
Name, size and the numbers the offsets follow from; picklable, and sent to workers.
record_round(r, actions, log_probs)
abstractmethod
¶
Scalars only: observations are already in place. O(n), no observation copy.
set_values(values)
abstractmethod
¶
(T+1, R) float32, from the whole-iteration critic pass.
set_n_legal(counts)
abstractmethod
¶
(T+1, R) integer: how many actions each cell's mask left, from the critic's pass.
Implementers keep the collected cycles and may drop the bootstrap row, in which no action was taken. The count is what tells the update which rows had a choice at all without unpacking an observation to find out.
set_final_values(cells, values)
abstractmethod
¶
V(final_obs) for truncated cells; cells is int64[(k, 2)] of (cycle, slot).
set_advantages(adv, ret)
abstractmethod
¶
(T, R) float32 each.
trainable_mask()
abstractmethod
¶
(T, R) bool: which cells reach the update. What decides it is the seat's group
and the cell's validity, never where the row was written.
batches(batch_size, minibatch_size, epochs, rng_for_epoch, *, choice_first=False)
abstractmethod
¶
Yield batches; each Batch knows its true sample count and iterates device-resident
minibatches. A batch never straddles an epoch boundary, and an epoch holds as many whole
batches as it can fill with the rows over spread one each across them, so every batch is
at least batch_size and an epoch is exactly n // batch_size optimizer steps. An
epoch with nothing trainable in it yields no batches at all. Each minibatch is weighted by
its share of its own batch. Gathers per MINIBATCH, never per batch.
choice_first reorders each batch's cells so that the ones with more than one legal
action come first, keeping the permutation's order inside each class. Implementers MUST
move no cell between batches: it is a reordering, so that a caller skipping the forced
rows skips whole minibatches of them, and every batch-level denominator is unchanged.
ObsCodec
¶
Bases: ABC
Quantisation of one observation row. The worker packs; the learner unpacks on the GPU.
Implementers MUST be exact round-trips for the integer-valued channels and MUST declare
row_bytes as a constant given an EnvSpec and its CodecTable.
codec_version
abstractmethod
property
¶
The RULE's version. The table it produces is data and travels separately.
table(spec, sample, *, min_states=MIN_TABLE_STATES)
abstractmethod
¶
Decide storage per key from the declared bounds and a sample of real observations.
Storage is decided from a sample; EXISTENCE is decided from the declaration. A plane the layout does not declare static is stored even when it is constant across the sample, because the tower planes are constant in any sample in which no tower falls.
Implementers MUST refuse a sample of fewer than min_states observations. The row
size of the whole run follows from this one decision, and a sample too small or too
idle to have reached the states a plane varies in decides it wrongly and in silence.
static_planes(obs)
abstractmethod
¶
The planes EnvSpec.spatial_layout declares static; stored once per seat, never
per row.
unpack_to_device(raw, statics, out)
abstractmethod
¶
Dequantise, scatter the static planes in, reshape the stored mask into the mask planes, and gather the frame-stack history.
Checkpointable
¶
Bases: Protocol
Anything that goes into a checkpoint folder of its own.
FORMAT_VERSION is the component's own, not the checkpoint's: a component may change its
files without the checkpoint format moving, and the manifest records each one so that a
load can say which component it could not read.
load_checkpoint(folder, *, strict)
¶
With strict=False, a missing file prints the exact path it wanted and continues
with a default. With strict=True it raises. strict defaults to True on resume:
tolerate-everything is right for a research tool and wrong for a harness that promises
the curve continues.
CheckpointStore
¶
Bases: ABC
Where checkpoints live, and the only thing that writes or reads them.
write(components, manifest)
abstractmethod
¶
Write atomically: into <name>.partial/, fsync each file, then os.replace.
The directory itself is not fsynced. os.fsync on a directory handle is a POSIX
guarantee and raises on Windows, where this harness's default profile runs, so the
durability step is per file and the atomic step is the rename -- which is atomic on both
platforms, and is the property recovery actually needs.
read(path, components, *, strict)
abstractmethod
¶
Verify every hash in the manifest, then load each component. A mismatch raises
CheckpointFormatError naming the first file that failed.
latest(run_dir)
abstractmethod
¶
The newest checkpoint, from the run index rather than from parsing directory names:
int(x) for x in os.listdir(...) crashes on a stray file, and the save is called from
the crash handler, which is where the user most needs it not to.
prune(run_dir, keep)
abstractmethod
¶
Remove all but the newest keep, and return what was removed.
Manifest
¶
Bases: Struct
What one checkpoint is, beside the folders that hold it.
config is the resolved config verbatim, so a checkpoint is self-describing without the
run directory around it, and files is every file in the checkpoint with its sha256, so a
truncated write from a crash six weeks ago is an error on load rather than a silently wrong
resume.
RngState
¶
Bases: Struct
Every random stream's position at the moment a checkpoint was written.
One reference learner saves none of this; the other re-seeds from the run's INITIAL seed on load, so a run resumed at ten million steps draws the same action noise it drew at step zero. Because every stream here is name-addressed, restoring the iteration counter and the shard positions restores the stream, not merely the parameters.
ConditionResult
¶
Bases: Struct
One gate condition, with the numbers that decided it rather than a bare pass or fail.
EvictionPolicy
¶
Bases: ABC
Which snapshots the sampler stops drawing.
select_for_eviction(*, pool, ratings, max_sampled)
abstractmethod
¶
Ids to remove FROM THE SAMPLER. The archive and the result log are never touched: eviction is about sampling cost, not about forgetting evidence.
GateDecision
¶
Bases: Struct
A candidate's audition, kept whole.
admit puts the snapshot in the pool, promote makes it the champion, and cycle
is the case where a snapshot is worth playing against without being the best: the three are
separate because a pool that only ever admits champions forgets everything it beat.
Matchmaker
¶
Bases: ABC, Checkpointable
What every battle plays next.
plan(iteration, pool, ratings, geometry)
abstractmethod
¶
The iteration's opening table. One Assignment per battle, each drawn at that
battle's current ordinal, so plan() is assign() applied across the geometry.
assign(battle, ordinal, pool, ratings)
abstractmethod
¶
The assignment for one battle's next episode.
Draws from match/battle/{battle}/ordinal/{ordinal}, so it is a pure function of the
master seed and its two arguments: the same episode of the same battle always meets the
same opponent, whichever iteration it happens to fall in and whichever worker holds it.
on_episode(record)
abstractmethod
¶
Record a finished episode, for the mixture's own bookkeeping.
PromotionGate
¶
Bases: ABC
Whether a candidate snapshot joins the pool, and whether it becomes the champion.
Rater
¶
Bases: ABC, Checkpointable
How a pile of results becomes a number per player.
fit(results)
abstractmethod
¶
id -> (rating in Elo units, standard error).
MUST be a pure function of results: same games in, same numbers out, in any order,
on any machine.
predict(a, b)
abstractmethod
¶
P(a scores against b), with draws counted as half a win.
transitivity_residual(results)
abstractmethod
¶
How badly one number per player fits: near zero on transitive results, large on synthetic rock-paper-scissors.
RatingTable
¶
Bases: Struct
One fit of the whole result log.
Ratings are in Elo units with the anchor pinned exactly, standard errors come from the
inverse observed Fisher information, and transitivity_residual says how much of the
result log a single scalar per player fails to explain -- which is the number that decides
whether a scalar rating is lying.
SnapshotStore
¶
Bases: ABC
Where frozen actors live. They outlive checkpoints: a rating means nothing without the player it rated.
get(snapshot_id, device)
abstractmethod
¶
LRU-cached: a shard-round touches at most ladder.max_resident_opponents of them,
so a cache that never thrashes is a small one.
Alarm
¶
Bases: ABC
A predicate over a metric row, with a severity and a patience.
severity="warn" logs and writes an alarms.jsonl row; severity="halt"
additionally writes a checkpoint and a diagnostic bundle, then raises AlarmHalt. An
alarm can stop a run and can never alter a value, which is why its thresholds are recorded
but excluded from the run identity.
AlarmResult
¶
Bases: Struct
One alarm's verdict for one iteration.
consecutive is how many iterations in a row the predicate has held, so a row written
before the patience is spent still records that something was building.
MetricsSink
¶
Bases: ABC, Checkpointable
Somewhere a row goes.
write(row)
abstractmethod
¶
One flat dict per iteration. Implementers MUST NOT mutate the row and MUST NOT raise on an unknown key.
write_episodes(rows)
¶
Finished episodes. Default: ignore.
write_alarms(alarms)
¶
Alarms that fired this iteration. Default: ignore.
write_artifact(name, path)
¶
A file produced this iteration -- a heatmap, a bundle. Default: ignore.
ActionDistribution
¶
Bases: ABC
A categorical distribution over the action space, with the illegal actions removed.
sample(uniforms)
abstractmethod
¶
(B,) int64. Driven by CALLER-SUPPLIED uniforms in [0, 1) so that the sampled
action is reproducible independently of batch composition and of torch's global RNG.
mode()
abstractmethod
¶
(B,) int64, mask-respecting argmax.
log_prob(actions)
abstractmethod
¶
(B,) float32.
entropy()
abstractmethod
¶
(B,) float32, over the whole legal set.
noop_entropy()
abstractmethod
¶
(B,) float32: the binary entropy of p(no-op) against p(play).
The leading indicator of no-op collapse, and bounded by 0.693 nats, which is why the coefficient that acts on it is not the joint entropy's.
n_legal()
abstractmethod
¶
(B,) int64. Diagnostics: raw entropy falling is ambiguous without it.
Actor
¶
Bases: ABC
The policy network. Implementations are also torch.nn.Module.
logits(obs)
abstractmethod
¶
(B, spec.n_actions) float32 ALWAYS, even inside autocast. Raw logits: no softmax,
no clamp, no mask.
distribution(obs)
¶
The default masked categorical over this actor's logits.
The import is here rather than at module scope because the shipped distribution is a torch module and this file must import without torch.
ActorCritic
¶
Bases: ABC
The pair, as the learner uses them.
Implementers may assume obs tensors are on the device and already dequantised, and that
obs.mask[:, 0] is True on every row (RoyaleGym sets mask[NOOP]=1 unconditionally,
action.py, including after game over). They MUST assert it.
ActResult
¶
Bases: Struct
What one rollout forward produces. Everything after the first two is diagnostics.
BackpropResult
¶
Bases: Struct
What one update forward produces, recomputed under the current parameters.
Critic
¶
Bases: ABC
The value network. Implementations are also torch.nn.Module.
value(obs)
abstractmethod
¶
(B,) float32.
NetworkFactory
¶
Bases: ABC
How an ActorCritic is built, and how two of them are told apart.
arch_digest(spec, arch)
abstractmethod
¶
sha256 of the canonical JSON of (arch, obs_space, frame_stack, num_cards, n_actions).
A snapshot or checkpoint built from a different architecture is refused by this digest, not shape-errored halfway through a load.
ObsBatch
¶
Bases: NamedTuple
One batch of observations on the device, already dequantised.
Shapes follow from EnvSpec and the frame stack k: spatial is
(B, k*S, H, W), mask_planes is (B, k*A_s, H, W) for the A_s per-slot action
planes, vector is (B, V) and is the CURRENT frame's only, and mask is
(B, n_actions) bool.
Assignment
¶
Bases: Struct
What one battle's next episode is.
Drawn by the Matchmaker at the battle's own episode boundary, from
match/battle/{b}/ordinal/{k}, and therefore a pure function of the master seed, the
battle index and the reset ordinal.
Close
¶
Bases: WorkerCommand
Shut the shard down.
Defer
¶
Bases: WorkerCommand
Nothing this round; hand back the same data.
EnvSpec
¶
Bases: Struct
Everything the learner must know about the environment before it builds anything.
Every field here is READ from the running environment at preflight. The harness holds no layout constant of its own: a width, a plane count and a field offset are all environment facts, and typing one into this repo would make a RoyaleGym change a silent wrong answer instead of a loud one.
n_planes
property
¶
Spatial planes the environment emits, static ones included.
n_grid_actions
property
¶
The no-op plus one action per (hand slot, tile): the part of the action space the pointer head lays out as a grid, and all of it for an environment without buttons.
n_buttons
property
¶
Ability buttons (a hero's, a champion's): the actions after the grid, action
n_grid_actions + k pressing button k. Their readiness is the observation's
ability_ready, which is the same bits as the mask's last n_buttons; 0 for an
observation without it.
static_planes
property
¶
The planes the layout DECLARES static -- never the ones a sample found constant.
EpisodeRecord
¶
Bases: Struct
One seat's finished episode: what it was, how it went, and enough to replay it.
The terminal scalars are copied straight out of the environment's final_info rather
than reconstructed from the states the worker saw. The environment already computes them,
and a second implementation of the same summary is a second answer. RoyaleGym's
EPISODE_STAT_KEYS is the authority on the set.
ObsKeySpec
¶
Bases: Struct
One key of the environment's observation space, read from the space itself.
low and high are per channel for a three-dimensional key and length one otherwise.
A space that declares scalar bounds broadcasts them over the whole array, so the per-channel
form is a reduction of the declared bounds over each plane rather than a second opinion about
them: it is what the codec needs in order to decide a plane's storage (section 7.2).
Plan
¶
RolloutRound
dataclass
¶
One shard-round of observations. Holds zero-copy views; not a Struct.
The observations themselves are not here: the worker packed them straight into their final
resting place in the experience buffer, and obs_rows says where each slot's row went.
What crosses the boundary is the scalars and the episodes that ended.
TWO TIMESTEPS, AND WHICH IS WHICH. A round is published after a step, so it carries the state that step reached and the step itself, and those are not the same timestep.
obs_rows,tickandgroupare the round's OWN cycle: where this state was written, the engine clock it stands at, and who holds the seat from here on.reward,terminated,truncated,deploy_statusandepisode_endare the step that ARRIVED here, which began one cycle earlier -- so they belong to the row below this one in the rectangle, which is the row whose action produced them.validis the publication: false says this round did not arrive, so neither the state nor the step it reports exists.
The round at cycle 0 has no step behind it and carries zeros for that half. The round at
cycle T has no row of its own and is the only carrier of row T - 1's step.
deploy_status is -1 for no command and 0 for an accepted one; 1..11 is a DeployStatus
refusal, which under a correct mask cannot happen and is therefore a mask bug rather than a
tolerance.
RolloutSource
¶
Bases: ABC
Where experience comes from. The one seam between the learner and the world.
Implementers may assume: begin_iteration precedes any next_round; exactly one
submit per next_round; Step.actions are legal under the mask that was handed out;
the caller does not retain a RolloutRound's views past the next next_round for the same
shard.
Implementers MUST guarantee: slots is ascending; every slot appears exactly once per
cycle; a dead worker's slots arrive with valid=False rather than not arriving.
stats()
¶
Per-round timing: env_ms, wait_ms, codec_ms, parent_wait_frac, bytes_out.
parent_wait_frac is the PARENT's share: time it spent blocked on a round that had
not been published, over that plus the workers' own env time. It rises when the workers
cannot keep up, which is the opposite reading from a number about workers idling, and
the opposite action. An inline source reports zero because there is nobody to wait for.
Default: {}.
SetState
¶
Bases: WorkerCommand
Start each battle from a recorded position; one blob per battle, None to leave it.
SlotPlan
¶
Bases: Struct
The iteration's opening assignment table: what every battle is playing at the moment the iteration begins.
Assignments change only at a battle's own episode boundary, where the parent draws a fresh one; nothing in an iteration's span reassigns a battle mid-episode, so a partially controlled trajectory cannot occur.
Spaces
¶
Bases: WorkerCommand
Re-read the spaces without stepping.
Step
¶
Bases: WorkerCommand
Step this shard by one decision.
group, opponent_ix and learner_seat carry the parent's assignment for every
battle whose episode started on the previous round, and are empty otherwise; the worker needs
them only to know which slots it fills with a scripted action.
WorkerCommand
¶
Bases: Struct
What the parent hands a shard for one round. Six variants, one per round.
A tagged union rather than a dict of flags: the worker's loop switches on the tag, so a command it does not know is a protocol failure at the boundary instead of a missing key deep inside a step.
WorkerFailure
¶
Bases: Struct
A worker that stopped answering, as a value the coordinator can act on.
kind is one of "exception", "crash", "timeout" or "protocol", and message is the
child's traceback verbatim. Both reference learners treat a dead worker as a permanent silent
hang; here it is a typed failure, which is what lets the farm restart the worker and the run
continue with the loss recorded rather than hidden in a throughput dip.
Schedule
¶
Bases: ABC
One scheduled scalar.
ScheduleState
¶
Bases: Struct
Every scheduled quantity, evaluated once at the top of an iteration.
Evaluated once and passed down, rather than evaluated where each is used: the discount the reward was computed with, the discount GAE used and the discount the metric row reports are then the same number by construction.
credit_horizon_seconds(decision_ms)
¶
1 / (1 - gamma * lambda) decisions, in seconds.
decision_ms comes from the environment, so the horizon follows a change in the
decision granularity instead of quietly meaning something else. Logged every iteration,
because it is the single highest-leverage number in the config and because it can
otherwise be ten seconds by accident.
ActorLossTerm
¶
Bases: Protocol
A term an extension adds to the actor's loss.
loss returns (coefficient, raw). The update adds coefficient * raw * scale, where
scale is actor_scale for a term whose raw value is a sum over the choice rows
(scaling == "rows", which keeps minibatch size a pure memory knob and matches the policy
term under every ppo.forced_rows value) and the minibatch weight for one whose raw value
is a mean of its own (scaling == "minibatch"). The gradient ratio, when asked for, is
measured on raw * scale, before the coefficient.
finish returns the term's keys for the row; every one must start with <extension>/.
actor_trained is False on a frozen iteration, when loss was not called at all.
state is what a checkpoint keeps (an empty dict keeps nothing); format_version is
saved beside it and a different one refuses the resume.
ActorTermInputs
¶
Bases: NamedTuple
What an extra actor-loss term is handed, identically in both of the update's actor paths.
log_probs and mask are the actor's masked log-probabilities and mask on the
minibatch's choice rows, in rows order, with the graph attached. actor is the live
actor, for a term that needs a forward of its own on other rows. actor_scale is the
per-batch scale the policy term is multiplied by, and weight the minibatch's share of
its batch.
cells names each row of the minibatch by its position in the iteration's whole batch, in
the minibatch's row order (cells[rows] names the choice rows). Every epoch trains every
cell once, so a term whose input on a row does not change within an iteration -- a frozen
reference's forward -- can compute it in the first epoch and look it up by cell after. None
from a caller that does not have them.
Update
¶
Bases: ABC, Checkpointable
The learner's optimisation step. Owns the optimizers, so it is checkpointable.
step(buffer, sched)
abstractmethod
¶
Consume one iteration's rectangle and return what happened.
Implementers may assume the buffer's advantages and returns are already computed and that every cell it hands out is trainable and valid. They MUST apply exactly the mask read from the buffer, never a recomputed one, and MUST leave the buffer unchanged.
UpdateResult
¶
Bases: Struct
What one update did, in the units the metric row uses.
kl_by_epoch and clip_fraction_by_epoch are per epoch rather than averaged: the rule
for n_epochs is read off them ("if epoch three's clip fraction is more than twice epoch
one's, lower it"), and an average cannot answer that. samples_unused_frac reads zero by
construction and is logged so that it can be seen to.