Skip to content

royalelearn

The training harness.

New, and barely tested

The training loop closed on 2026-09-22. python -m royalelearn train runs end to end, and it needs the torch extra. Every run so far has been a short test, ten iterations at most as of 2026-09-22, so treat anything below that touches the loop itself as young code. See The learner for what that means in practice.

Everything on this page is generated from the docstrings in the code, so it shows what exists rather than what is planned. It covers the main modules, not every one. LearningCoordinator, the object that runs a training run, lives in royalelearn.coordinator, which is not on this page. Read its docstrings in the code.

Some of these need torch, which is an optional extra: pip install -e "RoyaleLearn[torch]". The configuration tree, the run identity and the rollout workers are deliberately torch-free, and that is a property the package is tested for rather than an accident.

Configuration and identity

You write a config, and the run's identity is computed from it so that a resume can be checked against the run it is resuming.

royalelearn.config

The whole configuration of a run, as one tree.

There is one source of truth. The config object is it: no keyword-argument facade duplicates its leaves, because two sources of truth for one number is the kind of variable this harness exists to remove. Every struct forbids unknown fields, so a typo is an error at load rather than a setting that silently does nothing, and every number carries its unit in docs/harness-spec.md section 6 beside the reason it has the value it has.

msgspec rather than pydantic because it is already a RoyaleGym dependency, it gives forbidden unknown fields, tagged unions and a canonical JSON round-trip, and the package's dependency list stays royalegym, numpy, msgspec.

Profiles are functions rather than files: one per machine class, differing only in numbers, and a JSON config names one and overrides what it likes on top.

ConstantSpec

Bases: Struct

A value that does not move.

LinearSpec

Bases: Struct

start to end over over_env_steps game-steps, then end forever.

GeometricSpec

Bases: Struct

start to end geometrically, which is what a discount wants: 1 - gamma decays by a constant factor rather than gamma moving by a constant amount.

PiecewiseConstantSpec

Bases: Struct

(env_steps, value) breakpoints, held between them.

ObsConfig

Bases: Struct

What the policy sees beyond the current frame.

The observation carries positions and no motion, so a policy that needs to tell an advancing unit from a retreating one needs more than one frame. Stacking is a gather of buffer rows and costs no extra storage; it is 1 by default because the extra channels cost stem compute and the first run measures whether they buy anything.

RolloutConfig

Bases: Struct

The shape and the failure behaviour of the worker farm.

NetConfig

Bases: Struct

The architecture. Hashed into arch_digest, so a snapshot built from another one is refused rather than shape-errored.

LrBackoffConfig

Bases: Struct

Drop the learning rate when the KL stays above a threshold.

PPOConfig

Bases: Struct

The update.

AdvantageConfig

Bases: Struct

The discount and the credit horizon.

0.99 discounts the win to almost nothing by the end of regulation; the anneal follows OpenAI Five's, with an endpoint scaled to a four-minute match. gae_lambda at 0.99 gives a credit horizon of about 45 seconds, which matches the deploy-push-tower causal chain -- the references' 0.95 gives ten and cannot connect a deploy to the tower it takes.

GateConfig

Bases: Struct

The three conditions a candidate must pass to become the champion.

RaterConfig

Bases: Struct

The Bradley-Terry-Davidson fit.

SeedSnapshot

Bases: Struct

A frozen actor the pool holds from the run's first iteration and never evicts.

path is a folder holding actor.safetensors and spec.json: a run's own snapshot (<run>/snapshots/<digest>) or an actor artifact RoyaleImitate wrote. A relative path is read from the folder the command runs in. sha256 is of actor.safetensors, and it is what the run's identity records, so the same path holding other weights is another run; a run whose folder does not match is refused, and the refusal prints the digest it found.

SeatDecks

Bases: Struct

One deck the learner trains, dealt where the learner sits, against a field of others.

A mirror battle deals deck to both seats. A pool or scripted battle deals deck to the learner's seat, which is fixed per battle (even battles blue, odd red), and a deck from field to the other seat, where a frozen or seeded snapshot or a scripted opponent plays and whose rows are not trained on. field is drawn uniformly (list a deck twice to weight it), or None for a random deck of eight different cards. Decks are card NAMES, looked up in the engine's catalogue at the first reset; shuffle is RoyaleGym's ShuffleMode for a non-mirror deal. See royalelearn/ladder/seat_decks.py.

cls is the curriculum class each battle's deal is: RoyaleGym's by default, or a subclass that takes its keywords unchanged (one that also deals forms to the seats holding deck, say), resolved like any component, so its module must be in config.extra_component_modules. kwargs are that class's own keywords, merged into the ones the deal decides; they may not name those (deck, p, mirror_p, seat, pool, shuffle).

LadderConfig

Bases: Struct

Who the learner plays, and when a snapshot joins the pool.

SinkSpec

Bases: Struct

One metrics sink. JsonlSink is always installed: it is the file the resume test compares.

AlarmConfig

Bases: Struct

Thresholds for the alarm table of section 13.3.

Recorded and excluded from the run identity: an alarm can stop a run and can never alter a value, so two runs that differ only here produce the same numbers for as long as both run. The severities and patiences live with the alarms themselves; the overrides below are for an operator who wants one of them louder or quieter on their own machine.

DeterminismConfig

Bases: Struct

Section 5.1. The tier is part of the run identity; the thread count is part of the tier.

DoctorConfig

Bases: Struct

The start-up gates that can refuse a run.

RunConfig

Bases: Struct

One run, completely: RoyaleLearn's own keys.

A run that uses an extension (royalelearn.extensions) has one more top-level key per extension, its section. load_config returns such a config as a subclass of this one with the sections as fields after these, so a config without sections is exactly this type and encodes as it always did.

Geometry

Bases: Struct

The iteration's rectangle, derived from the config once and passed around whole.

cycles by n_slots, with slot -> (worker, shard, battle, seat) fixed for the run. Everything downstream -- the buffer's shape, the slot maps, the matchmaker's plan -- reads these numbers rather than recomputing them, so the arithmetic of section 2.2 lives in one place (config.geometry).

default_env_spec(engine=RUST_ENGINE, *, max_steps=480)

The environment a run uses unless it says otherwise.

The truncation is a step limit rather than a tick limit because it is decisions that the buffer counts; 480 of them is a full regulation match plus overtime at the default decision granularity, and it is the cap the episode-length histogram spikes at when two policies settle into the turtle equilibrium.

The reward is this package's potential composition rather than RoyaleGym's default_reward, whose elixir-trade term pays a player who never commits a card and whose tower term is one discount short of a potential (rewards.py says why each matters). The reward is part of the environment spec and so of the context digest: runs trained against the other one are a different objective and do not share a context with these.

role_counts(mix, n_battles)

How many battle slots play mirror, how many pool and how many scripted.

The mixture is a count of slots, not a probability each slot is drawn from, so this is the one place it is turned into whole battles. MixMatchmaker decides which slots get each role and geometry sizes the iteration from the same counts: two callers, one function, because the whole point of the exercise is that the number the iteration is sized for is the number the matchmaker fills.

It lives here rather than on the matchmaker only because of the import graph -- the matchmaker reaches config through api.ladder and identity, so config importing the matchmaker would close a cycle. The dependency runs the way it already did.

Rounding is resolved in two stages, each to the nearest whole battle with a half rounding up, and the last share takes whatever is left:

  • mirror = round(mix[0] * n_battles),
  • the remaining battles split between pool and scripted in the ratio mix[1] : mix[2], pool rounded and scripted taking the remainder.

So the three always sum to n_battles, and each lands within one battle of its exact share (the proof is short: each stage rounds by at most a half, and the second stage's input is already off by at most a half). Halves round up rather than to even, because round(0.5) is 0 in Python and a single battle under the shipped mixture would then be a rectangle with no mirror in it at all.

Two consequences are worth saying out loud. A share small enough that its quota is under a half rounds to zero: three battles cannot hold a tenth of one, so that role is simply not in the rectangle, and the mixture the run reports is the count, not the config. And an all-mirror mixture keeps every slot, which is what makes (1, 0, 0) collect twice the rows per cycle rather than one and a half.

geometry(config)

The iteration's rectangle, derived from the config.

Learner rows are the slots whose seat the learner plays: both seats of a mirror battle and one of every other. Because role_counts makes the mirror battles a count rather than a share, that is n_battles + mirror_battles exactly -- the same number on every cycle of every iteration of every run, whatever the master seed.

It used to be int(n_slots * (mirror + (1 - mirror) / 2)), the count's expectation. The two agree at the shipped laptop profile (96 + 48 = int(192 * 0.75) = 144), and what the expectation could not do was promise it: the realised count was Binomial(96, 0.5) + 96, with a standard deviation of 4.9 rows per cycle against an iteration whose slack over the trainable-row floor was 0.2%, so about one master seed in four halted the run.

The cycle count is then what it takes to reach timesteps_per_iteration learner transitions, rounded up -- an iteration collects at least what was asked for, never less, and now the "at least" is a fact about the rectangle rather than about its average.

laptop()

8 GB RAM, 4 GB VRAM, 8 threads. The shipped default.

The learner is the bottleneck on this machine by about 2.5 times, which is why there are three workers and not thirty-two, and why the network is four hundred thousand parameters and not four million.

workstation()

32 GB RAM, 12 GB VRAM, 16 threads.

It asked for rollout.overlap until 2026-09-22, which bought a second rectangle and no overlap, because nothing in the package honours the flag. It is refused now.

many_core()

64 GB RAM, 24 GB VRAM, 64 threads.

profile(name)

The shipped config for one machine class.

check_consistency(config)

Everything wrong with a config, in one list.

All of it at once rather than one exception at a time: on a machine where building an environment costs a minute, finding four mistakes takes four minutes otherwise.

validate(config)

Return the config, or raise PreflightError naming everything wrong with it.

retired_dropped(document)

A stored config with what the one-release shim drops without a word dropped: the six retired alarm keys at their old defaults, a null imitation and its null init and actor_lr_scale. For comparing a checkpoint's config with this run's; nothing is refused or said here.

load_config(source, *, say=print)

Read a config: a JSON string, a path to one, or an already-decoded mapping.

The document is applied ON TOP of the profile it names, so a file that sets three numbers gets the rest of that machine class rather than the laptop's, and a file that sets none is exactly the profile. Unknown fields are refused at every level of the tree: a top-level key RoyaleLearn does not own is a section, and one that no installed extension provides is refused by name (extensions.providers_for). A section that is null counts as absent.

dump_config(config, *, indent=None)

The canonical JSON of a config: keys in a fixed order, so a reformat is not a new run. A section set to None is absent, as it is on load.

config_hash(config)

sha256 of the canonical JSON. Recorded, printed and diffed on resume.

royalelearn.identity

What a run IS, as sixteen hex characters.

The identity is everything that can change a number the run produces. Two runs with the same identity produce the same metric stream under determinism.tier = "run_exact"; two runs with different identities are different experiments and are never plotted as one. A resume compares the identity field by field and refuses a difference by name.

What is in it, and why:

master_seed, determinism_tier, device_kind, torch_version   each changes the byte stream
the engine build and the card catalogue                     the game itself
env_spec_digest, obs_digest, action_digest                  what the policy sees and can do
frame_stack                                                 the trunk's input width
arch_digest, codec_version, codec_table_digest              the network and how a row is stored
algo_digest                                                 the PPO, GAE and schedule settings
rollout_digest                                              workers, games, shards, T and R
ladder_digest                                               who the learner plays
extensions                                                  each extension section the run
                                                            uses, by content, and the code
                                                            that provides it

What is recorded and NOT in it, because none of it changes what the run is: run_name, runs_dir, timestep_limit, the metrics sinks and their settings, checkpoint.keep, the alarm thresholds and severities, the doctor's budget, and the knobs that set how the collection is scheduled rather than what it collects. An alarm can stop a run; it cannot alter a number. A resume diffs those and prints every difference rather than refusing.

EngineBuild

Bases: Struct

Which engine, built from which data and which code, with which cards.

Every field comes from ClashParallelEnv.config(): calibration_digest over the calibration values as loaded, build_digest over the copies compiled into the extension, and binary_sha256 over the extension file itself. The harness hashes no data file of its own -- a digest computed here from the data directory would be a second opinion about what the engine is running on, and a second opinion is exactly what a stale build looks like. stale_build_differences MUST be empty; it is recorded so that a failure is auditable rather than reconstructed.

binary_sha256 is the one that was missing, and the data digests could not stand in for it. build_digest hashes the DATA compiled in, not the Rust: across a rebuild from an edited state.rs on 2026-09-23 it read 8952c1c4aa7d7923 on both sides while the file's hash went to 04d42611db5d6e5a, so two runs either side of the rebuild wrote identical blocks and played different games. It is NOT_STATED for an engine with no compiled file, and NOT_RECORDED -- the default, which only a decode ever reaches -- for an identity written before this field existed.

ExtensionRecord

Bases: Struct

One extension section in a run's identity: what it says, and whose code reads it.

digest is the section's identity value (Extension.identity_value: files by content, alarms left out). The other three name the providing package -- distribution, __version__ and commit -- and are None for a section RoyaleLearn provides itself, whose code royalelearn_git already names.

RunIdentity

Bases: Struct

The identity itself. run_id is the first sixteen hex characters of its sha256.

Encoded with omit_defaults so that a field added with a default of None leaves every identity that does not use it byte-identical, and so every run_id where it was. Every other field is always set when an identity is computed; the only ones that can sit at their default are user_code, on an identity decoded from before it existed, and extensions, on a run that uses none.

Decoding refuses a field this code does not know. An identity written by a build with a field since removed would otherwise decode with that field silently gone, and two identities that differ only in it would compare equal on a resume.

env_spec_digest_of(config)

The environment's identity: what it will be built as, not how its config was spelled.

action_digest_of(env, env_spec)

What the policy can do: the parser, the space it lays out and the decision clock.

One function for the identity and for anything else that must agree with a run about its action space -- a demonstration shard, an artifact -- so the two cannot compute it apart.

extension_records(config)

The identity's record of every extension section the run uses, or None for none.

The digest is of the section's identity value, which names every file by its content digest rather than its path; that the digest is TRUE of the file is the section's verify, run at every start. A section provided by an installed package also records the package: which distribution declared it, its __version__, and its commit (package_provenance). A package whose commit cannot be named -- installed from a wheel, or not at the top of its own repository -- is named by its content instead, content:<sha256> over its files (package_content_digest), so that two runs on two different builds of it stay two runs.

section_digest(value, *, drop=('path',))

The digest of a config section by value, with every key named in drop left out at every depth.

What the section's identity is when the section names files: each file is named beside its path by a content digest, so dropping the paths keeps the content in the identity and keeps a folder moved to another disk from being a new experiment. A section whose files are named under another key passes that key in drop.

run_id(identity)

The short name of a run: sha256 of the identity's canonical JSON, first sixteen hex.

identity_differences(a, b)

Every field in which two identities differ, so that a refusal can name all of them.

A field one side never recorded is not a difference: nobody can show the two disagree. It is not a match either, and unverified is where it is said. Only that one field is set aside; the rest of the block it sits in is compared as usual.

unverified(a, b)

What a comparison of two identities could not check, one sentence each.

catalogue_digest(cards)

sha256 over the card catalogue's identifying fields, in card order.

Name, id, elixir and placement kind: enough that a catalogue with a card renamed, repriced, reordered or added is a different game, and little enough that a change to a statistic the engine already hashes into its calibration digest does not count twice.

engine_build(env_config, cards, *, path_search=None, stale_build_differences=())

The EngineBuild for a built environment, from its own config() and catalogue.

user_packages(config)

{top-level package: where its source is} for the user code this run's env comes from.

A component counts when its class lives outside the three packages, and so does any string in a component's kwargs that names a module under extra_component_modules -- which is how a composition refers to its parts, and the only place such a reference can be. A module that is merely LISTED is not counted: listing cannot move a number, and the identity table excludes extra_component_modules for exactly that reason.

The whole top-level package, not the one file, because a reward reads its helpers: an edit to mybot/util.py moves mybot.rewards.MyReward as surely as an edit to the reward itself.

user_code(config)

The identity's record of the user's own code: {package: sha256 of its source}.

What this closes: the identity named every component by its dotted path, so a different body of mybot.rewards.MyReward was the same run, and the ladder pooled two objectives. What it does not: data files a component reads, and code reached by anything other than a component's class path or a string in its kwargs.

dirty_sources(config_path=None, config=None)

What this run would read that is not in any commit, each named with what changed.

A run's identity records each repository by commit and everything else about the environment by NAME: the reward, the observation builder and the action parser are dotted paths, hashed as strings. So two different bodies of royalegym.reward.default_reward produce the same env_spec_digest, the ladder files their games under one context, and a plot draws two objectives as one line. On 2026-09-22 an uncommitted change to a sibling repo sat under a live run for twenty minutes and nothing in the identity could have said so.

What this covers, stated rather than implied: the python packages this process resolved, the engine's data directory, the config file the run was started from, and -- given the config -- the user's own packages the env's components come from. What it does not: any other file in those checkouts, a package installed from a wheel rather than a checkout, and anything a component reads at call time from somewhere else. The identity records the user's packages by content as well (user_code), so an edit this cannot see because it was committed still makes a different run.

THE ENGINE'S COMPILED CODE is outside this function, and since 2026-09-24 it is inside the identity instead. The extension is installed rather than imported from a checkout, so there is no working tree here to ask. What the engine does publish is the hash of the file it loaded, and EngineBuild.binary_sha256 records it, so two runs on two builds are two runs. What is still not published is which COMMIT built that file, so an engine built from uncommitted source is identified but not attributable. Reaching into a sibling checkout to guess would put back the guessed path this function removed. First measured and put to the engine's session by the integrator, 2026-09-22; the binary hash was RoyaleGym's answer.

The first version of this checked git describe --dirty on two hardcoded packages. It missed the engine's data, missed RoyaleViser, and was blind to untracked files -- including, at the time it was written, the config of the run then in flight. Found by the integrator, who measured all three rather than arguing them.

package_provenance(package)

<commit sha> or <commit sha>-dirty of the repository package is tracked in, or "unknown".

Stricter than git describe from the package's folder, which walks UP into any repository that encloses it: from inside a checkout's virtual environment it names that checkout's own commit for an installed copy that is in no repository at all. So a commit is recorded only when the repository's top level IS the folder above the package and the package's __init__.py is tracked there. The dirty flag is git status over the package's folder, untracked files included, as dirty_sources reads it. A sha rather than a description, because a description changes when a tag is added to the same commit.

package_content_digest(package)

sha256 over every file in package's folder, by relative path and bytes, compiled caches left out; or "unknown" for a package with no folder.

What names a package installed from a wheel, which has no commit: the same files give the same name wherever they are installed, and any change to one gives another.

royalegym_provenance() cached

(version, git description) of the royalegym this process imported.

Read from the imported module rather than from a pinned number here, because an editable install from a sibling checkout is the normal case in this family and its version string alone does not say which commit.

torch_version()

The torch this process would train with, or "absent".

describe_device(device='cuda')

"cuda:<name>:sm_86" or "cpu:<machine>".

The device is in the identity because a different GPU is a different set of kernels, and run-exactness is promised within a device class rather than across all of them.

compute_identity(config, *, env_spec, build, arch_digest, codec_version, codec_table_digest, torch_version_string=None, device_kind=None)

The identity of the run config describes, against the environment that was built.

The arguments beyond the config are the facts only preflight knows: what the environment turned out to be, what the codec decided to do with it, and what device and torch build the learner will use. They are passed in rather than probed here so that the function is pure and a test can vary one field at a time.

Rewards

Where the shipped reward function lives, and the shape your own one takes.

royalelearn.rewards

The reward the harness trains against: one terminal objective and three potentials.

Every shaping term here is a difference of a potential, gamma * Phi(s') - Phi(s). That form is what keeps shaping from changing which policy is optimal (Ng, Harada and Russell, 1999), and RoyaleGym's own reward.py states the principle in its header. The terminal win/loss term is the only term that is the objective; everything else densifies the signal and must leave the optimum where it was.

The discount in that expression is the run's, not a constant: set_gamma is called once per iteration with the value the schedule is at, and the same number rides the Step command to every worker, so the discount the reward is computed with and the discount the learner bootstraps with are one number rather than two that agree at the start of a run. It is deliberately not in config(): config() feeds the run identity and the ladder's context, and a quantity that moves every iteration does not belong in either.

WHY THESE THREE AND NOT ROYALEGYM'S SHIPPED DEFAULTS

  • TowerHPReward computes Phi(s') - Phi(s). It is one gamma away from being exactly policy-invariant, and with the gamma in place the term never needs annealing.
  • ElixirTradeReward is not a potential and it rewards turtling: a player who never plays a card never incurs the negative half, while enemy units still die to its towers and earn the positive half. At a weight that makes it visible it is comparable to the terminal reward.
  • ElixirLeakPenalty is not zero-sum -- both players can leak at once -- which is the signature of a term standing in for a potential that is missing.

CommittedElixirPotential replaces both. Playing a unit card moves elixir from the bar to the board and is net zero (a spell is charged at the tap, a Mirror play loses its own one elixir, and a Tri Wizards play is not priced right yet: see the class); losing a unit costs what the unit cost; killing one gains it; and sitting at ten elixir is penalised on its own, because the opponent's side of the potential keeps rising while yours cannot. It replaces the leak penalty's coefficient with a property of the state, and it needs no annealing schedule, which is what RoyaleGym's house rule -- weights should settle, not drift -- asks for. Its weight in the composition is configurable (see default_potential_reward); that is a scale, set once for a run, not a schedule.

EXACT ARITHMETIC

Every potential is a Fraction and the discounting is done on fractions, with one conversion to float at the end. The seat-mirror property is the reason: Phi(s, red) is exactly -Phi(s, blue) as a rational, so the two seats' rewards are exact negatives of each other and "the reward is zero-sum" is a property a test can assert rather than a tolerance it has to choose. A running float sum of the same unit values returns plus and minus 1.1e-19 on a perfect mirror -- harmless to a gradient, and enough to make the property uncheckable.

PotentialReward

Bases: RewardFunction

gamma * Phi(s') - Phi(s) for a potential a subclass computes from one state.

The subclass returns a Fraction from the seat's own point of view, antisymmetric between the two seats, and this class does the discounting and the single conversion to float.

set_gamma(gamma)

Take the discount the schedule is at. Exact: a binary float is a rational.

potential(state, team) abstractmethod

Phi(s) from team's point of view, exact and antisymmetric in the seat.

get_reward(team, prev, state, results)

gamma * Phi(s') - Phi(s), with Phi of a finished battle taken to be zero.

The zero is taken rather than computed, and it is what makes the shaping harmless. Summed over an episode the term telescopes to gamma^T * Phi(s_T) - Phi(s_0); Phi(s_0) is the same whatever the policy does, so if Phi(s_T) is zero as well the shaping adds a constant and cannot move the optimum (Ng, Harada and Russell, 1999). Read off the final state instead, Phi(s_T) is the margin -- a 3-0 win has three times the crown potential of a 1-0 win -- and the shaping would quietly pay for the margin beside the objective, which is the one thing these terms were chosen not to do.

A truncation is the other case and it is not this one: game_over is the engine saying the battle was decided, while a step limit cuts a battle that still has a position worth something, and the estimator bootstraps from that position's value.

PotentialCrownReward

Bases: PotentialReward

Phi = (own crowns - enemy crowns) / 3.

The crown difference is the match's own scoreboard, so the potential is the scoreboard read as a fraction of a whole win. It is the coarsest of the three and the one that most directly anticipates the terminal term.

PotentialTowerHPReward

Bases: PotentialReward

Phi = (sum own_hp/own_max - sum foe_hp/foe_max) / 3.

The same scoreboard as the crowns, read continuously: a tower at a third of its health is a third of the way to the crown it will give up. Dividing by the tower count puts it on the crown potential's scale, so the two weights mean comparable things.

CommittedElixirPotential

Bases: PotentialReward

Phi = ((own bar + own board) - (foe bar + foe board)) / scale.

bar is the elixir a player is holding and board is the elixir value of what that player has on the field, a unit at Fraction(card.elixir, card.count) of the card that summoned it. Crown towers are excluded: they were never played, and the tower potential already owns what happens to them.

What the term says, in the four cases that matter: a unit card played moves elixir from the bar to the board and is worth nothing; a unit lost costs what it cost; a unit killed gains it; and holding a full bar loses ground, because the opponent's side keeps rising while yours cannot. The last is the one that replaces a leak penalty, and unlike a leak penalty it is zero-sum.

WHICH UNITS ARE PRICED, and why this is not every unit on the board. An engine reports each entity under a card, and the entity need not be that card's own unit. On the RustEngine measured on 2026-09-24, a Goblin Gang's spear goblins were filed under the Goblin Hut, at five elixir each; a dying Golem's golemites under the Golem, at eight; a dying Battle Ram's barbarians under the Battle Ram, at four. Priced by the card they were filed under, one Goblin Gang play read +15 elixir of potential, a Golem's death +8, a Battle Ram's +4 -- under every run up to then, at every weight. Today's engine files all six of a Gang's units under the Goblin Gang; its spear goblins still do not match the Gang's row. So a unit is priced only when it IS the unit its card's catalogue row describes: same hitpoints, same collision radius, same air or ground. Anything else a card produced scores zero. That is an understatement and it is deliberate, the rule RoyaleGym's ElixirTradeReward adopted in 2381149: a golemite is worth something, and this term says zero rather than eight. What it keeps true is the property the term needs. ONE PLAY OF A UNIT CARD PUTS EXACTLY THAT CARD'S ELIXIR ON THE BOARD: a card's summon count covers exactly the units its row describes, so a Goblin Gang still totals three. tests/test_elixir_pricing.py taps every card in each engine's catalogue, on both seats, and names any card no tap landed for, because the rule rests on that.

One card breaks it, and it is a strict expected failure in that file: the Tri Wizards. From RoyaleSim round 9 a Tri Wizards play puts its Electro Wizard and Ice Wizard down under their own card ids, and each is its own card's unit, so the play prices at 7 + 4 + 3 = 14 for a card of 7: a Tri Wizards play reads as a gain of 7 while those two live. The default catalogue holds it. It is an event-only card, and its fix waits with the other event-only cards.

A MIRROR PLAY LOSES EXACTLY ITS OWN ONE ELIXIR. It pays the copied card's elixir plus its own one, and puts down a copy one level above the copied card. The copy is priced as the copied card's own unit, and the extra elixir buys the level, which this term does not price. The copy's hitpoints are the higher level's (a mirrored Knight has 1938 against the Knight row's 1766), so the row is read at the level the engine reports for the unit (EntityState.level), from the engine's own rows (RustEngine.unit_hitpoints), never from a guessed level curve. A unit at the catalogue's level is matched against the row itself, which the engine's rows agree with at that level for every card; the tests check it. An engine that reports no level (MockEngine, and RoyaleSim before round 9) or gives no rows (RoyaleGym before RustEngine.unit_hitpoints) is matched against the row alone, and there a Mirror copy scores zero: while it lives, the play reads as its whole cost thrown away. test_a_mirror_play_puts_the_copied_cards_elixir_on_the_board taps the Mirror on both seats.

SPELLS ARE CHARGED AT THE TAP, through the bar, by construction: the bar drops by the spell's cost and nothing it leaves on the board is priced, so the elixir comes back only through what the spell kills. That is the true trade, and it is what ElixirTradeReward charges too. A spell that leaves units behind -- a barrel -- is charged the same way, and its units score zero by the rule above. It telescopes like everything else here, so it moves no optimum.

unit_value(entity)

What one entity on the board is worth: its card's per-unit elixir, or zero.

Zero for anything that is not the unit its card's row describes -- a spear goblin filed under the Goblin Hut, a golemite under the Golem -- because pricing it by the card it is filed under is how a three-elixir play came to read as eighteen. The hitpoints compared are the row's at the unit's own level, so a Mirror's copy, one level up, is its card's.

PotentialCombinedReward

Bases: CombinedReward

RoyaleGym's weighted sum, with the run's discount threaded into the terms that take one.

The worker calls set_gamma on whatever reward function its environment holds, once per iteration, so the composition has to be the thing that forwards it. A term that does not take a discount -- the terminal one -- is left alone.

It also says WHICH of its terms is the objective. CombinedReward files each term in the logged breakdown under its class name, and the metrics group has to tell the objective from the shaping to publish either one: shaping_dominates is the alarm that watches for the shaping taking the objective over, and it compares those two sums. Matching on a class name would put a rename of a class into the arithmetic of an alarm, so the composition names the objective instead, and the breakdown carries it under TERMINAL_REWARD_TERM.

terminal_class property

The class name the objective's term would otherwise be filed under.

set_gamma(reward, gamma)

Give gamma to every term of reward that takes one, however it is composed.

A free function as well as a method because a composition is a tree: a CombinedReward may hold another one, and the discount has to reach the leaves of whatever a bot creator assembled rather than only the terms this module shipped.

shaping_weight_problems(weights)

Everything wrong with a set of shaping weights, every one named, all at once.

Shared by default_potential_reward and config.check_consistency so that a config is refused at load for exactly what the reward would refuse at build.

default_potential_reward(*, crown=0.2, tower_hp=0.1, elixir=0.05)

The shipped composition: the objective, and three potentials under it.

The terminal term is the objective at 1.0; the crown potential anticipates it; the tower potential is the same scoreboard read continuously; and the elixir potential is the fastest-moving of the three.

The three shaping weights are keyword arguments, so a config can set them::

"reward_fn": {"cls": "royalelearn.rewards.default_potential_reward",
              "kwargs": {"crown": 0.2, "tower_hp": 0.1, "elixir": 0.05}}

Every term is a potential difference, so no weight can change which policy is optimal: a weight SCALES a term's per-step magnitude and changes nothing about when it arrives. That holds for any real weight, negative ones included, so it neither justifies raising a weight nor bounds it. The bound here is a choice, stated below. What a weight does change is how loud each term is step to step, and that is what to read before changing one: env/reward_terms_step_abs/<term>, the mean of each seat's sum |F_t| over an episode. NOT env/reward_terms_abs/<term>, which is the episode's SUM: for a potential it telescopes to 1 - gamma times how far the potential wandered, so it falls as the discount schedule rises whatever the weights are.

How loud the shipped weights are, measured 2026-09-24 on a 100-card RustEngine environment, ten random-legal battles at gamma 0.999, per seat per episode [M]: elixir 0.80, tower 0.15, crown 0.15, against a terminal of 0.80. The elixir term alone is already about as loud as the objective. Each magnitude is linear in its weight.

A weight must be a real number, finite and not negative: not a bool, which Python would quietly take as 1, and not a string, which constructs and then fails inside a worker on the first step. A negative potential weight pays a seat for losing ground, and a NaN reaches every return it touches; both would train, silently. config.check_consistency applies the same rule when a config is loaded, so a bad weight is named before any environment is built.

Collecting experience

royalelearn.rollout

The rollout side: the env description, the shared-memory byte layout, and the workers.

Names resolve lazily. api/rollout.py imports EnvFactorySpec from here, and layout.py imports EnvSpec from there, so an eager re-export in this file would close that loop at import time; resolving on first use keeps both directions working whichever module a caller imports first, and keeps the cost of importing the env description down to msgspec.

royalelearn.obs_layout

The vector fields the network needs, resolved by name.

The pointer policy head reads three runs of the flat observation vector: the one-hot of the card in each hand slot, each slot's cost, and whether each slot is affordable now. Their offsets are not written down here and are not computed from the vector's width. They are looked up by name in the layout the observation builder itself declares, so that a field added, removed or reordered upstream moves them, a field the head needs and the builder no longer emits is a refusal at start-up naming it, and the same code runs against a sixteen-card catalogue and a sixty-five-card one without knowing which it has.

HandFields

Bases: NamedTuple

Where the hand lives in the observation vector, and the shape its one-hot unfolds to.

card_onehot is hand_size consecutive one-hot blocks of onehot_width -- the catalogue plus one for an empty slot. The width is read off the field's own size divided by the hand size the action space declares, never from the vector's width or the card count, so a builder that widens the block says so by widening the field.

field_slice(spec, name)

The slice of the observation vector holding name.

A missing name is a PreflightError that lists what the layout does declare, because the failure it catches -- a renamed field -- is otherwise a silently wrong slice of somebody else's numbers.

resolve_fields(spec, names=REQUIRED_FIELDS)

Every named field's slice, or a PreflightError naming the first one that is missing.

hand_fields(spec)

The three hand fields, checked against each other and against the action space.

The learner

royalelearn.learn

The learner: the networks, the masked distribution, and everything that touches a device.

Every module under here imports torch. Names resolve lazily so that importing the package, the config tree or the CLI's read-only commands costs nothing and works in an environment where torch is not installed; asking for one of these names without torch raises that import's own error, which already says which package is missing.

Rating and checkpoints

royalelearn.ladder

The ladder: the rating, who plays whom, what a gate decides, and where the evidence lives.

Names resolve lazily, so that reading a result log -- which is what the rating tests and any offline analysis do -- does not import an environment, a snapshot store or torch.

royalelearn.checkpoint

Checkpoints: per-component folders, a manifest that proves the bytes, and an atomic write.

Every component owns its folder and its own pair of methods, so adding one to a checkpoint is adding a folder name and a dict entry rather than editing a central serialiser. Nothing about a component's state is described in two places.

Atomicity. Everything is written into <name>.partial/, every file is fsynced, and then the directory is renamed into place. The directory itself is not fsynced: os.fsync on a directory handle is a POSIX guarantee and raises on Windows, where this harness's default profile runs, so the durability step is per file and the atomic step is the rename -- which is atomic on both platforms and is the property recovery actually needs. A crash mid-write leaves the previous checkpoint intact and the partial directory obviously named.

No pickle on any path this module writes. safetensors for weights, msgspec JSON for everything else, and torch.load(weights_only=True) wherever a component reads an optimizer state back. A checkpoint is loaded six weeks later from a directory nobody has looked at since; it must not be able to run code.

IndexEntry

Bases: Struct

One checkpoint, as the index knows it.

CheckpointIndex

Bases: Struct

The run's checkpoints, newest last.

latest() reads this rather than parsing directory names. int(x) for x in os.listdir(...) crashes the save on any stray file, and the save is called from the crash handler, which is exactly where the user most needs it not to.

RngComponent

Every random stream's position, as one folder of a checkpoint.

One reference learner saves none of this; the other re-seeds all three generators from the run's initial seed on load, so a run resumed at ten million steps draws the same action noise it drew at step zero. Because every stream in this harness is name-addressed, restoring the iteration counter and the shard positions restores the stream itself rather than merely the parameters -- the torch and python states below are for the few draws that are not name-addressed, such as a dropout mask.

capture()

Everything that would have to be true again for the next draw to be the same one.

DirCheckpointStore

Bases: CheckpointStore

One directory per checkpoint, named by the cumulative env step that produced it.

verify(path, manifest)

Every file in the manifest, hashed and compared. Names the first that fails.

This is what turns a truncated write from a crash six weeks ago into an error at the moment of the resume rather than into a run that continues from something else.

prune(run_dir=None, keep=None)

Keep the newest keep and remove the rest, along with any abandoned partial.

HOUSEKEEPING MUST NOT BE ABLE TO END A RUN. shutil.rmtree raises PermissionError on Windows while any file inside the folder is open, and reading a checkpoint while a run continues is an ordinary thing to do. This is called at every checkpoint -- about 1,990 times in a long run -- and before 2026-09-23 a single open handle anywhere in the oldest folder would have ended it.

A folder that could not be removed STAYS IN THE INDEX, so the next prune tries it again rather than losing track of it, and it is not reported as removed. The returned list is what actually went.

Found by a platform-portability review, which measured it here rather than arguing it.

sha256_of(path, chunk=1 << 20)

The hash the manifest records, read in chunks so a buffer file costs no memory.

config_differences(before, after)

Every leaf in which two configs differ, keyed by its dotted path.

Over the whole tree rather than the top level, so that a changed learning rate reads as ppo.lr_actor instead of as the whole ppo block having moved.

check_resume(manifest, identity, config, *, allow_drift=False)

What a resume has to agree about, checked once before anything is loaded.

A difference in an identity field is a refusal rather than a warning: a policy trained on one catalogue's observation vector cannot load into another's, and the failure that follows is a shape error deep inside a forward pass rather than a sentence naming the field. Every other difference is printed and continued past, because a run that resumes with a different checkpoint interval is the same run.

Returns the config differences it printed. allow_drift records the identity differences instead of refusing them, which is what --allow-identity-drift is for.

check_described(manifest, rows)

Refuse a checkpoint whose learner no metric row of its iteration describes.

A checkpoint's weights are supposed to be ones the run's record reports: the row of the iteration its manifest names carries the same state digest. One that breaks this holds a learner trained past its row -- which is what the emergency save wrote while the update ran before the batch was judged -- and resuming it continues the run from a state its record does not contain, under counters that belong to a different learner.

Any row of that iteration will do: a run resumed from an earlier checkpoint writes the iterations after it a second time, and each copy describes a learner that existed. What cannot be compared is let through rather than refused -- no metric file, no row of that iteration, a line torn by a crash mid-write -- because an absence says nothing about the weights. Past iteration zero, which never has a row, it is printed: a resume the record could not vouch for should not read like one it did.

Seeding and determinism

Every random draw in a run descends from one seed tree, so a resumed run lands on the original's curve.

royalelearn.seeding

Name-addressed random streams.

Every generator in the harness is derived from the run's master seed and a PATH -- a string naming the consumer, such as "ppo/minibatch/iteration/4/epoch/1" -- rather than by spawning children of a root SeedSequence in the order the code happens to ask for them. Positional spawning makes every stream downstream of a new consumer move, so adding a diagnostic that draws one number changes the actions a policy takes. Here a stream is a pure function of its name, so a config change that ought to be irrelevant is irrelevant, and two runs of the same identity draw the same numbers however their code paths are ordered.

The namespace below is the whole of it. A new consumer adds a row rather than reusing a neighbour's path, because two consumers on one path advance each other's stream.

Stream

Bases: NamedTuple

One row of the namespace: a path template and what draws from it.

derive_seedseq(master_seed, path)

The SeedSequence for one named stream of one run.

The path is hashed to a 128-bit spawn key rather than appended to the entropy, so that the master seed stays the run's single number and the path cannot collide with a seed value. blake2b because it is in the standard library, is fast on short strings, and -- the property that matters -- is fixed forever, which a string hash with per-process randomisation is not.

derive_generator(master_seed, path)

A PCG64 generator for one named stream. PCG64 is numpy's default and is stable across versions and platforms, which is what makes a recorded run reproducible on another machine.

derive_int(master_seed, path)

A reproducible 63-bit integer, for env seeds.

63 rather than 64 bits because the seed crosses into ClashSelfPlayVecEnv.reset(seed=...) and gymnasium's seeding refuses a value that does not fit a signed 64-bit integer.

stream_path(template, /, **fields)

Fill one row of STREAMS.

Going through here rather than writing an f-string at the call site means a typo in a path is a KeyError naming the template at the moment it is drawn, instead of a private stream that quietly works and silently fails to be the stream anyone meant.

royalelearn.determinism

The determinism tiers, and the process settings each one needs.

Three tiers, of which the first is unconditional (docs/harness-spec.md section 5.1):

T1 env-exact given the run identity, every episode's engine state-hash sequence, every observation, every mask and every reward are bit-identical, on any machine, forever. It costs nothing and is a property of the engine and of name-addressed seeding, so there is no switch for it and apply does not mention it. T2 run-exact T1, and every gradient, parameter and metric row bit-identical on the same device class and torch/CUDA build. The default, at 10-20% of throughput. T3 throughput T1 only; the learner may pick nondeterministic kernels.

Two of the settings T2 needs cannot be applied from here, because they must be in place BEFORE torch initialises CUDA and before numpy imports its BLAS: CUBLAS_WORKSPACE_CONFIG and the three BLAS thread counts. They are environment variables, the entry point sets them, and apply asserts rather than sets, so a run that would have been silently non-reproducible dies at start-up naming the entry point instead of producing a curve that cannot be repeated.

This module imports torch inside the functions that need it. Importing it at module scope would put a 300 MB, three-second dependency on the import path of royalelearn.cli, which has to work in an environment without torch at all.

apply_cublas_workspace_config(env=None)

Set CUBLAS_WORKSPACE_CONFIG if it is unset, and return its value.

Called by the entry point before torch is imported. An existing value is left alone even when it is not one this harness would have chosen: the operator's own setting wins, and require_cublas_workspace_config is what decides whether it is good enough.

apply_blas_thread_env(env=None)

Pin the BLAS thread counts to one. Must run before numpy is imported.

require_cublas_workspace_config(entry_point='royalelearn.cli', env=None)

Refuse to run T2 without a deterministic cuBLAS workspace.

cuBLAS reads this variable once, when CUDA initialises, so setting it here would be too late and setting it silently would be worse than not setting it at all: the run would look reproducible and would not be.

set_fill_uninitialized(on)

Turn torch's fill of unwritten memory on or off (docs/harness-spec.md section 5.1).

It has an effect only while deterministic algorithms are on. Read at every allocation, so a change takes effect at the next one; nothing already allocated is touched.

apply(tier, *, torch_threads=1, entry_point='royalelearn.cli')

Put the process into tier and return what was applied, for the record.

T2 keeps bf16 autocast. Deterministic kernels are bit-reproducible run to run at any precision, so the tier costs nothing in precision: it buys exactly what it says, which is two runs of one identity agreeing row for row. tf32 is off because it is not bit-reproducible across the different batch shapes the rollout forward and the update forward use.

Metrics

royalelearn.metrics

Where a run's numbers go: the schema they are checked against, the sinks that write them, and the alarms that read them.

Names resolve lazily, so that reading the schema -- which is what the documentation build and the metric tests do -- does not import a sink, a socket or wandb.

The base classes

royalelearn.api

Every ABC and struct the harness is written against.

This subpackage imports numpy and msgspec and nothing else that matters: torch appears in type annotations only, behind TYPE_CHECKING. That is what lets a rollout worker, the CLI's config and identity commands, and import royalelearn itself run in an environment with no torch installed -- and what keeps a worker's resident memory three hundred megabytes smaller than the parent's.

The concrete implementations live outside api/ and may import whatever they need:

RolloutSource        rollout/farm.py, rollout/inline.py
ActorCritic          learn/actor_critic.py
ObsCodec             rollout/codec.py
ExperienceBuffer     learn/buffer.py
AdvantageEstimator   learn/gae.py
Update               learn/ppo.py
Schedule             learn/schedules.py
Matchmaker, Rater    ladder/matchmaker.py, ladder/rating.py
MetricsSink, Alarm   metrics/sinks.py, metrics/alarms.py
CheckpointStore      checkpoint.py

AdvantageEstimator

Bases: ABC, Checkpointable

Rewards and values in, advantages and returns out.

compute(*, rewards, values, final_values, terminated, truncated, trainable, gamma, lam) abstractmethod

rewards/terminated/truncated/trainable are (T, R); values is (T+1, R); final_values is (T, R) and is read only where truncated. Returns (advantages (T, R), returns (T, R), stats).

Implementers MUST bootstrap a terminated cell from 0 and a truncated cell from final_values, and MUST NOT carry the recursion across an episode boundary. A truncation is an episode that was cut, not one that was decided, and treating the two alike throws away the value of every position a step limit ended.

AdvantageStats

Bases: Struct

What the estimator saw, for the metric row.

reward_scale is the divisor the return scaler applied and clipped_reward_frac is how much of the batch the clip bound touched: both are how a scaled reward stays legible after the scaling.

CodecTable

Bases: Struct

How each observation key is stored, decided from EnvSpec.obs_space at preflight rather than from a list of plane indices. Logged, hashed and written into every snapshot.

plane is one entry per spatial plane: its name, its storage in {"uint8", "float16", "static", "derived"}, and the divisor that takes the stored integer back to the value the environment produced. "derived" is reserved: nothing rebuilds such a plane yet, so the codec refuses it at bind. Two runs whose tables differ are not comparable, and digest() is what says so.

digest()

sha256 of the canonical JSON of this table.

ExperienceBuffer

Bases: ABC, Checkpointable

A rectangle of (T + frame_stack) cycles x R slots: T collected cycles, one bootstrap row, and frame_stack - 1 history rows carried over from the previous iteration. Owns the shared-memory block the workers write into.

Implementers may assume each (cycle, slot) cell is written exactly once, by the worker that owns that slot; the learner only reads.

shared_handle() abstractmethod

Name, size and the numbers the offsets follow from; picklable, and sent to workers.

record_round(r, actions, log_probs) abstractmethod

Scalars only: observations are already in place. O(n), no observation copy.

set_values(values) abstractmethod

(T+1, R) float32, from the whole-iteration critic pass.

(T+1, R) integer: how many actions each cell's mask left, from the critic's pass.

Implementers keep the collected cycles and may drop the bootstrap row, in which no action was taken. The count is what tells the update which rows had a choice at all without unpacking an observation to find out.

set_final_values(cells, values) abstractmethod

V(final_obs) for truncated cells; cells is int64[(k, 2)] of (cycle, slot).

set_advantages(adv, ret) abstractmethod

(T, R) float32 each.

trainable_mask() abstractmethod

(T, R) bool: which cells reach the update. What decides it is the seat's group and the cell's validity, never where the row was written.

batches(batch_size, minibatch_size, epochs, rng_for_epoch, *, choice_first=False) abstractmethod

Yield batches; each Batch knows its true sample count and iterates device-resident minibatches. A batch never straddles an epoch boundary, and an epoch holds as many whole batches as it can fill with the rows over spread one each across them, so every batch is at least batch_size and an epoch is exactly n // batch_size optimizer steps. An epoch with nothing trainable in it yields no batches at all. Each minibatch is weighted by its share of its own batch. Gathers per MINIBATCH, never per batch.

choice_first reorders each batch's cells so that the ones with more than one legal action come first, keeping the permutation's order inside each class. Implementers MUST move no cell between batches: it is a reordering, so that a caller skipping the forced rows skips whole minibatches of them, and every batch-level denominator is unchanged.

ObsCodec

Bases: ABC

Quantisation of one observation row. The worker packs; the learner unpacks on the GPU.

Implementers MUST be exact round-trips for the integer-valued channels and MUST declare row_bytes as a constant given an EnvSpec and its CodecTable.

codec_version abstractmethod property

The RULE's version. The table it produces is data and travels separately.

table(spec, sample, *, min_states=MIN_TABLE_STATES) abstractmethod

Decide storage per key from the declared bounds and a sample of real observations.

Storage is decided from a sample; EXISTENCE is decided from the declaration. A plane the layout does not declare static is stored even when it is constant across the sample, because the tower planes are constant in any sample in which no tower falls.

Implementers MUST refuse a sample of fewer than min_states observations. The row size of the whole run follows from this one decision, and a sample too small or too idle to have reached the states a plane varies in decides it wrongly and in silence.

static_planes(obs) abstractmethod

The planes EnvSpec.spatial_layout declares static; stored once per seat, never per row.

unpack_to_device(raw, statics, out) abstractmethod

Dequantise, scatter the static planes in, reshape the stored mask into the mask planes, and gather the frame-stack history.

Checkpointable

Bases: Protocol

Anything that goes into a checkpoint folder of its own.

FORMAT_VERSION is the component's own, not the checkpoint's: a component may change its files without the checkpoint format moving, and the manifest records each one so that a load can say which component it could not read.

load_checkpoint(folder, *, strict)

With strict=False, a missing file prints the exact path it wanted and continues with a default. With strict=True it raises. strict defaults to True on resume: tolerate-everything is right for a research tool and wrong for a harness that promises the curve continues.

CheckpointStore

Bases: ABC

Where checkpoints live, and the only thing that writes or reads them.

write(components, manifest) abstractmethod

Write atomically: into <name>.partial/, fsync each file, then os.replace.

The directory itself is not fsynced. os.fsync on a directory handle is a POSIX guarantee and raises on Windows, where this harness's default profile runs, so the durability step is per file and the atomic step is the rename -- which is atomic on both platforms, and is the property recovery actually needs.

read(path, components, *, strict) abstractmethod

Verify every hash in the manifest, then load each component. A mismatch raises CheckpointFormatError naming the first file that failed.

latest(run_dir) abstractmethod

The newest checkpoint, from the run index rather than from parsing directory names: int(x) for x in os.listdir(...) crashes on a stray file, and the save is called from the crash handler, which is where the user most needs it not to.

prune(run_dir, keep) abstractmethod

Remove all but the newest keep, and return what was removed.

Manifest

Bases: Struct

What one checkpoint is, beside the folders that hold it.

config is the resolved config verbatim, so a checkpoint is self-describing without the run directory around it, and files is every file in the checkpoint with its sha256, so a truncated write from a crash six weeks ago is an error on load rather than a silently wrong resume.

RngState

Bases: Struct

Every random stream's position at the moment a checkpoint was written.

One reference learner saves none of this; the other re-seeds from the run's INITIAL seed on load, so a run resumed at ten million steps draws the same action noise it drew at step zero. Because every stream here is name-addressed, restoring the iteration counter and the shard positions restores the stream, not merely the parameters.

ConditionResult

Bases: Struct

One gate condition, with the numbers that decided it rather than a bare pass or fail.

EvictionPolicy

Bases: ABC

Which snapshots the sampler stops drawing.

select_for_eviction(*, pool, ratings, max_sampled) abstractmethod

Ids to remove FROM THE SAMPLER. The archive and the result log are never touched: eviction is about sampling cost, not about forgetting evidence.

GateDecision

Bases: Struct

A candidate's audition, kept whole.

admit puts the snapshot in the pool, promote makes it the champion, and cycle is the case where a snapshot is worth playing against without being the best: the three are separate because a pool that only ever admits champions forgets everything it beat.

Matchmaker

Bases: ABC, Checkpointable

What every battle plays next.

plan(iteration, pool, ratings, geometry) abstractmethod

The iteration's opening table. One Assignment per battle, each drawn at that battle's current ordinal, so plan() is assign() applied across the geometry.

assign(battle, ordinal, pool, ratings) abstractmethod

The assignment for one battle's next episode.

Draws from match/battle/{battle}/ordinal/{ordinal}, so it is a pure function of the master seed and its two arguments: the same episode of the same battle always meets the same opponent, whichever iteration it happens to fall in and whichever worker holds it.

on_episode(record) abstractmethod

Record a finished episode, for the mixture's own bookkeeping.

PromotionGate

Bases: ABC

Whether a candidate snapshot joins the pool, and whether it becomes the champion.

Rater

Bases: ABC, Checkpointable

How a pile of results becomes a number per player.

fit(results) abstractmethod

id -> (rating in Elo units, standard error).

MUST be a pure function of results: same games in, same numbers out, in any order, on any machine.

predict(a, b) abstractmethod

P(a scores against b), with draws counted as half a win.

transitivity_residual(results) abstractmethod

How badly one number per player fits: near zero on transitive results, large on synthetic rock-paper-scissors.

RatingTable

Bases: Struct

One fit of the whole result log.

Ratings are in Elo units with the anchor pinned exactly, standard errors come from the inverse observed Fisher information, and transitivity_residual says how much of the result log a single scalar per player fails to explain -- which is the number that decides whether a scalar rating is lying.

SnapshotStore

Bases: ABC

Where frozen actors live. They outlive checkpoints: a rating means nothing without the player it rated.

get(snapshot_id, device) abstractmethod

LRU-cached: a shard-round touches at most ladder.max_resident_opponents of them, so a cache that never thrashes is a small one.

Alarm

Bases: ABC

A predicate over a metric row, with a severity and a patience.

severity="warn" logs and writes an alarms.jsonl row; severity="halt" additionally writes a checkpoint and a diagnostic bundle, then raises AlarmHalt. An alarm can stop a run and can never alter a value, which is why its thresholds are recorded but excluded from the run identity.

holds(row) abstractmethod

Whether the condition is true for this row, ignoring patience.

message(row)

What to say when it fires. The default names the keys and their values.

AlarmResult

Bases: Struct

One alarm's verdict for one iteration.

consecutive is how many iterations in a row the predicate has held, so a row written before the patience is spent still records that something was building.

MetricsSink

Bases: ABC, Checkpointable

Somewhere a row goes.

write(row) abstractmethod

One flat dict per iteration. Implementers MUST NOT mutate the row and MUST NOT raise on an unknown key.

write_episodes(rows)

Finished episodes. Default: ignore.

write_alarms(alarms)

Alarms that fired this iteration. Default: ignore.

write_artifact(name, path)

A file produced this iteration -- a heatmap, a bundle. Default: ignore.

ActionDistribution

Bases: ABC

A categorical distribution over the action space, with the illegal actions removed.

sample(uniforms) abstractmethod

(B,) int64. Driven by CALLER-SUPPLIED uniforms in [0, 1) so that the sampled action is reproducible independently of batch composition and of torch's global RNG.

mode() abstractmethod

(B,) int64, mask-respecting argmax.

log_prob(actions) abstractmethod

(B,) float32.

entropy() abstractmethod

(B,) float32, over the whole legal set.

noop_entropy() abstractmethod

(B,) float32: the binary entropy of p(no-op) against p(play).

The leading indicator of no-op collapse, and bounded by 0.693 nats, which is why the coefficient that acts on it is not the joint entropy's.

(B,) int64. Diagnostics: raw entropy falling is ambiguous without it.

Actor

Bases: ABC

The policy network. Implementations are also torch.nn.Module.

logits(obs) abstractmethod

(B, spec.n_actions) float32 ALWAYS, even inside autocast. Raw logits: no softmax, no clamp, no mask.

distribution(obs)

The default masked categorical over this actor's logits.

The import is here rather than at module scope because the shipped distribution is a torch module and this file must import without torch.

ActorCritic

Bases: ABC

The pair, as the learner uses them.

Implementers may assume obs tensors are on the device and already dequantised, and that obs.mask[:, 0] is True on every row (RoyaleGym sets mask[NOOP]=1 unconditionally, action.py, including after game over). They MUST assert it.

backprop(obs, actions) abstractmethod

Recompute under the current parameters, applying exactly the mask read from the buffer -- never a recomputed one.

actor_state_dict_fp16() abstractmethod

What a pool snapshot stores: actor only, fp16, no optimizer.

ActResult

Bases: Struct

What one rollout forward produces. Everything after the first two is diagnostics.

BackpropResult

Bases: Struct

What one update forward produces, recomputed under the current parameters.

Critic

Bases: ABC

The value network. Implementations are also torch.nn.Module.

value(obs) abstractmethod

(B,) float32.

NetworkFactory

Bases: ABC

How an ActorCritic is built, and how two of them are told apart.

arch_digest(spec, arch) abstractmethod

sha256 of the canonical JSON of (arch, obs_space, frame_stack, num_cards, n_actions).

A snapshot or checkpoint built from a different architecture is refused by this digest, not shape-errored halfway through a load.

ObsBatch

Bases: NamedTuple

One batch of observations on the device, already dequantised.

Shapes follow from EnvSpec and the frame stack k: spatial is (B, k*S, H, W), mask_planes is (B, k*A_s, H, W) for the A_s per-slot action planes, vector is (B, V) and is the CURRENT frame's only, and mask is (B, n_actions) bool.

Assignment

Bases: Struct

What one battle's next episode is.

Drawn by the Matchmaker at the battle's own episode boundary, from match/battle/{b}/ordinal/{k}, and therefore a pure function of the master seed, the battle index and the reset ordinal.

Close

Bases: WorkerCommand

Shut the shard down.

Defer

Bases: WorkerCommand

Nothing this round; hand back the same data.

EnvSpec

Bases: Struct

Everything the learner must know about the environment before it builds anything.

Every field here is READ from the running environment at preflight. The harness holds no layout constant of its own: a width, a plane count and a field offset are all environment facts, and typing one into this repo would make a RoyaleGym change a silent wrong answer instead of a loud one.

n_planes property

Spatial planes the environment emits, static ones included.

n_grid_actions property

The no-op plus one action per (hand slot, tile): the part of the action space the pointer head lays out as a grid, and all of it for an environment without buttons.

n_buttons property

Ability buttons (a hero's, a champion's): the actions after the grid, action n_grid_actions + k pressing button k. Their readiness is the observation's ability_ready, which is the same bits as the mask's last n_buttons; 0 for an observation without it.

static_planes property

The planes the layout DECLARES static -- never the ones a sample found constant.

EpisodeRecord

Bases: Struct

One seat's finished episode: what it was, how it went, and enough to replay it.

The terminal scalars are copied straight out of the environment's final_info rather than reconstructed from the states the worker saw. The environment already computes them, and a second implementation of the same summary is a second answer. RoyaleGym's EPISODE_STAT_KEYS is the authority on the set.

ObsKeySpec

Bases: Struct

One key of the environment's observation space, read from the space itself.

low and high are per channel for a three-dimensional key and length one otherwise. A space that declares scalar bounds broadcasts them over the whole array, so the per-channel form is a reduction of the declared bounds over each plane rather than a second opinion about them: it is what the codec needs in order to decide a plane's storage (section 7.2).

Plan

Bases: WorkerCommand

The iteration's opening table, once per iteration.

RolloutRound dataclass

One shard-round of observations. Holds zero-copy views; not a Struct.

The observations themselves are not here: the worker packed them straight into their final resting place in the experience buffer, and obs_rows says where each slot's row went. What crosses the boundary is the scalars and the episodes that ended.

TWO TIMESTEPS, AND WHICH IS WHICH. A round is published after a step, so it carries the state that step reached and the step itself, and those are not the same timestep.

  • obs_rows, tick and group are the round's OWN cycle: where this state was written, the engine clock it stands at, and who holds the seat from here on.
  • reward, terminated, truncated, deploy_status and episode_end are the step that ARRIVED here, which began one cycle earlier -- so they belong to the row below this one in the rectangle, which is the row whose action produced them.
  • valid is the publication: false says this round did not arrive, so neither the state nor the step it reports exists.

The round at cycle 0 has no step behind it and carries zeros for that half. The round at cycle T has no row of its own and is the only carrier of row T - 1's step.

deploy_status is -1 for no command and 0 for an accepted one; 1..11 is a DeployStatus refusal, which under a correct mask cannot happen and is therefore a mask bug rather than a tolerance.

RolloutSource

Bases: ABC

Where experience comes from. The one seam between the learner and the world.

Implementers may assume: begin_iteration precedes any next_round; exactly one submit per next_round; Step.actions are legal under the mask that was handed out; the caller does not retain a RolloutRound's views past the next next_round for the same shard. Implementers MUST guarantee: slots is ascending; every slot appears exactly once per cycle; a dead worker's slots arrive with valid=False rather than not arriving.

stats()

Per-round timing: env_ms, wait_ms, codec_ms, parent_wait_frac, bytes_out.

parent_wait_frac is the PARENT's share: time it spent blocked on a round that had not been published, over that plus the workers' own env time. It rises when the workers cannot keep up, which is the opposite reading from a number about workers idling, and the opposite action. An inline source reports zero because there is nobody to wait for. Default: {}.

SetState

Bases: WorkerCommand

Start each battle from a recorded position; one blob per battle, None to leave it.

SlotPlan

Bases: Struct

The iteration's opening assignment table: what every battle is playing at the moment the iteration begins.

Assignments change only at a battle's own episode boundary, where the parent draws a fresh one; nothing in an iteration's span reassigns a battle mid-episode, so a partially controlled trajectory cannot occur.

Spaces

Bases: WorkerCommand

Re-read the spaces without stepping.

Step

Bases: WorkerCommand

Step this shard by one decision.

group, opponent_ix and learner_seat carry the parent's assignment for every battle whose episode started on the previous round, and are empty otherwise; the worker needs them only to know which slots it fills with a scripted action.

WorkerCommand

Bases: Struct

What the parent hands a shard for one round. Six variants, one per round.

A tagged union rather than a dict of flags: the worker's loop switches on the tag, so a command it does not know is a protocol failure at the boundary instead of a missing key deep inside a step.

WorkerFailure

Bases: Struct

A worker that stopped answering, as a value the coordinator can act on.

kind is one of "exception", "crash", "timeout" or "protocol", and message is the child's traceback verbatim. Both reference learners treat a dead worker as a permanent silent hang; here it is a typed failure, which is what lets the farm restart the worker and the run continue with the loss recorded rather than hidden in a throughput dip.

Schedule

Bases: ABC

One scheduled scalar.

ScheduleState

Bases: Struct

Every scheduled quantity, evaluated once at the top of an iteration.

Evaluated once and passed down, rather than evaluated where each is used: the discount the reward was computed with, the discount GAE used and the discount the metric row reports are then the same number by construction.

credit_horizon_seconds(decision_ms)

1 / (1 - gamma * lambda) decisions, in seconds.

decision_ms comes from the environment, so the horizon follows a change in the decision granularity instead of quietly meaning something else. Logged every iteration, because it is the single highest-leverage number in the config and because it can otherwise be ten seconds by accident.

ActorLossTerm

Bases: Protocol

A term an extension adds to the actor's loss.

loss returns (coefficient, raw). The update adds coefficient * raw * scale, where scale is actor_scale for a term whose raw value is a sum over the choice rows (scaling == "rows", which keeps minibatch size a pure memory knob and matches the policy term under every ppo.forced_rows value) and the minibatch weight for one whose raw value is a mean of its own (scaling == "minibatch"). The gradient ratio, when asked for, is measured on raw * scale, before the coefficient.

finish returns the term's keys for the row; every one must start with <extension>/. actor_trained is False on a frozen iteration, when loss was not called at all. state is what a checkpoint keeps (an empty dict keeps nothing); format_version is saved beside it and a different one refuses the resume.

ActorTermInputs

Bases: NamedTuple

What an extra actor-loss term is handed, identically in both of the update's actor paths.

log_probs and mask are the actor's masked log-probabilities and mask on the minibatch's choice rows, in rows order, with the graph attached. actor is the live actor, for a term that needs a forward of its own on other rows. actor_scale is the per-batch scale the policy term is multiplied by, and weight the minibatch's share of its batch.

cells names each row of the minibatch by its position in the iteration's whole batch, in the minibatch's row order (cells[rows] names the choice rows). Every epoch trains every cell once, so a term whose input on a row does not change within an iteration -- a frozen reference's forward -- can compute it in the first epoch and look it up by cell after. None from a caller that does not have them.

Update

Bases: ABC, Checkpointable

The learner's optimisation step. Owns the optimizers, so it is checkpointable.

step(buffer, sched) abstractmethod

Consume one iteration's rectangle and return what happened.

Implementers may assume the buffer's advantages and returns are already computed and that every cell it hands out is trainable and valid. They MUST apply exactly the mask read from the buffer, never a recomputed one, and MUST leave the buffer unchanged.

UpdateResult

Bases: Struct

What one update did, in the units the metric row uses.

kl_by_epoch and clip_fraction_by_epoch are per epoch rather than averaged: the rule for n_epochs is read off them ("if epoch three's clip fraction is more than twice epoch one's, lower it"), and an average cannot answer that. samples_unused_frac reads zero by construction and is logged so that it can be seen to.