Skip to content

royalegym

The environment layer. This is the package you import to get a battle your bot can play in.

Everything on this page is generated from the docstrings in the code, so it cannot drift from what the package actually does. If something here is wrong, the docstring is wrong. It covers the main modules, not every one. royalegym.evaluate, which scores one bot against another, and royalegym.opponents, the scripted opponents, are not on this page. Read their docstrings in the code.

If you are looking for where to start rather than what a name does, read the Quick Start Guide and the Overview first.

The environments

royalegym.env

Environments: PettingZoo ParallelEnv, Gymnasium single-agent wrapper, vectorised self-play.

DECOMPOSITION ObsBuilder, ActionParser, RewardFunction, StateMutator and two DoneConditions (one in the termination role, one in the truncation role) are constructor arguments. Changing a reward or an observation is a Python edit and never a recompile of the simulator, because that is where most research iteration happens.

DONE FLAGS Gymnasium's terminated comes from termination_cond (default GameOverCondition) and truncated from truncation_cond (default none). Both are consulted every step so counters advance; a termination on the same step as a truncation is reported as a termination only.

TIMING One env step = one decision = decision_ms of game time, rounded UP to a whole number of engine ticks (engine tick length comes from the engine's state, which reads calibration.json). Both players act simultaneously; the engine validates both commands against the same pre-step state.

ACTION MASKS Every observation dict carries action_mask (int8, as pettingzoo's parallel_api_test and gymnasium Discrete.sample(mask=...) expect), and mask_planes -- the same mask without the no-op, as [4, 32, 18] -- beside it. action_masks() returns the bool array sb3-contrib's MaskablePPO calls for. An action the engine rejects is turned into a no-op and reported in info["deploy_status"]. The info dict does NOT repeat the mask: it was the same array the observation already carried, batched a second time by every vector env for nothing.

EPISODE STATISTICS The info dict of the LAST step of an episode carries how that episode went -- length, crowns, tower hp, elixir leaked -- so a training run reads it out of infos["final_info"] instead of keeping a shadow copy of the state. EPISODE_STAT_KEYS names them. outcome and winner are there too, as before, but only when the engine ended the battle: a truncation has no winner.

It also carries the reward TERM BY TERM, as ``reward_sum/<TermName>`` summed
over the episode for that seat. That is the number which says which term a
policy is actually chasing, and without it a run logs one scalar and cannot tell
a win from a well-farmed shaping term. Those keys follow the reward function's
own terms rather than a fixed list, which is why they are not in
``EPISODE_STAT_KEYS``. ``log_reward_terms=True`` adds the same breakdown every
STEP as ``reward/<TermName>``, for watching one battle rather than training.

THE VIEWER ClashParallelEnv publishes only when it is HANDED a publisher, and reads no environment variable of its own: N envs each binding the viewer's one fixed UDP port is an OSError, not a feature. ClashSelfPlayVecEnv owns that decision instead and hands one publisher to one game (see its viser argument).

ClashParallelEnv

Bases: ParallelEnv[str, dict[str, ndarray], int]

A two-seat battle ("blue" and "red"), PettingZoo parallel API.

Each piece is a config object you can swap: the engine, the state mutator (how a battle starts), the obs builder (what each seat sees), the action parser (what it can do), the reward function, and the termination and truncation conditions. make_env builds one with sensible choices for all of them.

reward_terms(team)

This episode's reward, term by term, for one seat.

The number a training run plots and the one that says WHICH term a policy is actually chasing. It reaches a trainer through the terminal info as reward_sum/<TermName> -- and through final_info after an autoreset, which is the only place it survives -- so nothing has to reach into the env to get it. The keys follow the reward's own terms, so they are not in EPISODE_STAT_KEYS; the reward function is the authority on the set.

episode_stats(state, team)

How the episode that just ended went, from team's seat.

Written only on the terminal step, so a rollout buffer carries one of these per EPISODE rather than one per step; EPISODE_STAT_KEYS is the key list. Tower hp is the MEAN of the three towers' hp fractions (tower_hp_frac), which is TowerHPReward's potential over three -- so what a run logs and what its shaping term optimised are the same number. Crowns alone cannot tell a tower left at 1 hp from a tower never touched, which is why it is here.

config()

What this env IS, as a JSON-able dict, for a checkpoint to record.

Class names and constructor kwargs of every component, the decision granularity, the Reveal the observation was built with, and digests of the data underneath: calibration.json as it is on disk, plus the copy the compiled engine was built with when there is one. A checkpoint that pins this can say whether a policy is being evaluated on the env it was trained on -- including whether it was trained with hidden information revealed, which nothing about a weights file would otherwise show.

A description, not a constructor: EnvFactory is the picklable recipe that BUILDS one.

action_masks(agent='blue')

Boolean legal-action mask, the shape sb3-contrib MaskablePPO expects.

state()

Global state for centralised critics: Blue's observation, flattened.

Layout is state_space (every non-mask key of the observation Dict, in the Dict space's key order -- MASK_KEYS are left out). It inherits Blue's imperfect information: Red's elixir is the count Blue's builder keeps and Red's hand is not there at all, unless the builder's Reveal opens them.

ClashGymEnv

Bases: Env[dict[str, ndarray], int]

Single-agent Gymnasium env: one seat is the learner, the other an Opponent.

WHICH ENGINE YOU GET. With no engine= this builds a MockEngine, which is a readable reference implementation and not the game: different card table, spells that resolve instantly instead of travelling, no stuns or knockback. That is the right default for trying the API out, and the wrong thing to train a bot on believing it is the real one -- so it says so, once, rather than leaving the reader to find out from a result that does not transfer.

royalegym/ClashRoyaleMock-v0 and royalegym/ClashRoyaleRust-v0 say which they are in their names and neither warns. Pass engine= here and nothing warns either: the warning is about not having chosen, not about the mock.

ClashSelfPlayVecEnv

Bases: VectorEnv[Any, Any, Any]

N simultaneous games exposed as 2N agent slots: slot 2i = blue, 2i+1 = red, of game i.

Both seats feed one batch, which is how a single shared policy gets both players' experience in self-play. Autoreset is SAME_STEP: when a game ends, the returned observation is already the next game's first one, and the final observation / info are in infos["final_obs"] / infos["final_info"] (masked by infos["_final_obs"]). final_info is where an episode's statistics arrive -- EPISODE_STAT_KEYS.

THE VIEWER, and why it is decided here. A viewer watches ONE battle: it has one fixed UDP port, and N envs each binding it is OSError 10048, which is what ClashSelfPlayVecEnv(num_games>1) used to raise the moment ROYALEVISER was set. So the publisher is bound once, here, and handed to game 0 only. viser="env" (the default) means ViserPublisher.from_env(): a publisher when ROYALEVISER=host:port is set, and None -- costing nothing -- when it is not. Pass a ViserPublisher to bind one explicitly, or None never to publish.

ADDRESSABLE EPISODES, and why a run cannot resume without them. With autoreset_seed_fn unset, an autoreset calls reset() with no seed, and ClashParallelEnv.reset leaves its generator alone when the seed is None -- so the battle a game plays depends on how many battles that game has already played. A fresh worker cannot arrive at episode 400 without playing 399 first, which is why a resumed run diverges from the one it is continuing even when the learner itself came back byte for byte.

Set autoreset_seed_fn and each episode is named instead of counted: the seed for the nth episode of game g is fn(g, n), so any episode can be reached directly. episode_ordinals is the counter to checkpoint and set_episode_ordinals puts it back, after which the stream continues row for row rather than restarting::

fn = lambda game, n: (run_seed * 1_000_003 + game * 9973 + n) % 2**31
vec = ClashSelfPlayVecEnv(8, autoreset_seed_fn=fn)
...
saved = vec.episode_ordinals          # into the checkpoint

vec = ClashSelfPlayVecEnv(8, autoreset_seed_fn=fn)
vec.set_episode_ordinals(saved)       # out of it, before reset()
vec.reset()

The default is None and changes nothing: reset(seed=None) is what the autoreset already did. An explicit reset(seed=...) still wins over the function, and still consumes an ordinal, so the numbering keeps meaning "the nth episode this object started in game g" either way.

episode_ordinals property

Episodes started so far, per game. The number a checkpoint has to carry.

Without it a resumed run can name its episodes and still not know WHICH one to name next, so it starts at 0 and replays a run it has already played.

set_episode_ordinals(ordinals)

Put the counter back where a checkpoint left it, before reset().

The other half of episode_ordinals. Call it before reset(): reset starts an episode and therefore consumes an ordinal, so setting it after would skip one.

The five pieces you swap

These are the parts you write. Each one is a base class with a working default already shipped, so you can replace one and leave the other four alone.

royalegym.reward

Reward functions: (previous state, state, deploy results) -> scalar, per team.

Composable: CombinedReward([(WinLossReward(), 1.0), (TowerHPReward(), 0.3)]).

A WARNING ABOUT WEIGHTS A reward coefficient that gets re-tuned every time a new behaviour is measured is standing in for a missing reward term. If the agent learns to hoard elixir and you respond by nudging ElixirLeakPenalty up, then down, then up again as other behaviours shift, the weight is absorbing something the reward cannot express -- find that thing and give it a term (or remove a term that is fighting the terminal signal). Weights should settle, not drift. The terminal win/loss term is the only one that is the actual objective; everything else is shaping, and shaping that is not a difference of a potential function can change which policy is optimal.

ZERO-SUM Every term here is antisymmetric between seats on a mirrored transition (reward_blue == -reward_red), except ElixirLeakPenalty, IllegalActionPenalty and PlacementDepthReward, which score only the acting player's own behaviour. Tests check the antisymmetry, because a self-play reward that is not zero-sum rewards both players for colluding.

POSITION DeployResult carries the x and y a command was evaluated at, in the ENGINE frame. Any term that scores WHERE something was played must convert with protocol.to_own first, or it rewards Blue and punishes Red for the same placement and the self-play run learns the seat rather than the game. PlacementDepthReward is the worked example.

RewardFunction

Bases: ABC

What a bot is paid for. Write get_reward(team, prev, state, results): one number per seat per step, bigger is better. bind, reset and config are optional.

bind(engine)

Receive static engine data (card catalogue etc.). Optional.

reset(state)

Called at episode start. Optional.

config()

Constructor state, JSON-able, for ClashParallelEnv.config().

WinLossReward

Bases: RewardFunction

+1 on the transition into a win, -1 into a loss, draw for a draw.

CrownReward

Bases: RewardFunction

Change in (own crowns - enemy crowns).

TowerHPReward

Bases: RewardFunction

Enemy tower HP destroyed minus own tower HP lost, as fractions of max HP.

This is a difference of a potential (sum of normalised tower HP), so it does not change which policy is optimal for the terminal objective -- it only densifies the signal.

ElixirTradeReward

Bases: RewardFunction

Elixir the enemy spent and lost minus what this seat spent and lost, / MAX-ish scale.

Two things spend elixir and leave nothing behind: a unit that dies, and a spell that is cast. A unit is valued at card elixir / units summoned by that card. Entities that vanish count as killed whether by damage or by lifetime expiry -- a building that times out WAS spent elixir, so counting it as a loss is the honest accounting, not a bug. Crown towers are excluded (TowerHPReward covers them). A unit born and killed inside ONE decision step is in neither snapshot, so it is never priced for either seat; that has always been so and the spell charge below does not change it.

A SPELL IS CHARGED AT THE TAP, from the accepted deploy results, and never from the board. It is consumed the moment it is played, so its elixir is gone whether it killed anything or not, and a Fireball that hits air has to cost four. The board cannot say this: an engine whose spells resolve inside the tick they land reports no spell object at all, and the spell objects that do appear live for several ticks and carry no identity, so counting them once is not possible. Every engine reports an accepted tap exactly once, which is why it is read there. Without this the term paid for a spell's kills and charged nothing for the spell, which teaches that spells are free.

IT IS CHARGED WHAT THE PLAY COST, not the card's own elixir: the price the engine stated for that hand slot just before the tap (PlayerState.hand_costs), where it states one. For most cards the two are the same number. A Mirror's are not: it costs the card it copies plus its own one, and puts down a copy one level up, which matches no row and so scores nothing, like a Goblin Barrel's goblins. Charged its own one elixir, a Mirror of a Knight cost the term 1 where the engine took 4.

WHAT THE CATALOGUE DOES NOT PRICE. A card can put units on the board that are not the unit the card itself summons -- a hut and a Witch keep producing them, a Tombstone leaves more behind when it dies, a barrel releases them where it lands. The catalogue has one row per card and no row for any of those units. An engine before RoyaleSim 0.1.5 reports each of them under SOME card that can produce it, which need not be the card its owner played: a Tombstone's skeleton is reported under the Witch, a five-elixir card its owner may not even hold. From 0.1.5 an entity reports the card whose play put it down (the Tombstone's skeletons report the Tombstone). Either way a unit is paid for only when it is the unit its own card's row describes -- same hitpoints, same collision radius, same air or ground -- and anything else a card produced scores nothing. THAT IS AN UNDERSTATEMENT AND IT IS DELIBERATE: a Tombstone's skeleton is worth something, and this term says zero rather than five. Now that an entity says which card produced it, a produced unit could be priced instead; this term does not do that.

WHAT THAT LEAVES TRUE is the property the term actually needs: ONE PLAY OF A CARD IS WORTH EXACTLY THAT CARD'S ELIXIR, charged once, either at the tap or through the units it left behind, never both and never neither. Not every unit a tap puts down matches the row -- a Goblin Gang puts three spear goblins down beside its three goblins, a Rascals puts two girls down beside the boy, and the row describes one kind of each -- but the card's summon count covers exactly the units the row does describe, so the play still totals the card's price. What it costs is precision WITHIN a card: kill a Goblin Gang's spear goblins and the term pays nothing, kill its other three and it pays the whole three elixir. No produced unit matches the row of the card it is reported under. Both are measured in tests/test_rewards.py over every card either engine will place, on both seats, because the whole rule rests on them.

ONE CARD BROKE IT BEFORE ROYALESIM 0.1.5: THE TRI WIZARDS. From RoyaleSim 6909b6f to 0.1.4 their Electro Wizard and Ice Wizard are reported under THEIR OWN card ids (42, 23), not the Tri Wizards', so each matches its own card's row and is priced as that card: a 7-elixir play totals 14. From 0.1.5 every unit reports the card whose play put it down, and the play totals 7. tests/test_rewards.py grades the card on its own (PRICED_ELSEWHERE): an expected failure on an engine with the old labels, a pass on one with the new.

EXACT ARITHMETIC. Unit values are Fraction(elixir, count) and the sum is exact until the final division. A running float sum of the same values depends on entity iteration order, so on a perfectly rotation-mirrored transition (the same kills on both sides) it returns +1.1e-19 for one seat and -1.1e-19 for the other instead of 0 for both -- measured on both engines by tests/test_rust_engine.py's multi-unit rotation test. Harmless to a gradient, but it makes "equal rewards on a mirror" an uncheckable property.

unit_value(e)

The card's per-unit value, or zero for a unit the catalogue does not price.

ElixirLeakPenalty

Bases: RewardFunction

-1 per decision step spent sitting at full elixir (regeneration wasted).

Not zero-sum: both players can leak at once. max_milli defaults to the engine's MAX_MANA via calibration when bound through an env.

PlacementDepthReward

Bases: RewardFunction

How far up the board this team's accepted placements were, in own-frame tiles.

The template for every positional shaping term, and the reason DeployResult carries coordinates. Positive is forward: a placement in the enemy half scores above one behind your own towers, scaled so a placement at the far end is 1 and at your own back line is -1. weight on the aggressive side is the knob; a NEGATIVE weight rewards defending at home.

NOT IN default_reward AT ALL, on purpose -- not at zero weight, not at any weight. Whether pushing or defending is better is the thing a bot is supposed to learn, and a shaping term that answers it in advance is the weight-drift this module's docstring warns about. It is here to be copied and to prove the coordinates arrive, not to be switched on untested. test_the_positional_term_stays_out_of_the_default_reward keeps it out.

Antisymmetric between the seats on a mirrored transition, like the terms above: the own frame flips with the seat, so the same placement scores +d for one player and is scored by the other only through its own placements.

IllegalActionPenalty

Bases: RewardFunction

-1 for each of this team's commands the engine rejected this step.

With a correct mask and a masked policy this is always 0; a non-zero value in training logs is a mask bug or an unmasked policy, which is why it exists.

CombinedReward

Bases: RewardFunction

Weighted sum of terms. last_terms keeps each weighted term for logging.

KEYED BY TEAM, because the env calls get_reward once PER SEAT on the same transition. A single flat breakdown was cleared at the top of every call, so the first seat's numbers were destroyed by the second and whatever a training run logged as "the reward breakdown" was only ever Red's -- while the scalar rewards, which are returned rather than stored, were right for both. Read one seat's with terms_for(team).

terms_for(team)

The weighted breakdown of team's last reward; empty before the first.

default_reward()

Terminal objective plus light, potential-style shaping.

royalegym.obs

Observation builders: BattleState -> what the policy sees.

PERSPECTIVE Every observation is in the ACTING player's own frame (see protocol.py): Red's board is rotated 180 degrees and its towers/crowns/elixir are "own", Blue's are "enemy". A mirrored battle therefore yields a bit-identical observation for the other seat, which tests assert, so one policy plays both.

FAIR INFORMATION, AND THE Reveal Everything a builder writes by default is something a player watching the match could write down: the board, their own hand and cycle, the clock, and a COUNT of the opponent's elixir kept from the plays they saw and the regeneration rate everyone knows (MatchMemory). Nothing default-built is read out of the half of the state a player cannot see, with one known exception: the builders read no status flags, so a unit invisible to its enemy (a Ghost, STATUS_INVISIBLE) is still shown to the enemy seat where it stands. Hiding it or marking it waits on a measurement of what the live client's opponent sees.

``Reveal`` opens that half, one field at a time, for curriculum, distillation
and debugging. An enabled field ADDS its channels and vector slots; it is
never present-but-zero, so a fair observation and a cheating one do not even
have the same width, and a checkpoint cannot quietly be trained on one and
evaluated on the other. ``ClashParallelEnv.config()`` records the Reveal for
exactly that reason. The one field that changes a slot instead of adding one
is ``enemy_elixir``: the fair slot already holds the counted value, and the
reveal swaps in the value read from the state. The two agree on a played-out
battle, presses included (tests/test_env_obs.py, tests/test_press_memory.py); a
disagreement is a bug in the counter, or one of the misses ``MatchMemory`` names and
flags with ``exact``.

Floats appear here and only here-onwards (policy input). They are computed from integer state by the same operations for both seats, so the flip is exact.

NO FLOAT MAY DEPEND ON ENTITY LIST ORDER state.entities order is engine-private and is NOT seat-canonical: the Rust engine lists entities by storage slot, and slots are reused, so an exactly rotation-mirrored battle can list Blue's twins in a different order from Red's. Float addition is not associative, so any per-entity float accumulation (or any sort whose key omits a feature it then writes) makes obs[blue] != obs[red]. The rule: accumulate INTEGERS, convert once; sort rows by every input they are built from. Accumulating sp[hp channel] += e.hp / HP_SCALE in list order instead costs a 1 ulp seat mismatch on mirrored states (channels own_hp and enemy_hp), measured on both hand-built boards and boards the Rust engine played out. MatchMemory obeys the same rule: integer fine elixir units and integer counts, converted once, per seat.

MOST OF AN OBSERVATION IS STRUCTURALLY CONSTANT, AND THAT DECIDES HOW TO READ IT Across eight boards differing in a unit's position, a destroyed tower and the elixir, 51 of the spatial observation's 11 749 numbers differ. 0.4%. The rest is the arena's static planes, the standing towers and an unchanged hand. So the greatest pairwise cosine between two observations is about 0.9998 with nothing whatever wrong, and a collapse detector that puts a threshold under a cosine is measuring how much of the tensor is static rather than whether it discriminates. measure_variability reports both numbers for a given builder and set of states, so a consumer can calibrate against its own configuration instead of against this paragraph. On the cells that CAN move, those same eight boards sit at 0.984.

PLACEMENT LEGALITY IS THE MASK'S JOB, NOT A CHANNEL'S The action space is Discrete(2305) = no-op + 4 hand slots x 18 x 32 tiles, so the mask ALREADY states per-slot, per-tile legality exactly. own_troop_zone and own_building_zone restated a coarser version of it and cost two of the three PlacementOracle.point_grid calls the observation made per seat per step -- 6.66 calls per env.step before and 2.66 after, both seats, and env.step itself 766 -> 953 per second on the Rust engine; the mask itself is handed to the policy twice instead -- flat as action_mask for the head, and as mask_planes [4, 32, 18] (the mask minus the no-op, reshaped) for a convolutional trunk. enemy_troop_zone stays: it is about the OPPONENT's options and is in no mask.

SPELLS AND STATUS EFFECTS What the engine exposes, and nothing it does not: live spell objects (BattleState.spells: team, card, motion, centre, aim point, flight delay, roll progress, hits) and per-entity stun_ticks / knockback_ticks. Spell objects are visible information in the live game (a Fireball in the air, a Log rolling), so both seats see both teams' spells. The one exception is where an ENEMY spell is aimed: the player chose where to throw their own, but reading the opponent's landing point out of a projectile still in flight is a reveal (Reveal.enemy_spell_aim); own_spell_aim is unconditional. The spatial builder rasterises spells in the own_spells..enemy_stunned channels; the entity-list builder adds status features to each entity row and a separate spells array. Not exposed: a spell's hit radius or damage, which the engine does not export (use the card one-hot); and a unit's buffs and current target, which it does (EntityState.buffs, target_uid) and no builder reads yet. MockEngine resolves spells within a tick and has no status effects, so on it those channels and the spells array are always zero (mock_engine.py WHAT IT IS NOT); tests/test_rust_engine.py checks they carry information on RustEngine battles. KNOCKBACK: knockback_ticks is a per-entity feature only, never a spatial channel. Under the shipped calibration knockback.DURATION_MS = 0 the push is instant and the timer is 0 between ticks, so a channel for it would be constant zero and unverifiable by any coverage guard.

STUN IS ALMOST INVISIBLE AT THE DEFAULT DECISION RATE, and that is a measured
fact rather than a guess. A Zap's stun is 10 ticks; a decision is 10 ticks
(500 ms at TICK_MS 50). Measured on the Rust engine, a Zap cast at tick k of a
decision leaves stun_ticks = k (1, 4, 6, 10 for k = 0, 3, 5, 9) at the ONE
observation that follows, and 0 at every observation after it. So
``own_stunned`` / ``enemy_stunned`` fire for at most one step per Zap, and the
per-entity ``stun_ticks / 100`` feature reads between 0.01 and 0.10 for that one
step. A policy at 500 ms decisions can barely perceive a stun.
The features are left as "is stunned NOW", which is what the engine reports and
what the seat-flip and cell-by-cell tests can check exactly. "Was stunned since
the last observation" would be the feature a policy could actually use, but it
is a different thing -- it depends on the decision rate, not only on the state --
so it is recorded here as an open question rather than taken quietly.
(Measured with Simulator 1 on the 2026-09-21 build.)

Reveal dataclass

Which halves of the hidden state a builder is allowed to read.

All-False (the default) is the fair observation, apart from the invisible units the module doc names. Every True field ADDS channels or vector slots (see the module doc), except enemy_elixir, which swaps the source of the slot that already holds the counted value.

as_dict()

JSON-able, for ClashParallelEnv.config() and checkpoint metadata.

VectorField

Bases: NamedTuple

One run of slots in the flat vector. fair is False for a reveal's slots.

MatchMemory

One seat's memory of the match so far, kept from the states it is shown.

Everything here is derivable by a human watching: which cards the opponent has played and when, where their cycle therefore is, how much elixir they must have, and the player's own cycle, last play and wasted regeneration.

THE ENEMY ELIXIR COUNT is the reason this class exists. It is seeded once, at the start of the match, from the opponent's bar -- the starting amount is public, and a curriculum start that hands one side extra elixir is equally public -- and never read again. From there it is the engine's own arithmetic (protocol.ElixirLaw): pay for each play seen, then regenerate over the ticks that passed, clamped at the cap. Integer fine units throughout, so the count is bit-exact against the bar the engine keeps rather than approximately right (tests/test_env_obs.py plays a battle out and compares the two every step, in both directions).

HOW A PLAY IS SEEN. A play is a public event -- a unit appears, a spell is cast -- and it is read here from the player's hand changing between two observed states, which names the same event and names the CARD exactly. Two consequences, both deliberate: a deck holding the same card twice can hide a play (a real deck is eight distinct cards, and random_deck draws without replacement), and when states are sampled with GAPS rather than every decision, several plays inside one gap are attributed in hand-slot order -- own-frame, so still identical for the two seats of a mirrored battle, but not necessarily the order they happened in.

AND A PRESS. An ability press (a hero's, a champion's) is paid from the bar and changes no hand slot, so it is read from the side's abilities rows instead, which name the same public event: a hero's charge turns spent, or a champion's button goes dark while he still stands (presses_seen). It costs the row's price, dated like a play, and is no play: the cycle, the last play a Mirror copies and the play counts stay as they were. Uncharged, one enemy Golden Knight press left the count 1000 high for the rest of the match (found by the docs session, 2026-09-29).

AND IT SAYS WHEN IT CANNOT BE EXACT. exact goes False, and stays False for the match, the moment the count disagrees with the bar it is modelling -- on EITHER side. Three things make that happen: a play was missed (a deck that repeats a card can hide one, because a play that swaps a card for itself changes no hand slot), a press was missed (read off the rows, a hero or champion pressed and killed between two observations shows only its death; the env hands observe the presses it saw accepted, so its memories cannot miss one), or the engine's elixir law is not the one in calibration.json.

The enemy half of that check is the ONE place this class looks at the opponent's bar, and it does exactly one thing with it: set a boolean. The value is never read into foe_fine and never reaches an observation -- a wrong count is left wrong rather than quietly repaired, because repairing it is the cheat. Checking only the own bar was not enough and the gap was not theoretical: with a normal deck on one side and a repeating deck on the other, the own bar stays perfect while the enemy count drifts, and exact stayed True while the number it certified was wrong. The own bar IS resynced, since it is visible anyway, so it stops drifting further.

IDEMPOTENT BY TICK. observe advances nothing when the state's tick has not moved, so building the same state twice -- which the seat-flip and list-order gates do dozens of times -- cannot drift. A tick that moves BACKWARDS re-seeds, because in a running battle the clock does not go back.

THAT BACKWARDS-TICK RULE IS A BACKSTOP, NOT THE GUARANTEE. BattleState carries no episode identity, so nothing in it distinguishes a new battle that starts at a HIGHER tick -- a Snapshot resume, or a MatchSetup with start_tick past the last episode's end -- from the same battle continuing. The guarantee is ObsBuilder.reset, which the env calls on every episode and which forgets everything. A caller that drives builders itself must call it.

seed(state, team)

Start of a match: everything forgotten, the two bars read once.

start(tick, own_elixir_milli, enemy_elixir_milli, own_hand, next_card, enemy_hand=None)

seed without a state: the start of a match, from what a player sees at it.

Both bars are public at the start, and so is the player's own hand. The enemy hand is only used by observe to notice plays, so a caller that feeds dated plays to advance instead leaves it out.

observe(state, team, presses=None)

Advance to state. A no-op unless the clock moved forward.

presses, where the caller has them, are the ability presses accepted since the last observation, each (team, button): the env's own results, the same public event a player sees. They are charged at the price the button's row showed, and are exact at any gap. Without them the presses are read off the rows (presses_seen), which misses a unit pressed and killed between two observations.

show_own_hand(hand, next_card)

The player's own hand and next card, which the player always sees.

own_cycle is the queue of the player's cards OUTSIDE the hand, next card first: a play appends its card (advance) and a card arriving in the hand pops the head (here). The client refills one empty slot per period, so a played card's slot can hold EMPTY_CARD for a while. Through that window the queue holds one card more, and positions 6-8 stay what they are. With an instant refill the two happen in the same step and the queue keeps four cards, as it always did.

advance(tick, regular_ticks, overtime, own_plays, foe_plays, own_presses=(), foe_presses=())

Move to tick through the plays made since the last tick, each (tick, card), and the ability presses, each (tick, elixir): a press is paid from the bar like a play and is no play (see MatchMemory).

THE ONE PLACE A PLAY CHANGES THIS MEMORY. observe reads plays off hand slots and dates them all at the previous observation; a caller holding a timed log of plays dates each one at its own tick. Both come here, so the counts, the cycle and the leak have one set of formulas.

A play dated inside the interval splits the regeneration at that tick: the bar fills up to it, pays, and fills on. Splitting is exact, because regeneration is never negative, so clamping at a split point and again at the end gives the same bar and the same leak as clamping once. Plays dated at the start of the interval therefore cost the single ElixirLaw.advance call observe always made.

A play the counted bar cannot pay is counted in unaffordable (own, enemy) and the bar floors at zero. The engine refuses such a play, so on an engine's own log this stays 0; anywhere else it means a missed play or a different elixir law.

A MIRROR play is charged the card it copies plus its own elixir: its side's last play that was not a Mirror, which this memory has seen, since plays are public. A Mirror with no earlier play to copy is refused by the engine, so seeing one means a play was missed: it is charged its own elixir and exact goes False.

own_last_play_tick becomes tick, the moment the play is SEEN, not the moment it was made. That is what observe has always recorded, and a policy trained on it reads own_ticks_since_play that way.

enemy_elixir_milli()

The opponent's bar, counted. Exact while exact; an estimate after.

own_elixir_milli()

The same count for the player's own bar; only the law's tests read it.

enemy_possible_hand()

Cards that could be in the enemy hand right now, by the cycle rule.

A card the opponent played goes to the back of their 8-card cycle, so it cannot be in hand again until four more plays have happened: the last four cards they played are exactly the four behind their hand, and everything else is possible. "Everything else" is the whole catalogue until eight distinct cards have been seen, at which point the deck is known and the answer narrows to it. While a played card's slot waits for its refill, the card about to arrive is counted possible a step early: still a superset.

MatchClock

Bases: NamedTuple

Where a match is in time: everything the clock and elixir_rate fields read.

of(state) classmethod

The clock an engine reports.

at(tick, calibration=None) classmethod

The clock of a match still running at tick, from the rules alone.

A match decided at the end of regulation has no later ticks, so one still running there is in overtime. The rate is ElixirLaw.rate_at. Both rules are checked against MockEngine and RustEngine, tick by tick, in tests/test_fair_fields.py.

Variability

Bases: NamedTuple

How much of an observation can actually move. See measure_variability.

ObsBuilder

Bases: ABC

state -> observation dict. Always includes action_mask.

reset(state)

Called at the start of every episode: both seats forget the last one.

see_presses(presses)

The ability presses accepted since the last build, each (team, button), for the memories to charge (MatchMemory.observe); None reads them off the rows.

counts_are_exact(team)

Whether this seat's counted features are still provably right.

False once the opponent's elixir count has been caught disagreeing with the bar it models (MatchMemory). A training run should log it: a policy trained on a match where it went False was reading an estimate in a slot documented as exact, and the whole argument for putting that slot in the fair set is that it is not an estimate.

vector_layout()

The flat vector's fields, in order, for this builder's Reveal.

vector_offsets()

key -> slice into the flat vector this builder writes.

channel_names()

The names of the feature planes / columns this builder writes, in order.

None by default: a builder with no planes or columns need not write this.

spatial_layout()

(channel name, is static) per spatial plane, or empty if there are none.

STATIC means a function of the arena alone: the same numbers every tick, in every battle, for a given seat. A consumer that stores observations can hold those planes once instead of once per transition, and this says which they are rather than leaving it to be inferred by sampling states and hoping none of the others happened to be constant.

config()

Constructor state, JSON-able, for ClashParallelEnv.config().

SpellAimClock

When each enemy spell's target became readable, from what a player sees.

The client never draws an enemy spell's target. A thrown spell (FLIGHT) shows it through its arc once it has flown for a while, so its target counts as seen after_ticks ticks after it STARTS MOVING (not after the throw: a Goblin Barrel sits a while first). A spell seen while it still waits (delay_ticks > 0) starts moving at that tick plus its delay, exactly. One first seen already moving is dated at that sight, which can only show its target later than a player knew it, never sooner. A rolling spell's path is drawn on the ground and an area spell sits on its target, so those count from the first sight.

An engine that reports SpellState.ticks_flown (RoyaleSim 0.1.4 on) is read from that, exactly and per spell object. For an older one the clock below dates each spell by sight: a spell has no id in BattleState, so it is keyed by (team, card, target), which a flight keeps, and same-aim waves (Arrows' three) share one date, the last wave's.

SpatialObsBuilder

Bases: ObsBuilder

Dict(spatial [C, 32, 18], mask_planes [4, 32, 18], vector [V], action_mask [A]).

Entities are rasterised by the tile containing their centre in the own frame (x_own // SUBTILE, clamped). channel_names() lists the channels; spatial_channels(reveal) is the same list with each channel's meaning.

card_identity=True (D2; OFF by default) adds two things, as one switch because they are one change to what a network sees: a card_ids key, uint8 [2, 32, 18], naming the card on each tile (card_id_planes), and enemy_last_card in the vector. It is its OWN key and not a channel of spatial, because spatial is a float Box: a learner's codec stores a float box as a scaled half, a card id comes back as 6.997, and .long() reads card 6 with nothing failing.

THE VOCABULARY SIZE is num_cards + 2 from the LOADED table, and it is published as the card_ids Box's upper bound plus one, so a network sizes its embedding from observation_space["card_ids"].high.max() + 1 at construction.

CATALOGUE IDS ARE POSITIONAL: making one more card loadable renumbers every later id, and an embedding indexed by them would read a different game from the same checkpoint with nothing failing. So config() records the card NAMES in order, and a builder constructed with card_names REFUSES to bind to an engine whose catalogue differs.

config()

Constructor state, including the card names the card_ids ids refer to.

EntityListObsBuilder

Bases: ObsBuilder

Dict(entities [N, F], spells [M, S], mask_planes, vector [V], action_mask [A]).

Per-entity features (F = 18 + num_cards + 1), named by ENTITY_FEATURE_NAMES: 0 present, 1 own, 2 enemy, 3..6 kind one-hot (troop, building, king, princess), 7 x_own / width, 8 y_own / height, 9 hp / max_hp, 10 hp / 1000 (clipped to 1), 11 radius / tile, 12 flying, 13 deploying, 14 deploy_ticks / 100 (clipped), 15 stunned, 16 stun_ticks / 100 (clipped), 17 knockback slide in progress, 18.. card one-hot (last index = crown tower / no card). A unit a spell released (Goblin Barrel's Goblins) carries that spell's card. Rows are sorted canonically by entity_row_key: (enemy, y_own, x_own, kind, card, hp, max_hp, radius, flying, deploy_ticks, stun_ticks, knockback_ticks) -- EVERY entity field a row is built from, so two rows that tie are identical and the order is seat-invariant whatever order the engine listed them in. A shorter key -- (enemy, y_own, x_own, kind, card, hp) -- leaves stacked twins that differ only in deploy_ticks in engine list order. Entities beyond max_entities are dropped in that order; the drop is silent, so size max_entities generously.

Live spell objects, spells [max_spells, S] (S = 14 + num_cards), named by SPELL_FEATURE_NAMES: 0 present, 1 own, 2 enemy, 3..6 motion one-hot (flight, airborne, rolling, area; MOTION_BIT maps every engine motion onto these four: a fuse sets flight, and pulsing, strikes and scheduled set area; a motion code it does not name is refused, on every row and not only the kept ones), 7 x_own / width, 8 y_own / height, 9 and 10 the aim point of YOUR spells (0 on an enemy row), 11 delay_ticks / 100 (clipped; a pulsing area's life left, a fuse's time left), 12 travelled / length (0 when length is 0), 13 hits / 16 (clipped), 14.. card one-hot, and then two APPENDED columns under Reveal.enemy_spell_aim holding the opponent's aim point. Rows past max_spells are dropped in sort order. Positions are clipped to [0, 1] (a roll end point can lie past the arena edge).

THE SORT KEY PUTS THE AIM POINT LAST, AND THAT IS NOT COSMETIC. Row ORDER is
observable: it decides whose delay and hit count appear first. The key must
still name every field a row is built from, so the aim cannot leave it -- but
with the aim ranked early, two enemy spells alike in everything visible and
different in where they were going came out in an order set by where they
were going, and the fair observation changed when only the hidden aim
changed. Measured: two states differing ONLY in two enemy aim points gave
delay columns [0.03, 0.07] and [0.07, 0.03]. With the aim last, hidden data
can only order rows whose every visible field is equal, and those rows write
the same numbers, so their order cannot be seen.

channel_names()

Entity columns, then spell columns; the card one-hots are the trailing run.

vector_layout(num_cards, reveal=None, enemy_last_card=False, evolutions=False, evolution_progress=False)

The flat vector, field by field, in order. THE definition of the layout.

Fair fields come first and in a fixed order, so enabling a reveal never moves a fair feature: the slice holding "own elixir" is the same in a fair run and in a cheating one. tests/test_env_obs.py asserts the resulting width rather than trusting arithmetic done here.

enemy_last_card (D2, 2026-09-22; off by default) appends one FAIR field, the enemy's last play, at the END of the fair block -- after every fair field that existed before it, so turning it on moves no existing fair offset, and before the reveal fields, so the fair block stays contiguous. It is fair because anyone watching sees what the opponent just played.

evolutions (off by default) appends the OWN evolution fields after that, for the same reason: own_hand_evolved and own_next_evolved, and with evolution_progress also own_hand_evo_progress. They are fair because the client shows the charge on a player's own cards. The enemy's charge is not shown, so nothing here reads it.

vector_fields(num_cards, reveal=None, enemy_last_card=False, evolutions=False, evolution_progress=False)

(description, size) per field, in order. The suite checks the sizes add up.

vector_offsets(num_cards, reveal=None, enemy_last_card=False, evolutions=False, evolution_progress=False)

key -> slice into the flat vector, so nothing has to count slots by hand.

cards_that_left(before, after)

Cards played, read from hand slots changing. Slot order: see MatchMemory.

presses_seen(before, after, standing)

The costs of the ability presses one side made between two observations of its abilities rows, button by button: a hero's charge turning spent, or a champion's button going dark while a unit of its card still stands (standing: the card ids of that side's units on the board). A button gone dark with its unit gone is a death, not a press. See MatchMemory.

fair_fields(memory, clock, hand, next_card, own_elixir_milli, cards, max_mana, *, enemy_elixir_milli=None, enemy_last_card=False, hand_costs=None, own_pending=(), evolutions=False, evolution_progress=False, own_evo=())

Every fair vector field but the board's four, by name, from what a player sees.

THE CONTRACT. These are the exact numbers build_vector puts in the env's vector: it calls this function for them. So anything that can keep a MatchMemory without an engine -- dated plays through MatchMemory.advance -- gets the env's fields without building a state. The player supplies what a player sees anyway: the own hand, the next card, the own bar, and the clock (MatchClock.of a state, or MatchClock.at a tick). memory must already have been moved to clock.tick.

enemy_elixir_milli is the true enemy bar for Reveal.enemy_elixir only; left at None, the field is the memory's count, which is the fair one.

hand_costs is what each own hand slot costs right now (PlayerState.hand_costs: a Mirror costs the card it copies plus its own one, -1 when there is nothing to copy). Left at None, each slot is priced at its card's listed elixir, exactly as before it existed. A price p >= 0 is written p / MAX_MANA and is affordable when the own bar holds p; p < 0 keeps the listed elixir and is never affordable. The env passes the engine's prices; RoyaleImitate rebuilds the same ones from its log (from its cc78f2f).

own_pending is the player's OWN commands accepted and not run yet, under a command delay (PlayerState.pending: [kind, what, x, y, ticks_left, cost] rows). A player knows its own taps, so it is fair: a hand slot whose card has a play waiting is flagged and not affordable, and the bar a new play is paid from is the own bar less the waiting cost, as the engine accepts. The own bar itself stays unspent until a command runs, as the client's does. The enemy's waiting commands are never an input: the client shows an opponent's play only when it runs, and so does the enemy count.

own_evo is the player's OWN evolution counters (PlayerState.evo: [card_id, plays, next play evolved, cycle length] rows, one per evolved deck card), read only with evolutions. The client shows the charge on a player's own cards, so it is fair; the enemy's is not shown, and nothing takes it. evolution_progress needs the cycle length, the fourth column, and refuses rows without it.

Keys are FAIR_FIELDS in order, then enemy_last_card and the evolution fields when asked for. Each array is float32 and already clipped to [0, 1], as in the vector. They are views of one buffer, laid out in that order.

build_vector(state, team, cards, max_mana, reveal, memory, enemy_last_card=False, evolutions=False, evolution_progress=False)

The flat vector of vector_layout. memory must already have seen state.

The fields a player's own view decides are fair_fields' (both write them through _write_fair); this adds the four the board decides and any reveal, each at its vector_layout slot in one float32 buffer, and clips once.

measure_variability(builder, states, action_masks, team=0)

What fraction of this builder's observation responds to the battle at all.

WHY A CONSUMER WANTS THIS. A representation that has collapsed -- every board encoding to nearly the same vector -- looks exactly like slow learning, and the usual detector is a cosine between encoded states with a threshold under it. That threshold is meaningless without this number. Measured on the shipped spatial builder over eight boards differing in a unit's position, a destroyed tower and the elixir: 51 of 11 749 cells move, 0.4%, and the greatest pairwise cosine is 0.9998 with nothing whatever wrong. An encoder reporting 0.99 on that input is INCREASING discrimination, not losing it.

So measure the input the same way you measure the encoding, on the same states, and compare the two. A cosine threshold chosen without the input's own cosine is worse than no threshold, because it will fire on a healthy encoder or stay quiet on a dead one depending only on how much of the observation happens to be static.

THE SAME STATES IS A CONSTRAINT, NOT A CONVENIENCE. With 0.4% of cells moving, the baseline is dominated by which states were sampled: boards that differ only in elixir barely move the number, a board with a tower down moves it much more. Two people measuring the SAME encoder against baselines taken on different state sets will disagree about that encoder, and both will be right about what they measured. Take the baseline on the states the encoding was measured on, or the two numbers are not a pair.

WHAT THIS IS NOT FOR: comparing two DIFFERENT builders with each other. The number is built to compare one representation against ITSELF -- an encoding against the input it came from, on the same states, as a ratio. Across builders it is not measuring the same property twice. EntityListObsBuilder is mostly empty canonically-sorted rows where one unit moving can permute a whole row; SpatialObsBuilder is a dense grid of counts where the same unit touches two tiles. Different sparsity, different magnitudes, different response to a small change in the state, so a lower cosine on one may mean it discriminates less or may mean cosine reads a sorted sparse row-set differently from a dense grid, and nothing in the scalar separates those. Measured, for the record: on eight boards the entity-list builder scores 0.9997 on its moving cells against the spatial builder's 0.9841, and that difference is NOT evidence that one is worse.

What would settle it is whether a policy trained on each can tell the boards apart, which is a training question; a cheaper proxy is whether a small probe can recover a known state variable from each, which measures usable information rather than geometric spread. Neither is this function.

The masks are excluded: they are legality, they are handed to the policy separately, and their variability says nothing about the representation.

spatial_channels(reveal=None, evolutions=False, spell_aim=False)

The spatial channels: the fair ones, the evolved-unit pair and the readable enemy spell targets when asked for, then the revealed ones. The fair block keeps its offsets either way.

entity_channels(entities, team, arena)

The first ENTITY_CHANNELS channels, float32 [12, tiles_y, tiles_x], seen by team.

Every cell is an INTEGER sum (counts, raw hp) converted to float32 exactly once, so the result is a function of the entity SET and cannot depend on list order (module doc). hp: int64 / HP_SCALE in float64 (exact for any sum below 2**53), then one rounding to float32. Crown towers are counted in their OWN channel and not in own_buildings / enemy_buildings: a Cannon and a princess tower pose different problems, and one channel holding both could not say which it was. Module-level so a test can plant the old order-dependent float accumulation back in.

spell_channels(state, team, arena)

The SPELL_ROWS, float32 [6, tiles_y, tiles_x], seen by team.

Always all six rows, whatever the Reveal: the builder decides which of them reach the observation, so this stays one function with one layout that a test can plant a defect into. Integer counts converted once, like entity_channels, so the result is a function of the spell and entity SETS.

evolved_channels(entities, team, arena)

float32 [2, tiles_y, tiles_x], seen by team: own, then enemy, evolved units.

Counted on the centre tile, as entity_channels counts troops, from each unit's STATUS_EVOLVED bit. A hero's unit carries STATUS_HERO instead and is not counted. An engine that does not report the bits (status_flags -1) is refused: reading "not reported" as "not evolved" would hand a network zeros that look like an answer. Module-level so a test can plant a defect in it.

seen_aim_plane(state, team, arena, clock)

float32 [tiles_y, tiles_x], seen by team: the ENEMY's spells counted at their landing tile once clock says a player could read it. Module-level so a test can plant a defect.

card_id_planes(entities, team, arena, num_cards)

uint8 [2, tiles_y, tiles_x], seen by team: plane 0 own, plane 1 enemy.

Which CARD occupies each tile, which the float spatial planes cannot say: they count troops and sum hp, so a Giant and a Knight on one tile look alike (docs/observation-spec.md 3b; decision D2). Tiles are the same own-frame centre tiles entity_channels uses, so the two line up cell for cell.

TIES GO TO THE LOWEST UID, and that is not cosmetic. BattleState.entities is not in uid order -- a live battle gives [0, 2, 4, 1, 3, 5] -- so "whichever comes first" would make a plane a function of iteration order, and two runs of one seed could differ. A uid is unique for a whole battle and never reused.

Spells in flight never appear: BattleState.spells is a separate list, so a live spell has no entity and no tile. Units a spell releases do appear, under the releasing spell's catalogue id, because that is what the engine reports for them.

Module-level so a test can plant a defect in it.

entity_row_key(row)

Sort key of an EntityListObsBuilder row: every field but the entity itself. Module-level so a test can plant a shorter six-field key back in.

spell_row_key(row)

Sort key of a spells row: every field but the spell itself (as entity rows).

state.spells is in engine cast order, which is not seat-canonical. The row is built with the AIM POINT LAST, for the reason in EntityListObsBuilder: an enemy spell's aim is hidden, and a key that ranks it before a visible field orders visible content by hidden data.

royalegym.action

Action parsing and action masking.

Discrete(1 + 4 * 18 * 32) = 2305

Index 0 is NO-OP. Index 1 + slot * 576 + ty * 18 + tx plays hand slot slot at the centre of tile (tx, ty), in the ACTING player's own frame (own king at the bottom), so Blue and Red share one policy head. ability_buttons=True adds one action per ability button after these, n_buttons of them (the engine's count): Discrete(2308) on RoyaleSim 2245f9f on.

WHY THIS AND NOT THE ALTERNATIVES * Joint Discrete over (slot x tile) is the only shape in which a mask can say "Giant is legal here but Fireball is not" exactly. Legality is a property of the PAIR: elixir is per card, territory depends on placement type (spells go anywhere, a Goblin Barrel anywhere but water, a Log only where a troop may go but over buildings, buildings never into the pocket, a Miner anywhere on land but not on a building, a Mirror wherever the card it copies may go), footprints depend on card radius. sb3-contrib MaskablePPO applies a MultiDiscrete mask per dimension independently, so a MultiDiscrete([5, 18, 32]) head could only mask the marginals and would happily sample "Knight on the enemy king". An autoregressive factorised head (card, then position conditioned on card) CAN mask exactly, and is the better choice at scale, but it needs a custom policy; the joint Discrete works with stock MaskablePPO today and 2305 logits is small. * Tile resolution, not half-tile. The tilemap is half-tile (36 x 64). Half-tile would be Discrete(9217): 4x the logits, and under placement.TAP_SNAP = client16402_tile_centre the engine snaps every tap to its tile centre, spells too, so half-tile taps land where tile taps do. Measured by the docs session on RoyaleSim 0d0ccd6: a Knight's 992 legal half-tile moves put it on 215 points, against the tile grid's 213; a Fireball's 576 either way. HalfTileActionParser is provided so that trade can be measured instead of argued. * No "hold" or timing sub-action: timing is expressed by choosing NO-OP on a decision step, at decision_ms granularity (see env.py).

HOW THE MASK IS COMPUTED, AND WHY INDEPENDENTLY OF THE ENGINE PlacementOracle recomputes legality from the state snapshot, arena and DeployRules with numpy grids. The engine enforces the same rules on its own code path (Engine.check_deploy). The two are kept deliberately separate: tests compare them exhaustively, so a silently wrong mask -- which trains an agent to want illegal moves, or never to find legal ones -- fails a test instead of a training run.

A point on a half-cell boundary belongs to every cell it touches. A tile
centre touches four half-cells, so a tile is legal iff its whole 2x2 block
is. This closed-cell rule is what makes placement exactly invariant under
the 180-degree seat rotation.

TWO KINDS OF RULE, AS IN THE RUST ENGINE (arena.rs ``deploy_zone``)
CELL rules (water, no-deploy, the river band closed to troops, buildings'
own half) are grids over half-cells, and a point needs every cell it
touches. POINT rules are tested on the point itself: the bodies already on
the board, and the closed NoDeploySize rect of every alive enemy crown tower
(``DeployRules``). A troop's body is judged where the tap resolves -- its tile
centre under placement.TAP_SNAP -- and an own building's, tower's or live
bottle's tile does not refuse a troop tap, which the engine moves off it
(``PlacementOracle.own_tower_zone``), unless the move finds nowhere to go
(``GridActionParser.moved_taps_that_land``). The rect rule equals a cell rule
only while every rect edge lies on a half-cell boundary (the shipped sizes do);
testing the point keeps the mask right if a regenerated cards.json ever breaks
that.

A BUILDING IS ASKED A DIFFERENT QUESTION, AND IT IS NOT "DOES IT FIT HERE"
A building stands on a square of TILES, and a tap that does not fit is not
refused: the game moves the building to the nearest place it does fit. So
for a building card the mask answers "will a tap here build anything",
which stops depending on what is already on the board -- nothing can be in
the way, because being in the way relocates rather than refuses. Only the
cell rules remain. What a tap actually BUILDS, and where, is the engine's
``building_placement``, and the mask is not the place to ask it: an action
space over 2 304 tiles cannot say "here, but two tiles left" anyway.

Measured against the engine on 2026-09-22, tile centres, both seats, a board
with buildings and towers standing: the cell rules alone reproduce
``check_deploy`` for a Cannon on all 576 cells, where the old
body-overlap rule missed 18 of them. Before relocation shipped, those 18 were
right; a building really was refused for touching another body.

PlacementOracle

Placement legality grids from a state snapshot. Engine frame internally.

own_tower_boxes(state, team)

The placement box (engine frame, closed) of every ALIVE crown tower team owns.

own_building_boxes(state, team)

The placement box (engine frame, closed) of every ALIVE building team owns.

own_live_bottles(state, team) staticmethod

The points of the live bottles team owns: spell objects standing out a positive fuse (a Rage's bottle, a Lumberjack's death bottle), on the board from the cast to the release. A zero fuse (the Goblin Curse's area that makes an area) is no bottle; the engine reports it with no fuse left, so it is left out by its delay_ticks. Under spells.SUMMON_FUSE_START = death_bomb_flight, which does not ship, a bottle can stand one tick at zero, and that tick is missed.

own_tower_zone(state, team, px, py)

Where a team troop tap is moved off an own crown tower (the half-open arm), an own building (placement.TROOP_BUILDING_TAPS = as_tower_tap) or an own live bottle's tile (placement.LIVE_BOTTLE_TAPS = client16402_relocate, state.rs relocate_off_own_live_bottle), so no body blocks it.

The core snaps the tap to a one-tile box, floored in the PLACER's frame (placement.SNAP_EVEN_CORNER = placer_frame: in the own frame, then back) or in the arena's (absolute), and moves the troop when that box shares positive area with the placement box of an alive own crown tower or building (state.rs relocate_off_own_crown_tower). The two frames pick different tiles only for a point on a tile edge, which no point the mask asks about is. Broadcasts over px, py.

cell_grid(state, team, placement, troop_laws=True)

The CELL rules. Troop territory here is only 'not the river band'; the enemy tower rects are a point rule, applied in legal_points.

Per placement (the Rust core's Arena::deploy_zone by state.rs deploy_rule): SPELL no cell rule; SPELL_NOT_ON_WATER water only (no no-deploy, no territory: the king block is a legal Goblin Barrel target); TROOP and ROLLING water, no-deploy and the river band; BUILDING water, no-deploy and own half; TUNNEL water only, as SPELL_NOT_ON_WATER (the no-deploy strips and the enemy half are its ground). A MIRROR has no grid of its own: it is placed as the card it copies, which is what to ask with.

enemy_rects(state, team)

The rects a team troop may not be placed in: every ALIVE enemy crown tower's.

legal_points(state, team, card, px, py)

Legality of placing card at engine-frame points (px, py), any shape.

Ignores elixir and hand. A point must lie strictly inside the arena, every half-cell it touches must pass cell_grid, and it must pass the point rules (enemy tower rects for troops and rolling spells, footprints).

points(pitch_div)

Engine-frame subtile coordinates of the candidate points, [ny, nx].

grid_key(state, team, card, pitch_div)

Everything point_grid reads, as a hashable key -- or None, do not cache.

The grid is NOT a function of the card, and that is the whole point. It reads the placement class, the acting team, the pitch, which enemy crown towers are still standing, and the position and radius of every NON-troop entity (troops do not block a deploy). A building adds its own radius to each footprint, so that enters the key too; for a troop the term is zero, which is why one grid serves every troop card in a hand.

That collapses the calls that actually happen. Per env step both seats build a mask over up to four hand slots and an observation over the OPPONENT's troop zone -- and Blue's enemy_troop_zone is the identical grid Red's mask needs. Measured on MockEngine: 2.66 point_grid calls per env.step before this, and the observation's own call was 45% of the time spent building it.

Returns None for a placement whose grid this cannot key safely, so a future rule that reads something else is a cache MISS rather than a stale hit.

point_grid(state, team, card, pitch_div, moves=True)

Engine-frame legality of placing card at each candidate point.

moves=False: as if no tap were moved off an own tower, building or bottle, so a body under the tap refuses it. The parser compares the two to find the taps that are legal only because they are moved (GridActionParser.moved_taps_that_land).

pitch_div=1: tile centres [32, 18]. pitch_div=2: half-cell centres [64, 36]. Ignores elixir and hand; those are applied per slot by the parser.

The same rules as legal_points, specialised to the two regular grids because this is the hot path (4 mask slots + 3 obs zones per seat per env step): a tile centre touches exactly its tile's 2x2 half-cells and a half-cell centre exactly one, and every point rule separates into 1-D x and y tests that broadcast. Measured 2026-09-13: routing this through the general legal_points cost 170/251 us per call (tile/half, Knight, all towers up, MockEngine opening board) and dropped env.step through the full Python stack from about 1 012/s to 753/s on the Rust engine; this specialisation measures 82/91 us on the same board. tests/test_rust_engine.py holds this path equal to legal_points.

MEMOISED on everything it reads (grid_key). The returned array is marked READ-ONLY, because callers share it: writing to it would change what another seat's mask sees. Both shipped callers already copy on their way out -- the mask reshapes into its own buffer, the observation casts to float32 -- so the flag is a guard against a future one, not a change to either.

ActionParser

Bases: ABC

Agent action -> engine command, plus the legality mask for that action space.

action_mask(state, team) abstractmethod

int8 array of shape (space.n,): 1 = the engine will accept this action now.

mask_plane_shape()

Shape the mask MINUS the no-op reshapes to, or None if it does not.

obs.py hands the policy the mask twice: flat for the head, and as mask_planes for a convolutional trunk, which is only meaningful when the action space is a grid. A parser whose space is not one returns None and no mask_planes key appears in the observation.

config()

Constructor state, JSON-able, for ClashParallelEnv.config().

parse(action, state, team) abstractmethod

None means no-op.

GridActionParser

Bases: ActionParser

Discrete(1 + HAND_SIZE * ny * nx) over a regular grid of placement points.

buildings picks what a building tap means; see BUILDING_TAP_ARMS. It changes the ACTION SPACE and not the rules, so two parsers with different arms describe the same battles and a policy trained under one is not comparable to a policy trained under the other without saying so.

mask_plane_shape()

[hand slot, y, x]: the mask without index 0, in the acting seat's own frame.

encode lays the space out as 1 + slot * ny * nx + y * nx + x, so mask[1:].reshape(HAND_SIZE, ny, nx) is that same mask with no arithmetic -- a view, not a copy.

button_of(action)

The ability button action presses, or None for the no-op or a tile action.

moved_taps_that_land(state, team, slot, card, legal)

Own-frame legal without the MOVED troop taps the engine would still refuse.

A troop tap on an own crown tower, an own building or an own live bottle's tile is moved off it (PlacementOracle.own_tower_zone), so the mask offers it where a body would otherwise refuse it. But the move searches a bounded ring of tiles for one that fits, and when none does the tap stays where it was and the body under it refuses it (state.rs ring_nearest_fit). Measured on a board tiled with own Cannons (2026-09-28): 156 taps a seat offered and refused OCCUPIED, each a DeployRefused in a training run. So the engine is asked about exactly those cells, the ones legal only because the tap is moved, as buildable asks it where a building lands: the move is its rule, and a copy here would drift from it.

MEMOISED on what the move reads: the team, the card's placement and kind (every troop of one placement takes the same territory and the same footprint test, so one answer serves a whole hand of them; the Miner's and the Heal's classes are their own), the pitch, every non-troop body (the ring's fit is off every building's box, and off the troop's territory, which no enemy tower changes) and the own live bottles, and the cells asked about. Troops do not enter it. It reads the engine's CURRENT battle, as buildable does.

buildable(state, team, card, legal)

Own-frame grid of the taps this arm offers for card, asked of the engine.

Both arms need the engine, for different reasons. any_tap needs it because relocation does NOT always rescue a tap: when nothing fits within the engine's search the tap is refused, and the cell rules alone cannot tell, so a mask built from them offers actions the engine turns down. Measured on a Cannon lattice filling one half, 15 buildings was enough: the engine refused every tap and the cell-rule mask offered all 960. taps_where_the_building_stays needs it to know where the building would land.

Asks the engine, once per cell, where the building would land. It reads the engine's CURRENT battle, which is the state the caller is masking for; a parser handed a snapshot of some other battle would get a mask for the live one. The env masks the battle it just stepped, so this holds there, and tests/test_building_tap_arms.py grades the mask against the engine rather than assuming it.

MEMOISED, and it has to be. Without a memo this was 58% of env.step in a profile of 400 real steps: one engine call per offered cell, 96 000 of them, because a hand holding a building pays it every step for both seats. The first measurement missed that entirely, because the hand it sampled held no building.

The key is everything relocation reads and nothing else: the acting team, the card's own size, and every body that can block a box, which is the buildings and the ALIVE towers. Not the tick, not elixir, not troops -- a troop cannot block a building (DeployRules), so including it would miss the memo on every step for nothing. Buildings and towers change rarely, so this hits across steps rather than only across the two seats of one step.

TileActionParser

Bases: GridActionParser

The default: Discrete(2305), tile centres (plus n_buttons with ability_buttons).

HalfTileActionParser

Bases: GridActionParser

Discrete(9217), half-cell centres. For measuring whether resolution matters.

landing_tile(arena, snap_even_corner, team, x, y)

The own-frame TILE a building whose centre the engine put at engine-frame (x, y) stands on, for team.

An odd box's centre (a 3x3 Cannon's) is a tile centre, the same tile in any frame. An even box's (a 2x2 Tesla's) is a tile CORNER, and of the four tiles meeting there it stands on the one its snap floored the tap to: in the placer's own frame under placement.SNAP_EVEN_CORNER = placer_frame, in the arena's under absolute. Flooring the corner in the own frame under absolute put every Red Tesla one tile up and right of the tile it was tapped on, and the taps_where_the_building_stays arm offered Red 50 taps to Blue's 142 (measured on the placement batch's build, 2026-09-28).

mask_disagreements(engine, parser, state, team)

Every action where the mask and engine.check_deploy disagree, as (action, mask value, engine status). state must be engine.state().

Exhaustive over the whole action space. This is the check to run against any new engine (the Rust core included) before training on it.

royalegym.state_mutator

State mutators: how each episode begins.

A mutator returns either a MatchSetup (the engine builds a fresh battle from it) or a Snapshot (the engine loads an exact saved state). All randomness comes from the np.random.Generator the env passes in, which is itself seeded from env.reset(seed=...), so a seeded reset is reproducible end to end.

Curriculum is expressed here, not in the engine: start mid-game with a tower already down, start from a scripted defensive board, replay a saved position.

A battle here is described whole (MatchSetup is the complete starting state) and then handed to the engine, rather than edited in place, so mutators compose by choice (WeightedStateMutator picks one per episode) rather than by chaining, and there is no MutatorSequence. Variations on a start are subclasses of DefaultStateMutator that fill in more of the MatchSetup.

Snapshot dataclass

An exact engine state to resume from (Engine.save_state bytes).

StateMutator

Bases: ABC

How each episode starts. Write build(rng, cards), returning a MatchSetup (decks, shuffle, forms) or a Snapshot to start mid-battle. StateSetter is the same class under its older name.

build(rng, cards) abstractmethod

Describe the next episode's starting state.

config()

Constructor state, JSON-able, for ClashParallelEnv.config().

DefaultStateMutator

Bases: StateMutator

A normal battle from tick 0.

decks: fixed [blue, red] decks, each 8 card names (or catalogue ids, or a mix); None draws a random 8-card deck per team (at most one champion, as the ladder allows). Names are looked up in the engine's catalogue at every build, so a deck written by name means the same cards on every card table; an id is only a position in it. mirror: both teams get Blue's deck in the same order (ShuffleMode.MIRRORED) -- the setting for self-play symmetry checks.

MidGameStateMutator

Bases: DefaultStateMutator

Start partway through the match with randomised elixir and tower damage.

Ticks and elixir are given as inclusive integer ranges. tower_down_prob is the chance each princess tower starts destroyed (awarding the crown). Tower HP ranges are fractions in PERCENT of the engine's default max HP, which the mutator does not know -- so it is expressed as tower_hp_percent and converted using max_tower_hp supplied by the caller.

ScriptedBoardStateMutator

Bases: DefaultStateMutator

A fixed board: units already on the field (e.g. a defensive drill).

SnapshotStateMutator

Bases: StateMutator

Resume from saved positions, sampled uniformly (e.g. mined from replays).

WeightedStateMutator

Bases: StateMutator

Curriculum mix: pick a child mutator by weight each episode.

set_weights lets a training loop anneal the mix (e.g. from scripted drills toward full games) without rebuilding the env.

DeckCurriculumStateMutator

Bases: StateMutator

A deck curriculum: one deck you name, on one seat or both, against a pool.

Each episode
  • with probability p the deck is played. Then with probability mirror_p both seats get it in the same order (ShuffleMode.MIRRORED). Otherwise it goes to seat and the other seat draws from pool.
  • with probability 1 - p both seats draw from pool, independently.

seat is "blue", "red", or "either" (a fair coin each episode). With "either" and no mirror, the deck is on each seat in p / 2 of episodes.

pool is a list of decks, drawn uniformly (list a deck twice to weight it), or None for a random deck of eight different cards from the catalogue. Pool draws use shuffle; a mirror always uses ShuffleMode.MIRRORED.

Decks are card NAMES, looked up in the catalogue the env passes to build. A name the catalogue does not have fails the first reset, for every deck in the pool and not only the one drawn, so a typo cannot wait hours for its turn.

config() is exactly the constructor's keyword arguments, JSON-able, so from_config(m.config()) rebuilds a mutator that draws the same episodes from the same seed, and a config file can name the class and pass the dict as its kwargs. set_curriculum changes any of them between episodes without rebuilding the env::

main = ["HogRider", "Musketeer", "Cannon", "Skeletons",
        "Fireball", "Log", "Knight", "Zap"]
m = DeckCurriculumStateMutator(main, p=1.0, mirror_p=1.0)    # mirror matches
env = ClashParallelEnv(engine, state_mutator=m)
...
m.set_curriculum(p=0.9, mirror_p=0.0, pool=[deck_a, deck_b])  # then a pool

set_curriculum changes this object only. Envs built in other processes (an EnvFactory recipe run in a worker) hold their own copy and need the call too: send m.config() and call set_curriculum(**config) there.

from_config(config) classmethod

The mutator config() describes. Also takes the {"class", "params"} record ClashParallelEnv.config() keeps under "state_mutator".

set_curriculum(**changes)

Change any constructor argument, by name, from the next episode on.

All or nothing: a value that is refused leaves the mutator as it was.

random_deck(rng, cards)

Eight different cards of the catalogue, with at most MAX_CHAMPIONS champions.

A draw with more is drawn again, so every draw the rule allows is the deck the same seed dealt before the rule.

deck_ids(names, cards, where='deck')

Card ids for a deck written as card NAMES, looked up in an engine's catalogue.

A card id is a position in one catalogue, and positions move between card tables, so a deck kept as names means the same cards on every engine that has them. A name the catalogue does not have raises ValueError naming it: a typo, or a card this engine does not simulate.

royalegym.done_condition

Done conditions: state -> bool, in one of two roles.

A DoneCondition answers one question per env step: is this episode over? Which kind of over is decided by the slot the env receives it in:

  • termination_cond: the MDP really ended (the value of the next state is 0): a king tower fell, the clock ran out, a curriculum objective was met.
  • truncation_cond: we stopped watching (bootstrap from the next state): a step or tick budget was spent.

Mixing them up biases every value estimate near the cut, so the shipped conditions declare their role by subclassing TerminationCondition or TruncationCondition, and the env refuses a condition passed in the wrong slot. Both are thin subclasses of DoneCondition that add nothing but the name; AnyCondition / AllCondition stay role-neutral so they can combine conditions in either slot.

DoneCondition

Bases: ABC

When an episode ends. Write is_done(state); reset and config are optional.

The env takes one for termination (the battle is decided) and one for truncation (the episode is cut short). TerminalCondition is the same class under its older name.

reset(state)

Called at episode start. Optional.

config()

Constructor state, JSON-able, for ClashParallelEnv.config().

is_done(state) abstractmethod

Whether the episode is over as of state. Called once per env step.

TerminationCondition

Bases: DoneCondition, ABC

A DoneCondition for the termination slot: the episode's outcome is decided.

TruncationCondition

Bases: DoneCondition, ABC

A DoneCondition for the truncation slot: the episode is cut, not decided.

GameOverCondition

Bases: TerminationCondition

The engine declared the battle over (king down, time, sudden death).

StepLimitCondition

Bases: TruncationCondition

Truncate after max_steps env steps (decisions, not ticks).

TickLimitCondition

Bases: TruncationCondition

Truncate once the engine clock reaches max_tick (absolute tick).

FirstCrownCondition

Bases: TerminationCondition

Curriculum: end as soon as any crown changes hands since reset.

A termination, not a truncation: in this curriculum task the crown IS the outcome, so there is nothing to bootstrap from.

AnyCondition

Bases: _CompositeCondition

OR of several conditions. Every child is checked every step (so counters advance).

AllCondition

Bases: _CompositeCondition

AND of several conditions. Every child is checked every step (so counters advance).

The engines

RustEngine is the real one. MockEngine is a pure-Python stand-in that runs the whole API without the Rust build, which is useful before you have built anything and useless for anything about game fidelity.

royalegym.rust_engine

RustEngine: protocol.Engine over the compiled Rust core (../RoyaleSim/crates/royalesim).

WHAT IT IS The adapter that lets ClashParallelEnv and everything else in royalegym run on the real engine with no change to the RL layer. The Rust side is royalesim.Battle (../RoyaleSim/crates/royalesim/src/py.rs). A step is one Rust call (validate, apply, tick N times with the GIL released); the state snapshot is one JSON byte string decoded straight into protocol structs.

CONVENTIONS, AND HOW EACH IS KNOWN RATHER THAN BELIEVED * Positions: protocol.py's engine frame IS the Rust frame (Blue defends low y, origin at Blue's back-left corner, integer subtiles). They pass through _engine_xy unchanged. tests/test_rust_engine.py cross-checks tower centres, passability of every half-cell and deploy legality over a grid against MockEngine, which reads the same arena.json independently. * Tower slots: protocol TowerSlot is named in the OWNER's frame under a 180-degree seat rotation, so Red's own-LEFT princess is the engine's right-lane tower. The table is not typed in: _derive_slot_of_k matches the Rust tower centres to Arena.princess_centers / king_centers (protocol's naming) and refuses to start if they are not a bijection. * Card ids: position in the catalogue (Supercell internal names from data/derived/cards.json, e.g. "Archer" not "Archers"). * Deploy reasons: Rust returns indices into DEPLOY_REASONS, a list of DeployStatus NAMES, so no protocol number is copied into Rust. * Troop territory: the engine and the mask run the same shipped mechanic, the closed NoDeploySize rect of every alive enemy crown tower. The engine reads its sizes from the cards.json in the checkout it was built in, the mask from the one under data_dir(); construction compares Battle.tower_no_deploy_rects() and TERRITORY_MODEL with DeployRules and refuses on any difference (territory_differences). * MatchSetup: protocol.validate_setup runs FIRST in reset, before the core or this adapter touches anything, so a refused setup raises the Protocol's ValueError (same text as MockEngine) and the running battle is untouched. Checking only spawn_violation and the deck ids here would let start_tick -1 or 232, shuffle -1 and tower hp 231 reach PyO3 as an OverflowError, and a malformed tower_hp as an IndexError in the slot mapping -- neither of which MockEngine raises. The core's own checks (py.rs Battle.reset, state.rs scenario_spawn_now) are the authority the protocol rules were written from; tests/test_rust_engine.py and tests/test_parity_hardening.py hold the two together by bypassing this check. * Spells: the catalogue kind code IS the deploy rule, mapped to Placement by _PLACEMENT_OF_KIND (0 TROOP, 1 BUILDING, 2 SPELL, 3 ROLLING, 4 SPELL_NOT_ON_WATER, 5 TUNNEL, 6 MIRROR). A unit a spell releases (the Goblin Barrel's Goblins) is not a card; the core reports it under the releasing spell's catalogue id. state() decodes the core's live spell objects into BattleState.spells and the stun / knockback timers into EntityState.

STALE BUILDS ARE REFUSED The Rust crate compiles calibration.json and arena.json in (include_str!). A calibration edited after the last maturin develop would silently run the old constants, so construction compares every calibration value and the arena with the files on disk and raises on any difference. The same check refuses a Calibration.with_override the compiled engine cannot honour.

THE CARD TABLE IS READ, NOT COMPILED IN Every royalesim.Battle reads data/derived/cards.json from the RoyaleSim checkout the engine was built in, when it is constructed (card.rs load_repo). Re-running tools/extract_cards.py --vintage 2018 --out data/derived/cards.json in that checkout changes the next RustEngine's card table at once: no rebuild, and no stale-build refusal, because there is nothing compiled to be stale. ROYALESIM_DATA_DIR does not move that file; it moves only what this package reads. So build in the checkout whose data you want. Which table an engine got is in its config(): cards_json_fnv1a64 (the hash RoyaleSim's replay fixtures record), cards_vintage and cards_json_hash_source (see RustEngine.card_table_stamp).

WHAT DIFFERS FROM MockEngine ON PURPOSE (engine mechanics, not adapter choices) Cards and towers run at the engine's unified level (card_level, 9 on the 2018 rarity table) where the mock uses CSV level 1; spells travel, roll, stun and knock back here and resolve instantly there; crowns_from_destroyed_towers=False is refused.

RustEngine

Deterministic Rust battle engine behind protocol.Engine.

set_command_delay_ticks(delay)

THE COMMAND DELAY, in ticks, from the next reset on: an int for both seats, or (Blue, Red). A play or press accepted on tick T runs on T + delay, checked again in full then; until it runs the card stays in hand, the bar unspent, and the card or button refuses another command (CARD_PENDING). 0, the default, runs every command at once. The live client's is measured at 21-22 ticks (RoyaleSim df69520).

A delay above 0 needs an engine that has it and can name its refusal: refused here, not at the first refused command.

unit_hitpoints(card_id, level)

Every unit a play of card_id at unified level puts on the board, as the engine's (role, unit name, hitpoints) rows: the card's own row first (role "own", the unit its catalogue row describes), then each unit it puts down, each at the level it takes (roles "second_summon", "spawn", "death_spawn", "release"). The engine's rows as they are (RoyaleSim Battle.unit_hitpoints), so a consumer can tell a unit's hitpoints at the level it was played at, which the catalogue row gives only at the catalogue's level (a Knight: 1766 at 11, 1938 at 12).

ValueError for an unknown card id or a level the card's ladder lacks, naming it. NotImplementedError from an engine build before the rows (RoyaleSim 6909b6f). Not part of protocol.Engine: MockEngine has no levels, so a consumer asks for it with getattr and keeps its own rule where it is absent.

reset(seed, setup)

protocol.validate_setup first -- nothing is read into the core, and no adapter table is indexed, until the whole setup has passed. See the module doc.

building_placement(team, card_name, x, y)

Where a building tapped at ENGINE-frame (x, y) would end up, or None.

(centre_x, centre_y, (x0, y0, x1, y1)), engine frame, subtiles. A tap whose box does not fit is not refused: the engine moves the building to the nearest place it does fit, so the centre this returns is often not the point tapped. None means the tap is refused outright, which a tap outside the arena, on water, on a no-deploy cell or outside the card's territory still is.

This is on the engine because WHERE a building lands is the engine's rule. Working it out here would be a second copy of that rule, and the two would part. action.GridActionParser asks it for the taps_where_the_building_stays arm, and an engine that cannot answer refuses that arm rather than pretending.

config()

Constructor state, JSON-able, for ClashParallelEnv.config().

Everything that makes two RustEngines run different battles from the same commands: which cards are in the catalogue, the level they run at, and which pathfinder and deploy-clamp arms were selected. SymmetricRustEngine is a different engine for this purpose and a checkpoint that does not say so cannot be reproduced -- and the clamp has to be named separately from the pathfinder, because selecting one does not select the other. The card NAMES do not say which card table they were read from, so card_table_stamp is in here too.

engine_identity()

Which engine this is: the data compiled in (build_digest) and the binary loaded in THIS process (engine_binary_sha256). A trace header records both, taken while it records, which is the moment the docstring of engine_binary_digest asks for.

card_table_stamp()

Which card table this engine read, taken when it was constructed.

cards_json_fnv1a64: FNV-1a 64 of the cards.json bytes, 16 hex digits, the value RoyaleSim's replay fixtures record under the same name. cards_vintage: that file's provenance.vintage. cards_json_hash_source: "engine" when the compiled engine reported the hash of the table it loaded (ENGINE_CARD_HASH_NAMES), and then the vintage is "unknown" unless cards_json_path hashes to the same value; "disk_at_construction" when it was hashed from cards_json_path, the file the engine reads, just before and just after the engine read it (a file that changed in between is refused); "unavailable" when that file could not be found, and then the other two are "" and "unknown".

debug_nudge(uid, dx, dy)

TEST-ONLY: move a live entity by (dx, dy) subtiles.

SymmetricRustEngine

Bases: RustEngine

RustEngine under the frame-planned pathfinder (path_search="trace_fitted_astar"), which also selects the fixed-distance knockback, AND the own-frame deploy clamp (ground_y_clamp="deploy_column_range_own_frame"), which it does not (see RustEngine.__init__).

FOR ROTATION-MIRROR GATES ONLY. The shipped search is the game's own, measured on client 16.402, and it is not seat-symmetric: its goal scan and neighbour order run in absolute arena coordinates, so a Red unit and its rotated Blue twin can publish different equal-cost routes (measured: 20 of 54 twin problems on the shipped arena). That is the real game and RustEngine reproduces it. The tests that assert a mirrored battle stays a rotation to the end exist to catch seat bias in the ENGINE'S OTHER SYSTEMS and in this adapter, so they run under this class, whose routes are exact rotations by construction.

THE DEPLOY CLAMP IS THE SAME KIND OF THING AND HAD TO BE ASKED FOR SEPARATELY. formation.GROUND_Y_CLAMP's shipped arm is measured per side and is deliberately not the rotation of itself, so a multi-unit card deployed at mirrored points does not land at mirrored positions -- measured through the engine, Skeletons over the 230 legal mirrored tile centres broke the rotation 22 times under the shipped arm and 0 times under the own-frame one. Until the core exposed the key this class set only path_search, so the rotation gates ran against the asymmetric clamp and correctly reported an asymmetry that is real and intended.

A KEY WITH NO KEYWORD IS SELECTED THROUGH calibration_overrides: SYMMETRIC_ARMS. The first was targeting.FIRST_TOWER_PICK (RoyaleSim 126992a). Its shipped arm reads a summon member's lane in the arena's frame, measured that way on both seats, so a Skeleton Army on the centre line splits one member differently for Blue and for Red. Three rotation gates went red on it at tick 140 while this class's one-tick probe read clean: a lane pick shows only when the members start walking. This class took the old current_x arm until RoyaleSim 1d661b0 added client_spawn_lane_own_frame, the client's lane rule read in the placer's own frame, which it now prefers.

symmetry_problems()

Where this vehicle is NOT a rotation mirror, measured. Empty means it is.

Not raised from __init__ on purpose. Some callers want this class for its determinism rather than its symmetry -- the spawn-ordinal parity gates, for instance -- and refusing to construct would fail six tests whose claim has nothing to do with seats. The check belongs where the claim is made: rotation_divergence calls it before measuring anything, and one test asserts it is empty so the state of the vehicle is visible on its own.

symmetry_report()

symmetry_problems with what to do about it. Empty when there is nothing.

core_import_message(error)

What a reader is told when the engine is not installed: the install page, whose line works before PyPI too, then the page for a source build.

core_available()

Whether the battle engine (royalesim) is installed; CORE_IMPORT_ERROR says why not when it is not.

engine_binary_digest() cached

A short hash of the compiled extension FILE that is loaded, or "unknown".

build_digest identifies the DATA an engine was built with. Nothing identified the CODE: the extension exposes no version, and its distribution version is a constant, so an engine rebuilt from a modified or uncommitted Rust tree was indistinguishable from the one before it. Every guard here, and the ones in the sibling repos, compared data and would have passed.

This does not say which commit built it, which is the engine's to publish. It says whether the binary is the same binary, which is the part that was invisible: two builds from different sources do not produce the same file. Read it beside build_digest -- one moving without the other is the interesting case, and a result recorded against a binary nobody can identify is worth less than one that names it.

Cached: about 4 ms over 2.3 MB, paid once per process. "unknown" rather than raising when the module is not a file on disk, because a missing provenance stamp should not stop a battle.

TAKE IT IN THE PROCESS THAT PRODUCED THE RESULT. This describes the binary THIS process loaded. Fetched later from a second process it describes whatever is on disk by then, which is not the same claim and is the easier one to make by accident: a recording, a picture or a metrics row stamped after the fact says which engine exists now, not which engine made it. Record it beside the result, not beside the report.

build_digest()

A short hash of the data the loaded extension was COMPILED with.

The same hash protocol.calibration_digest takes of calibration.json on disk, over the copy compiled into the extension, plus the arena it was built with. A checkpoint pins this next to the on-disk digest: equal means the policy was trained on an engine built from the data that is there now, and unequal says which way to look without needing the two files side by side. Module-level and also reachable as RustEngine.build_digest so it can be read without constructing an engine -- construction is exactly what refuses on a stale build. The card table is not in it, because it is not compiled in: RustEngine.card_table_stamp says which one an engine read.

stale_build_differences(calibration, arena_path=None)

What differs between the compiled-in data and calibration / arena.json.

engine_cards_json_path()

(path, how it was found) of the cards.json the compiled engine reads.

Not data_dir(): ROYALESIM_DATA_DIR moves what this package reads and never what the engine reads. In order: the path the engine states, if it states one (ENGINE_CARD_PATH_NAMES), found by "engine"; the build directory compiled into the extension, "build checkout"; the sibling checkout of the documented layout, "sibling checkout", a fallback that is right only when the engine was built there.

cards_json_stamp(path)

(FNV-1a 64 of the file's bytes, its provenance.vintage), hashed once per version.

The hash is protocol.fnv1a64 over the raw bytes, the one RoyaleSim's replay fixtures record as cards_json_fnv1a64. Keyed on size and modification time, so a regenerated file is hashed again and an unchanged one is not.

catalogue_vintage_split(rust_cards, mock_cards, engine_vintage=None)

Why the two engines are reading DIFFERENT card tables, or None.

DECIDED BY THE TABLE, NOT BY THE ROWS. The engine's cards.json names the table it is (provenance.vintage, or engine_vintage when given). When that is the table extracted from MockEngine's own pack (RAW_CARD_PACK_TABLE_VINTAGE), the two read the same data, and any row difference is a defect in an engine or in this adapter, which builds the rows: None is returned and the caller's comparison FAILS on it. Until 2026-09-27 a row difference alone meant "different tables", so an adapter bug in one of these fields (a radius scaled in Python) read as a table split and skipped.

The compiled engine reads data/derived/cards.json from the RoyaleSim checkout it was built in, each time an engine is constructed; MockEngine reads the raw CSVs under data_dir(), every time. In a public checkout that is built there, the two are the same vintage by construction: the newer client packs are not redistributed, so the extractor can only produce the tracked table, and any difference here is then a real defect. On a machine that HAS a newer pack and regenerated cards.json from it, the two are simply different tables and every cross-engine comparison is measuring the data rather than the engines.

THE TWO HALVES ARE READ FROM DIFFERENT PLACES, which is the part that catches people out. Pointing ROYALESIM_DATA_DIR at a 2018 data directory moves MockEngine's half and does not move the engine's at all -- measured: with a pure 2018 data dir the extension still reported 95 cards and Goblins at 4, because it went on reading the cards.json of the checkout it was built in. Regenerating THAT file moves the engine's half at once, with no rebuild. So build in the checkout whose data you want. stale_build_differences does not cover this: it compares calibration.json and arena.json, the two files compiled in, and the skip is what handles the catalogue. RustEngine.card_table_stamp names the table an engine actually read.

THAT IS OBSERVED, NOT PREDICTED. A clean clone of the four repos, its data generated from the tracked 2018 tables and royalesim built in it, runs the tests this function guards: 96 passed, 0 skipped, so the two engines agree over the shared card set and the contract holds. On a machine that also holds a private client pack they skip instead, which is this function working rather than the contract going unchecked.

Returned as a ready reason string so a test can skip on it and SAY SO. A skip is not a pass: the comparison that skipped still has to run somewhere.

territory_differences(battle, rules, arena, slot_of_k)

What differs between the compiled engine's troop territory and rules.

Reads the numbers the ENGINE holds (Battle.tower_no_deploy_rects, engine tower k order, mapped to protocol slots through slot_of_k), never the adapter's own copy: a check that compares the Python rules with themselves cannot fail.

check_field_order(core)

Refuse an engine whose positional columns do not line up with this package's.

WHY THIS EXISTS. EntityState, SpellState and ProjectileState are array_like: the engine's state_json sends each row as a JSON ARRAY and it is decoded BY POSITION. A type mismatch raises on its own, but two adjacent fields of one type do not -- on 2026-09-24 the engine gained target_uid and attack_phase, two ints side by side, and a swap between sim's order and this package's would have drawn a uid as an attack phase without a word. The day before, DEPLOY_REASONS had failed exactly this way: a positional list grew in one repo and not the other.

THE RULE. The engine's list must be a PREFIX of ours. Ours may be longer -- a newer royalegym reading an older engine, whose missing trailing columns decode to their "not reported" defaults. The engine's may not be longer, and no shared column may be in a different place.

Returns, per export name, whether that layout was actually checked. An engine that exports no field list passes UNCHECKED, and says so, rather than passing silently.

rotation_probe(engine, tiles=PROBE_TILES)

Deploy each multi-unit card at the same own-frame tile from both seats; report drift.

THE DIRECT MEASUREMENT of the property a rotation gate assumes. Both seats command the same tile in their own frame, so after one tick the two groups must be exact rotations of each other. Anything else means some system under the engine is keyed on the seat or on the absolute arena frame.

Multi-unit cards only, because a single unit lands on the tap and a formation is laid out AROUND it -- which is where a per-side offset shows up. Troops only: a building or a spell does not have a ring.

This exists because the alternative failed twice. Deciding from a list of which keys are asymmetry sources means the list is right only until the engine gains a key, and both times it gained one the gates went red pointing at the wrong thing -- at the observation builder, and at the engine. Measuring the property instead is immune to a key nobody has named yet.

troop_ruled_spells(engine, battle_args, overrides)

The names of the catalogue's kind-3 spells (placement ROLLING) that the core judges by a TROOP's rule: accepted on free ground in the caster's own half, refused OCCUPIED on its own princess tower. Measured on a probe battle built exactly as engine's, both seats, and cached per build, card table, catalogue and constructor arguments. Both seats must agree, or the spell is left as it is and the every-card mask gate names it.

battle_takes(keyword)

Whether the core's Battle constructor takes keyword, from its text signature (pyo3 #[pyo3(signature = ...)]). False without a core or a signature.

tunnelling_cards(engine, battle_args, overrides)

The names of the catalogue's troop and building cards (codes 0 and 1) that the core accepts in the ENEMY half, on free land well inside the enemy's tower rects, where it refuses every other troop and building OUT_OF_TERRITORY: the cards that travel under ground (placement.SPAWN_PATHFIND_TERRITORY). Measured on a probe battle built exactly as engine's, both seats, and cached per build, card table, catalogue and constructor arguments. Both seats must agree, and the same card must be accepted on free land in its own half, or it is left as it is and the every-card gate names it. A card the opening deal never puts in hand is not asked, and keeps its code.

symmetric_overrides(ledger)

The symmetric arm of each SYMMETRIC_ARMS key this compiled ledger ships some OTHER arm of: the most preferred one its candidates list (the first, if it lists none).

A key the ledger lacks is left out, so an engine older than the key still constructs; one already shipping the arm is left out, so nothing is overridden for nothing; one listing none of the symmetric arms is left out, and the rotation probe reports it.

royalegym.mock_engine

MockEngine: a pure-Python stand-in that satisfies protocol.Engine.

WHAT IT IS FOR Building and testing the whole RL layer before the Rust engine exists. It is deliberately simple, but it is not a toy in the three ways that matter to the RL layer: it uses the REAL arena (water, bridges, no-deploy, king block) from data/derived/arena.json, it reads every physics constant it uses from calibration.json (or globals.csv when the registry does not have it yet), and it obeys the three invariants the real engine must obey.

WHAT IT IS NOT A fidelity claim. No unit collision, no projectile travel (ranged hits are instant), no stun/slow, spells hit once, splash is centred on the target, all cards and towers at CSV base level, overtime that ends level is decided by calibration match.OVERTIME_TIEBREAK (the same rule table the Rust engine reads; the rule itself is not yet measured on a recorded match). None of this should be used to argue about how the real game behaves.

ITS SPELL EFFECTS ARE NOT THE ENGINE'S. The Rust core (../RoyaleSim/crates/royalesim
spell.rs) flies Fireball / Arrows / Goblin Barrel to their target, lands and
rolls The Log, stuns with Zap and knocks units back. Here every spell resolves
in the tick it materialises: circle spells and the Log's strip hit once at the
tap, no stun, no knockback, and the Goblin Barrel's Goblins appear at the tap at
once. So ``BattleState.spells`` is always empty and every ``stun_ticks`` /
``knockback_ticks`` is 0. What IS held to the engine is the part the RL layer
must agree on regardless of mechanics: which cards are spells, each spell's
placement class, every deploy verdict and reason code (tests/test_rust_engine.py,
three-way at every half-cell), the catalogue row shape (count/radius/hp 0), and
that a released Goblin reports its barrel's card id. Do not read a spell trade
off a MockEngine battle.

THE INVARIANTS, AS THEY APPLY HERE 1. No floats: integer subtiles, integer elixir with an exact rational scale, truncation-toward-zero division (_tdiv) so that a Red unit's step is exactly the negation of the mirrored Blue unit's step. Python's // floors, and floor(-a/b) != -floor(a/b): using it here breaks the mirror test, which is how that test was checked (see tests/test_env_mock_engine.py). 2. Simultaneity: every phase reads start-of-phase state and writes into buffers (targets, proposed moves, damage, deaths, spawns) applied in one pass. Every tie-break is in the ACTING team's own frame, so nothing prefers Blue. 3. No baked numbers: see MockEngine.__init__.

Pcg32

Bases: Struct

Python port of royalesim::Rng (lib.rs). Same stream for the same seed.

Cross-language equality has NOT been verified against the compiled crate; a parity test for it belongs here once royalesim exposes Rng.

MockEngine

Deterministic, seeded, integer-only. Satisfies protocol.Engine.

config()

Constructor state, JSON-able, for ClashParallelEnv.config().

The card NAMES, because a catalogue subset is what makes one run's card ids mean something different from another's, and card_level for the same reason the Rust adapter reports it. footprint_model because it decides where a building may stand (FOOTPRINT_MODEL). card_table_stamp because the names do not say which numbers were read for them.

card_table_stamp()

Which card table this engine read, as RustEngine.card_table_stamp says it.

cards_vintage: the raw client pack the stats come from (RAW_CARD_PACK). cards_loaded_fnv1a64: FNV-1a 64 (protocol.fnv1a64) of every template built from that pack -- units, spells and cards, msgpack in catalogue order -- so it moves with any number read from the CSVs and not with a column nobody reads. There is no cards_json_fnv1a64: this engine reads no cards.json.

reset(seed, setup)

Validate EVERYTHING (protocol.setup_violation), then build the new battle off to the side, then commit it. A refused setup raises ValueError and the running battle is untouched, as in the Rust core's Battle.reset (which builds a local BattleState and assigns it last).

The order matters. Validating decks and spawns, assigning self._s and only then refusing a destroyed king mid-construction leaves a refused reset with tick 0, the new elixir and hands and zeroed Red towers in place.

load_state(blob)

Restore a snapshot, matching its cards to this catalogue BY NAME.

Refuses (ValueError, running battle untouched) when a hand, a queue, a pending deploy or a non-tower board entity names a card this catalogue does not have -- the Rust catalogue_violation rule and messages. A snapshot from a permuted superset catalogue is re-indexed and plays the same cards. Decoding the blob straight into self._s reads its ids through whatever catalogue loaded them: a Knight snapshot plays as a Valkyrie, and a one-card catalogue crashes.

tiebreak_drain_step(lowest)

The hp every standing crown tower loses this tick, from the lowest crown tower hp of all six read before it (state.rs tiebreak_drain_step).

snapshot_catalogue_violation(s, catalogue)

Why snapshot s cannot run behind a catalogue holding catalogue names.

The Rust py.rs::catalogue_violation rule, in its order and with its messages: every card a battle can still produce -- a hand slot, the cycle queue, a pending deploy, a non-tower entity on the board -- must be in the catalogue. Crown towers are never catalogue cards. Module-level so a test can plant it away.

Recording and watching

royalegym.replay

Battle traces: record, save, load, and VERIFY by re-simulation.

A trace is two things at once
  • a replay in the sense calibration.json means -- seed + initial setup + the command log + periodic state hashes -- which verify_trace re-simulates and checks hash by hash; and
  • enough per-frame entity data (positions, hp, radius) for render.py to draw the battle as a self-contained HTML page without an engine (python -m royalegym.render trace.msgpack -o battle.html).

Encoding is msgspec msgpack (.msgpack, compact) or JSON (.json, for eyeballing). Per-entity rows use array_like structs, so a frame is a list of short arrays; the header names the columns so the viewer reads them by name (render.py refuses a trace that lacks a column it draws). State hashes are hex strings, because a u64 does not survive a JavaScript Number.

SPELLS (2026-09-13). A frame also records the engine's live spell objects (BattleState.spells) and, through EntityState, each entity's stun and knockback timers; the header names the spell columns in spell_fields. Both are trailing fields with defaults, so a version-1 trace recorded before they existed still decodes (with no spells and zero timers) and the version number did not move. A MockEngine trace has no spell objects: its spells resolve inside a tick.

ReplayRecorder

Attach to an env (recorder=). Keeps the most recent trace in trace.

frame_every_tick=True records a frame per engine tick (smooth viewing, and the tightest determinism check); False records one per decision. on_finish is called with each completed trace, e.g. to save it.

SavingReplayRecorder

Bases: ReplayRecorder

A recorder that WRITES completed traces, so a rollout worker can produce them.

WHY THIS EXISTS RATHER THAN on_finish ReplayRecorder already takes on_finish, and for a script that is the better tool. It is unreachable from a training run: the recorder is an EnvFactory COMPONENT, given as an importable class plus JSON kwargs, and a callable is not something JSON can express. So a worker that gets a pickled recipe can be handed this class and a directory, and cannot be handed a function. Every argument here is a string, an int or a bool for that reason.

Asked for by train, who had built and measured the rest: a 3-minute battle is 3,601
frames and 1.6 MB on disk, and their viewer plays one to a real viewer at wall clock
in 57 MB. Watching a live run cannot work -- `time/collection` is 4.35 s of a 62.94 s
iteration, and those 4.35 seconds hold about 82,000 engine ticks, so it is a firehose
between silences rather than slow gameplay. Replaying a saved trace is the form that
works, and the training run producing them was the one missing piece.

WHAT IT COSTS THE RUN One write every every completed episodes, and nothing else. Recording itself is whatever frame_every_tick already costs; this adds serialisation of a finished trace, not per-tick work.

THREE THINGS IT DOES THAT A NAIVE VERSION WOULD NOT, each because a rollout worker is not a script: * The PID is in the name. Workers are separate PROCESSES with separate counters, so two of them would otherwise write the same 000050 file and one would win. * The write is atomic, via a temporary file and os.replace. A viewer reading the directory while a worker writes would otherwise find a half-written trace, and msgpack does not fail loudly on one. * The directory is created and tested at CONSTRUCTION, so a bad path fails when the config is loaded rather than after the first episode of a five-hour run. A recorder that silently saves nothing is the failure train hit from the other side: a publisher that reported "3601 frames, 0 sent" and looked like it worked.

save_trace(trace, path)

Write a recorded battle to path (msgpack, or JSON when it ends in .json) and return the path. royaleviser PATH opens it.

load_trace(path)

Read a battle save_trace wrote (msgpack, or JSON when it ends in .json).

verify_trace(trace, engine)

Re-simulate trace on engine and list every divergence (empty = identical).

Checks the catalogue first, then the deploy statuses of every step, every recorded frame hash, and the final hash. The first divergence is usually the informative one; later ones are consequences.

WHEN THE REPLAY DIVERGES AND THE ENGINE IS ANOTHER BUILD, the first line says so, naming both. It is said only on a divergence: a different binary that replays exactly is ordinary (a CI runner compiles its own), and the note is there so that a list of hash lines is not read as a changed battle when it may be a changed build.

THE CARDS A SETUP DEALS ARE COMPARED BY NAME first. A setup names its cards by POSITION (a deck is a list of catalogue ids), and the catalogue renumbers whenever a card becomes loadable, so the same setup on an engine with a different table deals different cards: the first deploy diverges and the hash list says nothing about why. So when a dealt id names another card on this engine, or none, the result is one line naming it, and nothing is replayed. Only the ids the setup uses are compared: an id no deck or spawn holds changes nothing a replay reads. A trace that starts from a snapshot is not compared, since the snapshot holds the engine's own card indices.

royalegym.viser

The engine side of RoyaleViser: publish battle state over UDP only while a viewer watches.

The viewer's rule: a separate process, never in the tick loop, zero cost when nobody is watching. This module is the whole of what royalegym knows about the viewer; it imports nothing from RoyaleViser (dependency direction stays RoyaleLearn -> RoyaleGym -> RoyaleSim) and nothing graphical.

from royalegym.env import ClashParallelEnv, ClashSelfPlayVecEnv
from royalegym.viser import ViserPublisher
env = ClashParallelEnv(viser=ViserPublisher())        # 127.0.0.1:9870
# or, for self-play, with the ROYALEVISER environment variable set to 127.0.0.1:9870:
vec = ClashSelfPlayVecEnv(8)                          # binds ONE, watches game 0
# then, in another process:  python -m royaleviser --stream 127.0.0.1:9870

WHO READS ROYALEVISER ViserPublisher.from_env(), and only ClashSelfPlayVecEnv calls it. ClashParallelEnv publishes when it is HANDED a publisher and reads no environment variable of its own: a viewer watches ONE battle and has one fixed port, so N envs each building their own from the variable is OSError 10048, which is what self-play with more than one game used to raise. The decision belongs to whoever knows how many battles there are.

PROTOCOL The viewer sends the heartbeat datagram HELLO to (host, port) once a second while it is open. publish costs one clock read while nobody has said hello in ATTACH_TIMEOUT_S: the socket is polled for heartbeats at most once a second (a non-blocking recvfrom). While attached, every call encodes one frame with msgspec msgpack and sends ONE datagram to the last heartbeat's address. A datagram over MAX_DATAGRAM bytes (UDP over IPv4) is resent with every unit's path emptied, then counted in dropped if still too big. Measured 2026-09-20 (MockEngine): 2.6 KB per frame with 12 units + 6 towers, 11 KB with 60 entities, far under the limit.

WIRE FORM frame_dict is royaleviser.model.Frame as a plain dict with that dataclass's field names, so royaleviser.model.decode_frame reads the datagram without a converter and royaleviser.sources.TraceSource draws a trace with the same unit/spell rows. The field list is repeated here on purpose (../RoyaleViser/tests/test_sources.py round-trips a published frame through the viewer's decoder and its contract check, so a drift fails there). Positions stay in engine subtiles; units_per_tile says how many per tile.

ViserPublisher

Sends frames to an attached viewer; detached, publish is one clock read.

port 0 binds an ephemeral port (tests); address is the bound (host, port). publish(state, cards, arena) builds the frame dict, adds the spawn/death lines it derives from the previous published state (only while attached, so a viewer that attaches mid-battle does not see every unit "spawn"), and sends it. publish_dict sends a ready dict (royaleviser.sources.Publisher's path).

from_env() classmethod

A publisher for ROYALEVISER=host:port, or None when the variable is unset/empty.

ROYALEVISER=1 (or true, on, yes) means the default address, HOST:PORT.

ROYALEVISER_RUN names the run these frames come from. The ports are fixed, so two runs on one machine reach the same viewer and it cannot tell whose frames it is drawing beside whose learning panel; the name in every frame is what lets it say. Unset is the empty string, and then no run key is sent at all.

publish(state, cards, arena, decks=None, events=(), meta=None, forms=None)

One datagram for state if a viewer is attached. Returns whether one was sent.

publish_dict(d)

Send one frame dict to the attached viewer (the size policy lives here).

names_of(cards)

card id -> name for one catalogue; unknown ids print as #<id> (never silent).

unit_dict(e, name_of)

royaleviser.model.Unit as a dict. Path is not in an engine state.

footprint is the engine's own box for a building or tower, passed through as it came: None for a troop and for an engine that reports no box, so the viewer draws its marked stand-in rather than a box made up here.

NOT REPORTED IS None, NEVER A VALUE. target, direction and state were hard-coded None until the engine exported them (2026-09-24). An engine that still exports nothing decodes to EntityState's "did not say" defaults, and each maps back to None here: no target, a zero facing, and attack_phase -1. A zero facing in particular must not reach the viewer as [0, 0], because the viewer normalises the direction's length and a zero vector has none.

status is the buff list as the engine names it, names whole: the viewer splits sim's "|"-joined names itself.

projectile_dict(p, name_of)

One projectile in flight, for the viewer. ENGINE frame, subtiles, like units.

name is what fired it: the card, "tower" for a crown tower (firer -1), and None when the engine does not know (firer -2, a projectile restored from a snapshot older than the field). Sim keeps -2 apart from -1 so that "unknown" never reads as "a tower fired this", and mapping every negative id to "tower" would undo exactly that.

player_dict(p, name_of, deck, forms=None)

royaleviser.model.Player: the engine knows everything but the cycle beyond next_card.

An empty hand slot (EMPTY_CARD) is the empty string: known to be empty, not unknown.

The special forms, by NAME: evo is the engine's rows with the card's name for its id; abilities is each button's [name, available, spent, cost], the viewer's four columns. The name is the row's card where the engine states it (a champion's button does), else the k-th hero entry of the deck, else "" where the deck's forms are unknown. ability_cooldowns is its own key, parallel to abilities: each button's cooldown ticks left (0 when ready, and 0 while a champion's ability runs, the button then dark), -1 where the engine's row does not say. Only when some row says, so a frame from an engine without the column is what it always was.

frame_dict(state, name_of, units_per_tile, decks=None, events=(), meta=None, forms=None)

The viewer's Frame for one BattleState (see WIRE FORM).

play_event(tick, team, name, x, y, units_per_tile)

One events-panel line for an accepted deploy: "t120 Blue plays Knight (3.5, 14.5)".

Self-play bookkeeping

royalegym.selfplay

Self-play: opponents, an opponent pool, sampling and Elo bookkeeping.

OPPONENTS An Opponent maps (observation, action mask, rng) -> action. The rng is the env's own generator, so a seeded env with a stochastic opponent is still reproducible (gymnasium's step-determinism check relies on this).

POOL SAMPLING uniform every snapshot equally likely -- broad, protects against forgetting latest always the newest -- pure self-play, fastest to cycle into rock- paper-scissors loops pfsp prioritised fictitious self-play (AlphaStar): weight each snapshot by f(P[learner beats it]); hard f(p)=(1-p)^power focuses on opponents the learner still loses to, variance f(p)=p(1-p) focuses on even matchups. Win probabilities use a Beta(1,1) prior, so an unplayed snapshot reads 0.5 instead of dividing by zero or being ignored forever.

ELO Standard logistic Elo with base-10 / 400 scale. It is bookkeeping for humans and for PFSP diagnostics, not a training signal. Elo assumes transitive strength; a Clash meta is not transitive, so read it with that in mind.

NoopOpponent

Never plays a card. The easiest possible scripted opponent.

RandomLegalOpponent

Uniform over legal actions, taking no-op with probability noop_prob.

Pure uniform over 2305 actions picks no-op ~0% of the time and spams cards the instant they are affordable, which is a strange opponent to learn against.

CallableOpponent

Wrap a frozen policy function fn(obs, mask) -> action.

OpponentPool

Snapshot registry with sampling strategies and Elo/head-to-head tracking.

The learner is an ordinary entry (conventionally id "learner") so its Elo is tracked on the same scale as the snapshots it plays.

record_result(a, b, score_a)

Record one game. score_a is 1 (a won), 0.5 (draw) or 0 (b won).

Returns the new (elo_a, elo_b). The update is zero-sum: the pool's total Elo is conserved, checked by tests/test_selfplay_pool.py::test_the_elo_update_is_zero_sum. It was not checked by anything for the life of this class, while this line said it was -- which is worse than saying nothing, because it stops the next reader looking.

win_rate(a, b)

P[a beats b] with a Beta(1,1) prior.