royalegym¶
The environment layer. This is the package you import to get a battle your bot can play in.
Everything on this page is generated from the docstrings in the code, so it cannot drift from
what the package actually does. If something here is wrong, the docstring is wrong. It covers
the main modules, not every one. royalegym.evaluate, which scores one bot against another,
and royalegym.opponents, the scripted opponents, are not on this page. Read their docstrings
in the code.
If you are looking for where to start rather than what a name does, read the Quick Start Guide and the Overview first.
The environments¶
royalegym.env
¶
Environments: PettingZoo ParallelEnv, Gymnasium single-agent wrapper, vectorised self-play.
DECOMPOSITION ObsBuilder, ActionParser, RewardFunction, StateMutator and two DoneConditions (one in the termination role, one in the truncation role) are constructor arguments. Changing a reward or an observation is a Python edit and never a recompile of the simulator, because that is where most research iteration happens.
DONE FLAGS
Gymnasium's terminated comes from termination_cond (default
GameOverCondition) and truncated from truncation_cond (default
none). Both are consulted every step so counters advance; a termination on
the same step as a truncation is reported as a termination only.
TIMING
One env step = one decision = decision_ms of game time, rounded UP to a
whole number of engine ticks (engine tick length comes from the engine's
state, which reads calibration.json). Both players act simultaneously; the
engine validates both commands against the same pre-step state.
ACTION MASKS
Every observation dict carries action_mask (int8, as pettingzoo's
parallel_api_test and gymnasium Discrete.sample(mask=...) expect), and
mask_planes -- the same mask without the no-op, as [4, 32, 18] -- beside
it. action_masks() returns the bool array sb3-contrib's MaskablePPO calls
for. An action the engine rejects is turned into a no-op and reported in
info["deploy_status"]. The info dict does NOT repeat the mask: it was the
same array the observation already carried, batched a second time by every
vector env for nothing.
EPISODE STATISTICS
The info dict of the LAST step of an episode carries how that episode went --
length, crowns, tower hp, elixir leaked -- so a training run reads it out of
infos["final_info"] instead of keeping a shadow copy of the state.
EPISODE_STAT_KEYS names them. outcome and winner are there too, as
before, but only when the engine ended the battle: a truncation has no winner.
It also carries the reward TERM BY TERM, as ``reward_sum/<TermName>`` summed
over the episode for that seat. That is the number which says which term a
policy is actually chasing, and without it a run logs one scalar and cannot tell
a win from a well-farmed shaping term. Those keys follow the reward function's
own terms rather than a fixed list, which is why they are not in
``EPISODE_STAT_KEYS``. ``log_reward_terms=True`` adds the same breakdown every
STEP as ``reward/<TermName>``, for watching one battle rather than training.
THE VIEWER
ClashParallelEnv publishes only when it is HANDED a publisher, and reads no
environment variable of its own: N envs each binding the viewer's one fixed UDP
port is an OSError, not a feature. ClashSelfPlayVecEnv owns that decision
instead and hands one publisher to one game (see its viser argument).
ClashParallelEnv
¶
Bases: ParallelEnv[str, dict[str, ndarray], int]
A two-seat battle ("blue" and "red"), PettingZoo parallel API.
Each piece is a config object you can swap: the engine, the state mutator (how a battle
starts), the obs builder (what each seat sees), the action parser (what it can do), the
reward function, and the termination and truncation conditions. make_env builds one
with sensible choices for all of them.
reward_terms(team)
¶
This episode's reward, term by term, for one seat.
The number a training run plots and the one that says WHICH term a policy is
actually chasing. It reaches a trainer through the terminal info as
reward_sum/<TermName> -- and through final_info after an autoreset,
which is the only place it survives -- so nothing has to reach into the env
to get it. The keys follow the reward's own terms, so they are not in
EPISODE_STAT_KEYS; the reward function is the authority on the set.
episode_stats(state, team)
¶
How the episode that just ended went, from team's seat.
Written only on the terminal step, so a rollout buffer carries one of these
per EPISODE rather than one per step; EPISODE_STAT_KEYS is the key list.
Tower hp is the MEAN of the three towers' hp fractions (tower_hp_frac),
which is TowerHPReward's potential over three -- so what a run logs and what
its shaping term optimised are the same number. Crowns alone cannot tell a
tower left at 1 hp from a tower never touched, which is why it is here.
config()
¶
What this env IS, as a JSON-able dict, for a checkpoint to record.
Class names and constructor kwargs of every component, the decision
granularity, the Reveal the observation was built with, and digests of
the data underneath: calibration.json as it is on disk, plus the copy the
compiled engine was built with when there is one. A checkpoint that pins
this can say whether a policy is being evaluated on the env it was trained
on -- including whether it was trained with hidden information revealed,
which nothing about a weights file would otherwise show.
A description, not a constructor: EnvFactory is the picklable recipe
that BUILDS one.
action_masks(agent='blue')
¶
Boolean legal-action mask, the shape sb3-contrib MaskablePPO expects.
state()
¶
Global state for centralised critics: Blue's observation, flattened.
Layout is state_space (every non-mask key of the observation Dict, in
the Dict space's key order -- MASK_KEYS are left out). It inherits
Blue's imperfect information: Red's elixir is the count Blue's builder
keeps and Red's hand is not there at all, unless the builder's Reveal
opens them.
ClashGymEnv
¶
Bases: Env[dict[str, ndarray], int]
Single-agent Gymnasium env: one seat is the learner, the other an Opponent.
WHICH ENGINE YOU GET. With no engine= this builds a MockEngine, which is a
readable reference implementation and not the game: different card table, spells
that resolve instantly instead of travelling, no stuns or knockback. That is the
right default for trying the API out, and the wrong thing to train a bot on
believing it is the real one -- so it says so, once, rather than leaving the reader
to find out from a result that does not transfer.
royalegym/ClashRoyaleMock-v0 and royalegym/ClashRoyaleRust-v0 say which they
are in their names and neither warns. Pass engine= here and nothing warns
either: the warning is about not having chosen, not about the mock.
ClashSelfPlayVecEnv
¶
Bases: VectorEnv[Any, Any, Any]
N simultaneous games exposed as 2N agent slots: slot 2i = blue, 2i+1 = red, of game i.
Both seats feed one batch, which is how a single shared policy gets both
players' experience in self-play. Autoreset is SAME_STEP: when a game ends,
the returned observation is already the next game's first one, and the final
observation / info are in infos["final_obs"] / infos["final_info"]
(masked by infos["_final_obs"]). final_info is where an episode's
statistics arrive -- EPISODE_STAT_KEYS.
THE VIEWER, and why it is decided here. A viewer watches ONE battle: it has
one fixed UDP port, and N envs each binding it is OSError 10048, which is
what ClashSelfPlayVecEnv(num_games>1) used to raise the moment ROYALEVISER
was set. So the publisher is bound once, here, and handed to game 0 only.
viser="env" (the default) means ViserPublisher.from_env(): a publisher
when ROYALEVISER=host:port is set, and None -- costing nothing -- when it is
not. Pass a ViserPublisher to bind one explicitly, or None never to
publish.
ADDRESSABLE EPISODES, and why a run cannot resume without them. With
autoreset_seed_fn unset, an autoreset calls reset() with no seed, and
ClashParallelEnv.reset leaves its generator alone when the seed is None --
so the battle a game plays depends on how many battles that game has already
played. A fresh worker cannot arrive at episode 400 without playing 399 first,
which is why a resumed run diverges from the one it is continuing even when
the learner itself came back byte for byte.
Set autoreset_seed_fn and each episode is named instead of counted: the
seed for the nth episode of game g is fn(g, n), so any episode can be
reached directly. episode_ordinals is the counter to checkpoint and
set_episode_ordinals puts it back, after which the stream continues row
for row rather than restarting::
fn = lambda game, n: (run_seed * 1_000_003 + game * 9973 + n) % 2**31
vec = ClashSelfPlayVecEnv(8, autoreset_seed_fn=fn)
...
saved = vec.episode_ordinals # into the checkpoint
vec = ClashSelfPlayVecEnv(8, autoreset_seed_fn=fn)
vec.set_episode_ordinals(saved) # out of it, before reset()
vec.reset()
The default is None and changes nothing: reset(seed=None) is what the
autoreset already did. An explicit reset(seed=...) still wins over the
function, and still consumes an ordinal, so the numbering keeps meaning "the
nth episode this object started in game g" either way.
episode_ordinals
property
¶
Episodes started so far, per game. The number a checkpoint has to carry.
Without it a resumed run can name its episodes and still not know WHICH one to name next, so it starts at 0 and replays a run it has already played.
set_episode_ordinals(ordinals)
¶
Put the counter back where a checkpoint left it, before reset().
The other half of episode_ordinals. Call it before reset(): reset
starts an episode and therefore consumes an ordinal, so setting it after
would skip one.
The five pieces you swap¶
These are the parts you write. Each one is a base class with a working default already shipped, so you can replace one and leave the other four alone.
royalegym.reward
¶
Reward functions: (previous state, state, deploy results) -> scalar, per team.
Composable: CombinedReward([(WinLossReward(), 1.0), (TowerHPReward(), 0.3)]).
A WARNING ABOUT WEIGHTS
A reward coefficient that gets re-tuned every time a new behaviour is
measured is standing in for a missing reward term. If the agent learns to
hoard elixir and you respond by nudging ElixirLeakPenalty up, then down,
then up again as other behaviours shift, the weight is absorbing something
the reward cannot express -- find that thing and give it a term (or remove a
term that is fighting the terminal signal). Weights should settle, not drift.
The terminal win/loss term is the only one that is the actual objective;
everything else is shaping, and shaping that is not a difference of a
potential function can change which policy is optimal.
ZERO-SUM
Every term here is antisymmetric between seats on a mirrored transition
(reward_blue == -reward_red), except ElixirLeakPenalty,
IllegalActionPenalty and PlacementDepthReward, which score only the
acting player's own behaviour. Tests check the antisymmetry, because a self-play
reward that is not zero-sum rewards both players for colluding.
POSITION
DeployResult carries the x and y a command was evaluated at, in the
ENGINE frame. Any term that scores WHERE something was played must convert with
protocol.to_own first, or it rewards Blue and punishes Red for the same
placement and the self-play run learns the seat rather than the game.
PlacementDepthReward is the worked example.
RewardFunction
¶
Bases: ABC
What a bot is paid for. Write get_reward(team, prev, state, results): one number
per seat per step, bigger is better. bind, reset and config are optional.
WinLossReward
¶
CrownReward
¶
TowerHPReward
¶
Bases: RewardFunction
Enemy tower HP destroyed minus own tower HP lost, as fractions of max HP.
This is a difference of a potential (sum of normalised tower HP), so it does not change which policy is optimal for the terminal objective -- it only densifies the signal.
ElixirTradeReward
¶
Bases: RewardFunction
Elixir the enemy spent and lost minus what this seat spent and lost, / MAX-ish scale.
Two things spend elixir and leave nothing behind: a unit that dies, and a spell that is cast. A unit is valued at card elixir / units summoned by that card. Entities that vanish count as killed whether by damage or by lifetime expiry -- a building that times out WAS spent elixir, so counting it as a loss is the honest accounting, not a bug. Crown towers are excluded (TowerHPReward covers them). A unit born and killed inside ONE decision step is in neither snapshot, so it is never priced for either seat; that has always been so and the spell charge below does not change it.
A SPELL IS CHARGED AT THE TAP, from the accepted deploy results, and never from the board. It is consumed the moment it is played, so its elixir is gone whether it killed anything or not, and a Fireball that hits air has to cost four. The board cannot say this: an engine whose spells resolve inside the tick they land reports no spell object at all, and the spell objects that do appear live for several ticks and carry no identity, so counting them once is not possible. Every engine reports an accepted tap exactly once, which is why it is read there. Without this the term paid for a spell's kills and charged nothing for the spell, which teaches that spells are free.
IT IS CHARGED WHAT THE PLAY COST, not the card's own elixir: the price the engine
stated for that hand slot just before the tap (PlayerState.hand_costs), where it
states one. For most cards the two are the same number. A Mirror's are not: it costs
the card it copies plus its own one, and puts down a copy one level up, which matches
no row and so scores nothing, like a Goblin Barrel's goblins. Charged its own one
elixir, a Mirror of a Knight cost the term 1 where the engine took 4.
WHAT THE CATALOGUE DOES NOT PRICE. A card can put units on the board that are not the unit the card itself summons -- a hut and a Witch keep producing them, a Tombstone leaves more behind when it dies, a barrel releases them where it lands. The catalogue has one row per card and no row for any of those units. An engine before RoyaleSim 0.1.5 reports each of them under SOME card that can produce it, which need not be the card its owner played: a Tombstone's skeleton is reported under the Witch, a five-elixir card its owner may not even hold. From 0.1.5 an entity reports the card whose play put it down (the Tombstone's skeletons report the Tombstone). Either way a unit is paid for only when it is the unit its own card's row describes -- same hitpoints, same collision radius, same air or ground -- and anything else a card produced scores nothing. THAT IS AN UNDERSTATEMENT AND IT IS DELIBERATE: a Tombstone's skeleton is worth something, and this term says zero rather than five. Now that an entity says which card produced it, a produced unit could be priced instead; this term does not do that.
WHAT THAT LEAVES TRUE is the property the term actually needs: ONE PLAY OF A CARD IS WORTH EXACTLY THAT CARD'S ELIXIR, charged once, either at the tap or through the units it left behind, never both and never neither. Not every unit a tap puts down matches the row -- a Goblin Gang puts three spear goblins down beside its three goblins, a Rascals puts two girls down beside the boy, and the row describes one kind of each -- but the card's summon count covers exactly the units the row does describe, so the play still totals the card's price. What it costs is precision WITHIN a card: kill a Goblin Gang's spear goblins and the term pays nothing, kill its other three and it pays the whole three elixir. No produced unit matches the row of the card it is reported under. Both are measured in tests/test_rewards.py over every card either engine will place, on both seats, because the whole rule rests on them.
ONE CARD BROKE IT BEFORE ROYALESIM 0.1.5: THE TRI WIZARDS. From RoyaleSim 6909b6f to
0.1.4 their Electro Wizard and Ice Wizard are reported under THEIR OWN card ids (42,
23), not the Tri Wizards', so each matches its own card's row and is priced as that
card: a 7-elixir play totals 14. From 0.1.5 every unit reports the card whose play put
it down, and the play totals 7. tests/test_rewards.py grades the card on its own
(PRICED_ELSEWHERE): an expected failure on an engine with the old labels, a pass on
one with the new.
EXACT ARITHMETIC. Unit values are Fraction(elixir, count) and the sum is
exact until the final division. A running float sum of the same values depends
on entity iteration order, so on a perfectly rotation-mirrored transition (the
same kills on both sides) it returns +1.1e-19 for one seat and -1.1e-19 for the
other instead of 0 for both -- measured on both engines by
tests/test_rust_engine.py's multi-unit rotation test. Harmless to a gradient,
but it makes "equal rewards on a mirror" an uncheckable property.
unit_value(e)
¶
The card's per-unit value, or zero for a unit the catalogue does not price.
ElixirLeakPenalty
¶
Bases: RewardFunction
-1 per decision step spent sitting at full elixir (regeneration wasted).
Not zero-sum: both players can leak at once. max_milli defaults to the
engine's MAX_MANA via calibration when bound through an env.
PlacementDepthReward
¶
Bases: RewardFunction
How far up the board this team's accepted placements were, in own-frame tiles.
The template for every positional shaping term, and the reason
DeployResult carries coordinates. Positive is forward: a placement in the
enemy half scores above one behind your own towers, scaled so a placement at the
far end is 1 and at your own back line is -1. weight on the aggressive side
is the knob; a NEGATIVE weight rewards defending at home.
NOT IN default_reward AT ALL, on purpose -- not at zero weight, not at any
weight. Whether pushing or defending is better is the thing a bot is supposed to
learn, and a shaping term that answers it in advance is the weight-drift this
module's docstring warns about. It is here to be copied and to prove the
coordinates arrive, not to be switched on untested.
test_the_positional_term_stays_out_of_the_default_reward keeps it out.
Antisymmetric between the seats on a mirrored transition, like the terms above: the own frame flips with the seat, so the same placement scores +d for one player and is scored by the other only through its own placements.
IllegalActionPenalty
¶
Bases: RewardFunction
-1 for each of this team's commands the engine rejected this step.
With a correct mask and a masked policy this is always 0; a non-zero value in training logs is a mask bug or an unmasked policy, which is why it exists.
CombinedReward
¶
Bases: RewardFunction
Weighted sum of terms. last_terms keeps each weighted term for logging.
KEYED BY TEAM, because the env calls get_reward once PER SEAT on the same
transition. A single flat breakdown was cleared at the top of every call, so
the first seat's numbers were destroyed by the second and whatever a training
run logged as "the reward breakdown" was only ever Red's -- while the scalar
rewards, which are returned rather than stored, were right for both. Read one
seat's with terms_for(team).
terms_for(team)
¶
The weighted breakdown of team's last reward; empty before the first.
default_reward()
¶
Terminal objective plus light, potential-style shaping.
royalegym.obs
¶
Observation builders: BattleState -> what the policy sees.
PERSPECTIVE Every observation is in the ACTING player's own frame (see protocol.py): Red's board is rotated 180 degrees and its towers/crowns/elixir are "own", Blue's are "enemy". A mirrored battle therefore yields a bit-identical observation for the other seat, which tests assert, so one policy plays both.
FAIR INFORMATION, AND THE Reveal
Everything a builder writes by default is something a player watching the
match could write down: the board, their own hand and cycle, the clock, and a
COUNT of the opponent's elixir kept from the plays they saw and the
regeneration rate everyone knows (MatchMemory). Nothing default-built is
read out of the half of the state a player cannot see, with one known exception:
the builders read no status flags, so a unit invisible to its enemy (a Ghost,
STATUS_INVISIBLE) is still shown to the enemy seat where it stands. Hiding it
or marking it waits on a measurement of what the live client's opponent sees.
``Reveal`` opens that half, one field at a time, for curriculum, distillation
and debugging. An enabled field ADDS its channels and vector slots; it is
never present-but-zero, so a fair observation and a cheating one do not even
have the same width, and a checkpoint cannot quietly be trained on one and
evaluated on the other. ``ClashParallelEnv.config()`` records the Reveal for
exactly that reason. The one field that changes a slot instead of adding one
is ``enemy_elixir``: the fair slot already holds the counted value, and the
reveal swaps in the value read from the state. The two agree on a played-out
battle, presses included (tests/test_env_obs.py, tests/test_press_memory.py); a
disagreement is a bug in the counter, or one of the misses ``MatchMemory`` names and
flags with ``exact``.
Floats appear here and only here-onwards (policy input). They are computed from integer state by the same operations for both seats, so the flip is exact.
NO FLOAT MAY DEPEND ON ENTITY LIST ORDER
state.entities order is engine-private and is NOT seat-canonical: the Rust
engine lists entities by storage slot, and slots are reused, so an exactly
rotation-mirrored battle can list Blue's twins in a different order from Red's.
Float addition is not associative, so any per-entity float accumulation (or any
sort whose key omits a feature it then writes) makes obs[blue] != obs[red]. The
rule: accumulate INTEGERS, convert once; sort rows by every input they are
built from. Accumulating sp[hp channel] += e.hp / HP_SCALE in list order
instead costs a 1 ulp seat mismatch on mirrored states (channels own_hp and
enemy_hp), measured on both hand-built boards and boards the Rust engine
played out. MatchMemory obeys the same rule: integer fine elixir units and
integer counts, converted once, per seat.
MOST OF AN OBSERVATION IS STRUCTURALLY CONSTANT, AND THAT DECIDES HOW TO READ IT
Across eight boards differing in a unit's position, a destroyed tower and the
elixir, 51 of the spatial observation's 11 749 numbers differ. 0.4%. The rest is
the arena's static planes, the standing towers and an unchanged hand. So the
greatest pairwise cosine between two observations is about 0.9998 with nothing
whatever wrong, and a collapse detector that puts a threshold under a cosine is
measuring how much of the tensor is static rather than whether it discriminates.
measure_variability reports both numbers for a given builder and set of
states, so a consumer can calibrate against its own configuration instead of
against this paragraph. On the cells that CAN move, those same eight boards sit
at 0.984.
PLACEMENT LEGALITY IS THE MASK'S JOB, NOT A CHANNEL'S
The action space is Discrete(2305) = no-op + 4 hand slots x 18 x 32 tiles,
so the mask ALREADY states per-slot, per-tile legality exactly. own_troop_zone
and own_building_zone restated a coarser version of it and cost two of the
three PlacementOracle.point_grid calls the observation made per seat per
step -- 6.66 calls per env.step before and 2.66 after, both seats, and
env.step itself 766 -> 953 per second on the Rust engine; the mask itself is
handed to the policy twice instead -- flat as action_mask for the head, and
as mask_planes [4, 32, 18] (the mask minus the no-op, reshaped) for a
convolutional trunk. enemy_troop_zone stays: it is about the OPPONENT's
options and is in no mask.
SPELLS AND STATUS EFFECTS
What the engine exposes, and nothing it does not: live spell objects
(BattleState.spells: team, card, motion, centre, aim point, flight delay, roll
progress, hits) and per-entity stun_ticks / knockback_ticks. Spell
objects are visible information in the live game (a Fireball in the air, a Log
rolling), so both seats see both teams' spells. The one exception is where an
ENEMY spell is aimed: the player chose where to throw their own, but reading
the opponent's landing point out of a projectile still in flight is a reveal
(Reveal.enemy_spell_aim); own_spell_aim is unconditional. The spatial
builder rasterises spells in the own_spells..enemy_stunned channels; the
entity-list builder adds status features to each entity row and a separate
spells array. Not exposed: a spell's hit radius or damage, which the engine
does not export (use the card one-hot); and a unit's buffs and current target,
which it does (EntityState.buffs, target_uid) and no builder reads yet.
MockEngine resolves spells within a tick and has no status effects, so on it
those channels and the spells array are always zero (mock_engine.py WHAT IT
IS NOT); tests/test_rust_engine.py checks they carry information on RustEngine
battles.
KNOCKBACK: knockback_ticks is a per-entity feature only, never a spatial
channel. Under the shipped calibration knockback.DURATION_MS = 0 the push is
instant and the timer is 0 between ticks, so a channel for it would be constant
zero and unverifiable by any coverage guard.
STUN IS ALMOST INVISIBLE AT THE DEFAULT DECISION RATE, and that is a measured
fact rather than a guess. A Zap's stun is 10 ticks; a decision is 10 ticks
(500 ms at TICK_MS 50). Measured on the Rust engine, a Zap cast at tick k of a
decision leaves stun_ticks = k (1, 4, 6, 10 for k = 0, 3, 5, 9) at the ONE
observation that follows, and 0 at every observation after it. So
``own_stunned`` / ``enemy_stunned`` fire for at most one step per Zap, and the
per-entity ``stun_ticks / 100`` feature reads between 0.01 and 0.10 for that one
step. A policy at 500 ms decisions can barely perceive a stun.
The features are left as "is stunned NOW", which is what the engine reports and
what the seat-flip and cell-by-cell tests can check exactly. "Was stunned since
the last observation" would be the feature a policy could actually use, but it
is a different thing -- it depends on the decision rate, not only on the state --
so it is recorded here as an open question rather than taken quietly.
(Measured with Simulator 1 on the 2026-09-21 build.)
Reveal
dataclass
¶
Which halves of the hidden state a builder is allowed to read.
All-False (the default) is the fair observation, apart from the invisible units
the module doc names. Every True field ADDS
channels or vector slots (see the module doc), except enemy_elixir, which
swaps the source of the slot that already holds the counted value.
as_dict()
¶
JSON-able, for ClashParallelEnv.config() and checkpoint metadata.
VectorField
¶
Bases: NamedTuple
One run of slots in the flat vector. fair is False for a reveal's slots.
MatchMemory
¶
One seat's memory of the match so far, kept from the states it is shown.
Everything here is derivable by a human watching: which cards the opponent has played and when, where their cycle therefore is, how much elixir they must have, and the player's own cycle, last play and wasted regeneration.
THE ENEMY ELIXIR COUNT is the reason this class exists. It is seeded once, at
the start of the match, from the opponent's bar -- the starting amount is
public, and a curriculum start that hands one side extra elixir is equally
public -- and never read again. From there it is the engine's own arithmetic
(protocol.ElixirLaw): pay for each play seen, then regenerate over the
ticks that passed, clamped at the cap. Integer fine units throughout, so the
count is bit-exact against the bar the engine keeps rather than approximately
right (tests/test_env_obs.py plays a battle out and compares the two every
step, in both directions).
HOW A PLAY IS SEEN. A play is a public event -- a unit appears, a spell is
cast -- and it is read here from the player's hand changing between two
observed states, which names the same event and names the CARD exactly. Two
consequences, both deliberate: a deck holding the same card twice can hide a
play (a real deck is eight distinct cards, and random_deck draws without
replacement), and when states are sampled with GAPS rather than every decision,
several plays inside one gap are attributed in hand-slot order -- own-frame, so
still identical for the two seats of a mirrored battle, but not necessarily the
order they happened in.
AND A PRESS. An ability press (a hero's, a champion's) is paid from the bar and changes
no hand slot, so it is read from the side's abilities rows instead, which name the
same public event: a hero's charge turns spent, or a champion's button goes dark while
he still stands (presses_seen). It costs the row's price, dated like a play, and
is no play: the cycle, the last play a Mirror copies and the play counts stay as they
were. Uncharged, one enemy Golden Knight press left the count 1000 high for the rest of
the match (found by the docs session, 2026-09-29).
AND IT SAYS WHEN IT CANNOT BE EXACT. exact goes False, and stays False for
the match, the moment the count disagrees with the bar it is modelling -- on
EITHER side. Three things make that happen: a play was missed (a deck that repeats
a card can hide one, because a play that swaps a card for itself changes no hand
slot), a press was missed (read off the rows, a hero or champion pressed and killed
between two observations shows only its death; the env hands observe the presses
it saw accepted, so its memories cannot miss one), or the engine's elixir law is not
the one in calibration.json.
The enemy half of that check is the ONE place this class looks at the
opponent's bar, and it does exactly one thing with it: set a boolean. The value
is never read into foe_fine and never reaches an observation -- a wrong
count is left wrong rather than quietly repaired, because repairing it is the
cheat. Checking only the own bar was not enough and the gap was not theoretical:
with a normal deck on one side and a repeating deck on the other, the own bar
stays perfect while the enemy count drifts, and exact stayed True while the
number it certified was wrong. The own bar IS resynced, since it is visible
anyway, so it stops drifting further.
IDEMPOTENT BY TICK. observe advances nothing when the state's tick has not
moved, so building the same state twice -- which the seat-flip and list-order
gates do dozens of times -- cannot drift. A tick that moves BACKWARDS re-seeds,
because in a running battle the clock does not go back.
THAT BACKWARDS-TICK RULE IS A BACKSTOP, NOT THE GUARANTEE. BattleState
carries no episode identity, so nothing in it distinguishes a new battle that
starts at a HIGHER tick -- a Snapshot resume, or a MatchSetup with
start_tick past the last episode's end -- from the same battle continuing.
The guarantee is ObsBuilder.reset, which the env calls on every episode and
which forgets everything. A caller that drives builders itself must call it.
seed(state, team)
¶
Start of a match: everything forgotten, the two bars read once.
start(tick, own_elixir_milli, enemy_elixir_milli, own_hand, next_card, enemy_hand=None)
¶
seed without a state: the start of a match, from what a player sees at it.
Both bars are public at the start, and so is the player's own hand. The enemy
hand is only used by observe to notice plays, so a caller that feeds dated
plays to advance instead leaves it out.
observe(state, team, presses=None)
¶
Advance to state. A no-op unless the clock moved forward.
presses, where the caller has them, are the ability presses accepted since the
last observation, each (team, button): the env's own results, the same public
event a player sees. They are charged at the price the button's row showed, and are
exact at any gap. Without them the presses are read off the rows (presses_seen),
which misses a unit pressed and killed between two observations.
show_own_hand(hand, next_card)
¶
The player's own hand and next card, which the player always sees.
own_cycle is the queue of the player's cards OUTSIDE the hand, next card first: a
play appends its card (advance) and a card arriving in the hand pops the head
(here). The client refills one empty slot per period, so a played card's slot can hold
EMPTY_CARD for a while. Through that window the queue holds one card more, and
positions 6-8 stay what they are. With an instant refill the two happen in the same
step and the queue keeps four cards, as it always did.
advance(tick, regular_ticks, overtime, own_plays, foe_plays, own_presses=(), foe_presses=())
¶
Move to tick through the plays made since the last tick, each (tick, card),
and the ability presses, each (tick, elixir): a press is paid from the bar like a
play and is no play (see MatchMemory).
THE ONE PLACE A PLAY CHANGES THIS MEMORY. observe reads plays off hand slots
and dates them all at the previous observation; a caller holding a timed log of
plays dates each one at its own tick. Both come here, so the counts, the cycle
and the leak have one set of formulas.
A play dated inside the interval splits the regeneration at that tick: the bar
fills up to it, pays, and fills on. Splitting is exact, because regeneration is
never negative, so clamping at a split point and again at the end gives the same
bar and the same leak as clamping once. Plays dated at the start of the interval
therefore cost the single ElixirLaw.advance call observe always made.
A play the counted bar cannot pay is counted in unaffordable (own, enemy) and
the bar floors at zero. The engine refuses such a play, so on an engine's own log
this stays 0; anywhere else it means a missed play or a different elixir law.
A MIRROR play is charged the card it copies plus its own elixir: its side's last
play that was not a Mirror, which this memory has seen, since plays are public. A
Mirror with no earlier play to copy is refused by the engine, so seeing one means a
play was missed: it is charged its own elixir and exact goes False.
own_last_play_tick becomes tick, the moment the play is SEEN, not the
moment it was made. That is what observe has always recorded, and a policy
trained on it reads own_ticks_since_play that way.
enemy_elixir_milli()
¶
The opponent's bar, counted. Exact while exact; an estimate after.
own_elixir_milli()
¶
The same count for the player's own bar; only the law's tests read it.
enemy_possible_hand()
¶
Cards that could be in the enemy hand right now, by the cycle rule.
A card the opponent played goes to the back of their 8-card cycle, so it cannot be in hand again until four more plays have happened: the last four cards they played are exactly the four behind their hand, and everything else is possible. "Everything else" is the whole catalogue until eight distinct cards have been seen, at which point the deck is known and the answer narrows to it. While a played card's slot waits for its refill, the card about to arrive is counted possible a step early: still a superset.
MatchClock
¶
Bases: NamedTuple
Where a match is in time: everything the clock and elixir_rate fields read.
of(state)
classmethod
¶
The clock an engine reports.
at(tick, calibration=None)
classmethod
¶
The clock of a match still running at tick, from the rules alone.
A match decided at the end of regulation has no later ticks, so one still running
there is in overtime. The rate is ElixirLaw.rate_at. Both rules are checked
against MockEngine and RustEngine, tick by tick, in tests/test_fair_fields.py.
Variability
¶
Bases: NamedTuple
How much of an observation can actually move. See measure_variability.
ObsBuilder
¶
Bases: ABC
state -> observation dict. Always includes action_mask.
reset(state)
¶
Called at the start of every episode: both seats forget the last one.
see_presses(presses)
¶
The ability presses accepted since the last build, each (team, button), for
the memories to charge (MatchMemory.observe); None reads them off the rows.
counts_are_exact(team)
¶
Whether this seat's counted features are still provably right.
False once the opponent's elixir count has been caught disagreeing with the
bar it models (MatchMemory). A training run should log it: a policy
trained on a match where it went False was reading an estimate in a slot
documented as exact, and the whole argument for putting that slot in the
fair set is that it is not an estimate.
vector_layout()
¶
The flat vector's fields, in order, for this builder's Reveal.
vector_offsets()
¶
key -> slice into the flat vector this builder writes.
channel_names()
¶
The names of the feature planes / columns this builder writes, in order.
None by default: a builder with no planes or columns need not write this.
spatial_layout()
¶
(channel name, is static) per spatial plane, or empty if there are none.
STATIC means a function of the arena alone: the same numbers every tick, in every battle, for a given seat. A consumer that stores observations can hold those planes once instead of once per transition, and this says which they are rather than leaving it to be inferred by sampling states and hoping none of the others happened to be constant.
config()
¶
Constructor state, JSON-able, for ClashParallelEnv.config().
SpellAimClock
¶
When each enemy spell's target became readable, from what a player sees.
The client never draws an enemy spell's target. A thrown spell (FLIGHT) shows it through its
arc once it has flown for a while, so its target counts as seen after_ticks ticks after it
STARTS MOVING (not after the throw: a Goblin Barrel sits a while first). A spell seen while it
still waits (delay_ticks > 0) starts moving at that tick plus its delay, exactly. One first
seen already moving is dated at that sight, which can only show its target later than a player
knew it, never sooner. A rolling spell's path is drawn on the ground and an area spell sits on
its target, so those count from the first sight.
An engine that reports SpellState.ticks_flown (RoyaleSim 0.1.4 on) is read from that,
exactly and per spell object. For an older one the clock below dates each spell by sight:
a spell has no id in BattleState, so it is keyed by (team, card, target), which a flight
keeps, and same-aim waves (Arrows' three) share one date, the last wave's.
SpatialObsBuilder
¶
Bases: ObsBuilder
Dict(spatial [C, 32, 18], mask_planes [4, 32, 18], vector [V], action_mask [A]).
Entities are rasterised by the tile containing their centre in the own frame
(x_own // SUBTILE, clamped). channel_names() lists the channels;
spatial_channels(reveal) is the same list with each channel's meaning.
card_identity=True (D2; OFF by default) adds two things, as one switch because
they are one change to what a network sees: a card_ids key, uint8 [2, 32, 18],
naming the card on each tile (card_id_planes), and enemy_last_card in the
vector. It is its OWN key and not a channel of spatial, because spatial is a
float Box: a learner's codec stores a float box as a scaled half, a card id comes
back as 6.997, and .long() reads card 6 with nothing failing.
THE VOCABULARY SIZE is num_cards + 2 from the LOADED table, and it is published
as the card_ids Box's upper bound plus one, so a network sizes its embedding from
observation_space["card_ids"].high.max() + 1 at construction.
CATALOGUE IDS ARE POSITIONAL: making one more card loadable renumbers every later id,
and an embedding indexed by them would read a different game from the same checkpoint
with nothing failing. So config() records the card NAMES in order, and a builder
constructed with card_names REFUSES to bind to an engine whose catalogue differs.
config()
¶
Constructor state, including the card names the card_ids ids refer to.
EntityListObsBuilder
¶
Bases: ObsBuilder
Dict(entities [N, F], spells [M, S], mask_planes, vector [V], action_mask [A]).
Per-entity features (F = 18 + num_cards + 1), named by ENTITY_FEATURE_NAMES:
0 present, 1 own, 2 enemy, 3..6 kind one-hot (troop, building, king, princess),
7 x_own / width, 8 y_own / height, 9 hp / max_hp, 10 hp / 1000 (clipped to 1),
11 radius / tile, 12 flying, 13 deploying, 14 deploy_ticks / 100 (clipped),
15 stunned, 16 stun_ticks / 100 (clipped), 17 knockback slide in progress,
18.. card one-hot (last index = crown tower / no card). A unit a spell
released (Goblin Barrel's Goblins) carries that spell's card.
Rows are sorted canonically by entity_row_key: (enemy, y_own, x_own, kind,
card, hp, max_hp, radius, flying, deploy_ticks, stun_ticks, knockback_ticks) --
EVERY entity field a row is built from, so two rows that tie are identical and
the order is seat-invariant whatever order the engine listed them in. A shorter
key -- (enemy, y_own, x_own, kind, card, hp) -- leaves stacked twins that differ
only in deploy_ticks in engine list order. Entities beyond
max_entities are dropped in that order; the drop is silent, so size
max_entities generously.
Live spell objects, spells [max_spells, S] (S = 14 + num_cards), named by
SPELL_FEATURE_NAMES:
0 present, 1 own, 2 enemy, 3..6 motion one-hot (flight, airborne, rolling,
area; MOTION_BIT maps every engine motion onto these four: a fuse sets
flight, and pulsing, strikes and scheduled set area; a motion code it does not
name is refused, on every row and not only the kept ones), 7 x_own / width,
8 y_own / height, 9 and 10 the aim point of YOUR spells (0 on an enemy row),
11 delay_ticks / 100 (clipped; a pulsing area's life left, a fuse's time left),
12 travelled /
length (0 when length is 0), 13 hits / 16 (clipped), 14.. card one-hot, and
then two APPENDED columns under Reveal.enemy_spell_aim holding the
opponent's aim point. Rows past max_spells are dropped in sort order.
Positions are clipped to [0, 1] (a roll end point can lie past the arena
edge).
THE SORT KEY PUTS THE AIM POINT LAST, AND THAT IS NOT COSMETIC. Row ORDER is
observable: it decides whose delay and hit count appear first. The key must
still name every field a row is built from, so the aim cannot leave it -- but
with the aim ranked early, two enemy spells alike in everything visible and
different in where they were going came out in an order set by where they
were going, and the fair observation changed when only the hidden aim
changed. Measured: two states differing ONLY in two enemy aim points gave
delay columns [0.03, 0.07] and [0.07, 0.03]. With the aim last, hidden data
can only order rows whose every visible field is equal, and those rows write
the same numbers, so their order cannot be seen.
channel_names()
¶
Entity columns, then spell columns; the card one-hots are the trailing run.
vector_layout(num_cards, reveal=None, enemy_last_card=False, evolutions=False, evolution_progress=False)
¶
The flat vector, field by field, in order. THE definition of the layout.
Fair fields come first and in a fixed order, so enabling a reveal never moves a fair feature: the slice holding "own elixir" is the same in a fair run and in a cheating one. tests/test_env_obs.py asserts the resulting width rather than trusting arithmetic done here.
enemy_last_card (D2, 2026-09-22; off by default) appends one FAIR field, the
enemy's last play, at the END of the fair block -- after every fair field that
existed before it, so turning it on moves no existing fair offset, and before the
reveal fields, so the fair block stays contiguous. It is fair because anyone
watching sees what the opponent just played.
evolutions (off by default) appends the OWN evolution fields after that, for the
same reason: own_hand_evolved and own_next_evolved, and with
evolution_progress also own_hand_evo_progress. They are fair because the
client shows the charge on a player's own cards. The enemy's charge is not shown, so
nothing here reads it.
vector_fields(num_cards, reveal=None, enemy_last_card=False, evolutions=False, evolution_progress=False)
¶
(description, size) per field, in order. The suite checks the sizes add up.
vector_offsets(num_cards, reveal=None, enemy_last_card=False, evolutions=False, evolution_progress=False)
¶
key -> slice into the flat vector, so nothing has to count slots by hand.
cards_that_left(before, after)
¶
Cards played, read from hand slots changing. Slot order: see MatchMemory.
presses_seen(before, after, standing)
¶
The costs of the ability presses one side made between two observations of its
abilities rows, button by button: a hero's charge turning spent, or a champion's
button going dark while a unit of its card still stands (standing: the card ids of
that side's units on the board). A button gone dark with its unit gone is a death, not a
press. See MatchMemory.
fair_fields(memory, clock, hand, next_card, own_elixir_milli, cards, max_mana, *, enemy_elixir_milli=None, enemy_last_card=False, hand_costs=None, own_pending=(), evolutions=False, evolution_progress=False, own_evo=())
¶
Every fair vector field but the board's four, by name, from what a player sees.
THE CONTRACT. These are the exact numbers build_vector puts in the env's vector:
it calls this function for them. So anything that can keep a MatchMemory without
an engine -- dated plays through MatchMemory.advance -- gets the env's fields
without building a state. The player supplies what a player sees anyway: the own
hand, the next card, the own bar, and the clock (MatchClock.of a state, or
MatchClock.at a tick). memory must already have been moved to clock.tick.
enemy_elixir_milli is the true enemy bar for Reveal.enemy_elixir only; left
at None, the field is the memory's count, which is the fair one.
hand_costs is what each own hand slot costs right now (PlayerState.hand_costs:
a Mirror costs the card it copies plus its own one, -1 when there is nothing to copy).
Left at None, each slot is priced at its card's listed elixir, exactly as before it
existed. A price p >= 0 is written p / MAX_MANA and is affordable when the own bar
holds p; p < 0 keeps the listed elixir and is never affordable. The env passes the
engine's prices; RoyaleImitate rebuilds the same ones from its log (from its cc78f2f).
own_pending is the player's OWN commands accepted and not run yet, under a command
delay (PlayerState.pending: [kind, what, x, y, ticks_left, cost] rows). A player
knows its own taps, so it is fair: a hand slot whose card has a play waiting is flagged
and not affordable, and the bar a new play is paid from is the own bar less the waiting
cost, as the engine accepts. The own bar itself stays unspent until a command runs, as
the client's does. The enemy's waiting commands are never an input: the client shows
an opponent's play only when it runs, and so does the enemy count.
own_evo is the player's OWN evolution counters (PlayerState.evo: [card_id,
plays, next play evolved, cycle length] rows, one per evolved deck card), read only
with evolutions. The client shows the charge on a player's own cards, so it is
fair; the enemy's is not shown, and nothing takes it. evolution_progress needs
the cycle length, the fourth column, and refuses rows without it.
Keys are FAIR_FIELDS in order, then enemy_last_card and the evolution fields
when asked for. Each array is float32 and already clipped to [0, 1], as in the
vector. They are views of one buffer, laid out in that order.
build_vector(state, team, cards, max_mana, reveal, memory, enemy_last_card=False, evolutions=False, evolution_progress=False)
¶
The flat vector of vector_layout. memory must already have seen state.
The fields a player's own view decides are fair_fields' (both write them through
_write_fair); this adds the four the board decides and any reveal, each at its
vector_layout slot in one float32 buffer, and clips once.
measure_variability(builder, states, action_masks, team=0)
¶
What fraction of this builder's observation responds to the battle at all.
WHY A CONSUMER WANTS THIS. A representation that has collapsed -- every board encoding to nearly the same vector -- looks exactly like slow learning, and the usual detector is a cosine between encoded states with a threshold under it. That threshold is meaningless without this number. Measured on the shipped spatial builder over eight boards differing in a unit's position, a destroyed tower and the elixir: 51 of 11 749 cells move, 0.4%, and the greatest pairwise cosine is 0.9998 with nothing whatever wrong. An encoder reporting 0.99 on that input is INCREASING discrimination, not losing it.
So measure the input the same way you measure the encoding, on the same states, and compare the two. A cosine threshold chosen without the input's own cosine is worse than no threshold, because it will fire on a healthy encoder or stay quiet on a dead one depending only on how much of the observation happens to be static.
THE SAME STATES IS A CONSTRAINT, NOT A CONVENIENCE. With 0.4% of cells moving, the baseline is dominated by which states were sampled: boards that differ only in elixir barely move the number, a board with a tower down moves it much more. Two people measuring the SAME encoder against baselines taken on different state sets will disagree about that encoder, and both will be right about what they measured. Take the baseline on the states the encoding was measured on, or the two numbers are not a pair.
WHAT THIS IS NOT FOR: comparing two DIFFERENT builders with each other. The
number is built to compare one representation against ITSELF -- an encoding
against the input it came from, on the same states, as a ratio. Across builders
it is not measuring the same property twice. EntityListObsBuilder is mostly
empty canonically-sorted rows where one unit moving can permute a whole row;
SpatialObsBuilder is a dense grid of counts where the same unit touches two
tiles. Different sparsity, different magnitudes, different response to a small
change in the state, so a lower cosine on one may mean it discriminates less or
may mean cosine reads a sorted sparse row-set differently from a dense grid, and
nothing in the scalar separates those. Measured, for the record: on eight boards
the entity-list builder scores 0.9997 on its moving cells against the spatial
builder's 0.9841, and that difference is NOT evidence that one is worse.
What would settle it is whether a policy trained on each can tell the boards apart, which is a training question; a cheaper proxy is whether a small probe can recover a known state variable from each, which measures usable information rather than geometric spread. Neither is this function.
The masks are excluded: they are legality, they are handed to the policy separately, and their variability says nothing about the representation.
spatial_channels(reveal=None, evolutions=False, spell_aim=False)
¶
The spatial channels: the fair ones, the evolved-unit pair and the readable enemy spell targets when asked for, then the revealed ones. The fair block keeps its offsets either way.
entity_channels(entities, team, arena)
¶
The first ENTITY_CHANNELS channels, float32 [12, tiles_y, tiles_x], seen by team.
Every cell is an INTEGER sum (counts, raw hp) converted to float32 exactly once,
so the result is a function of the entity SET and cannot depend on list order
(module doc). hp: int64 / HP_SCALE in float64 (exact for any sum below 2**53),
then one rounding to float32. Crown towers are counted in their OWN channel and
not in own_buildings / enemy_buildings: a Cannon and a princess tower
pose different problems, and one channel holding both could not say which it
was. Module-level so a test can plant the old order-dependent float
accumulation back in.
spell_channels(state, team, arena)
¶
The SPELL_ROWS, float32 [6, tiles_y, tiles_x], seen by team.
Always all six rows, whatever the Reveal: the builder decides which of them
reach the observation, so this stays one function with one layout that a test
can plant a defect into. Integer counts converted once, like
entity_channels, so the result is a function of the spell and entity SETS.
evolved_channels(entities, team, arena)
¶
float32 [2, tiles_y, tiles_x], seen by team: own, then enemy, evolved units.
Counted on the centre tile, as entity_channels counts troops, from each unit's
STATUS_EVOLVED bit. A hero's unit carries STATUS_HERO instead and is not
counted. An engine that does not report the bits (status_flags -1) is refused:
reading "not reported" as "not evolved" would hand a network zeros that look like an
answer. Module-level so a test can plant a defect in it.
seen_aim_plane(state, team, arena, clock)
¶
float32 [tiles_y, tiles_x], seen by team: the ENEMY's spells counted at their landing
tile once clock says a player could read it. Module-level so a test can plant a defect.
card_id_planes(entities, team, arena, num_cards)
¶
uint8 [2, tiles_y, tiles_x], seen by team: plane 0 own, plane 1 enemy.
Which CARD occupies each tile, which the float spatial planes cannot say: they
count troops and sum hp, so a Giant and a Knight on one tile look alike
(docs/observation-spec.md 3b; decision D2). Tiles are the same own-frame centre tiles
entity_channels uses, so the two line up cell for cell.
TIES GO TO THE LOWEST UID, and that is not cosmetic. BattleState.entities is not in
uid order -- a live battle gives [0, 2, 4, 1, 3, 5] -- so "whichever comes first" would
make a plane a function of iteration order, and two runs of one seed could differ. A
uid is unique for a whole battle and never reused.
Spells in flight never appear: BattleState.spells is a separate list, so a live
spell has no entity and no tile. Units a spell releases do appear, under the releasing
spell's catalogue id, because that is what the engine reports for them.
Module-level so a test can plant a defect in it.
entity_row_key(row)
¶
Sort key of an EntityListObsBuilder row: every field but the entity itself. Module-level so a test can plant a shorter six-field key back in.
spell_row_key(row)
¶
Sort key of a spells row: every field but the spell itself (as entity rows).
state.spells is in engine cast order, which is not seat-canonical. The row is
built with the AIM POINT LAST, for the reason in EntityListObsBuilder: an
enemy spell's aim is hidden, and a key that ranks it before a visible field
orders visible content by hidden data.
royalegym.action
¶
Action parsing and action masking.
Discrete(1 + 4 * 18 * 32) = 2305
Index 0 is NO-OP. Index 1 + slot * 576 + ty * 18 + tx plays hand slot
slot at the centre of tile (tx, ty), in the ACTING player's own frame
(own king at the bottom), so Blue and Red share one policy head.
ability_buttons=True adds one action per ability button after these,
n_buttons of them (the engine's count): Discrete(2308) on RoyaleSim 2245f9f on.
WHY THIS AND NOT THE ALTERNATIVES
* Joint Discrete over (slot x tile) is the only shape in which a mask can say
"Giant is legal here but Fireball is not" exactly. Legality is a property of
the PAIR: elixir is per card, territory depends on placement type (spells go
anywhere, a Goblin Barrel anywhere but water, a Log only where a troop may go
but over buildings, buildings never into the pocket, a Miner anywhere on land
but not on a building, a Mirror wherever the card it copies may go), footprints
depend on card radius.
sb3-contrib MaskablePPO applies a MultiDiscrete mask per dimension
independently, so a MultiDiscrete([5, 18, 32]) head could only mask the
marginals and would happily sample "Knight on the enemy king". An
autoregressive factorised head (card, then position conditioned on card)
CAN mask exactly, and is the better choice at scale, but it needs a custom
policy; the joint Discrete works with stock MaskablePPO today and 2305
logits is small.
* Tile resolution, not half-tile. The tilemap is half-tile (36 x 64).
Half-tile would be Discrete(9217): 4x the logits, and under placement.TAP_SNAP =
client16402_tile_centre the engine snaps every tap to its tile centre, spells
too, so half-tile taps land where tile taps do. Measured by the docs session on
RoyaleSim 0d0ccd6: a Knight's 992 legal half-tile moves put it on 215 points,
against the tile grid's 213; a Fireball's 576 either way. HalfTileActionParser
is provided so that trade can be measured instead of argued.
* No "hold" or timing sub-action: timing is expressed by choosing NO-OP on
a decision step, at decision_ms granularity (see env.py).
HOW THE MASK IS COMPUTED, AND WHY INDEPENDENTLY OF THE ENGINE
PlacementOracle recomputes legality from the state snapshot, arena and
DeployRules with numpy grids. The engine enforces the same rules on its
own code path (Engine.check_deploy). The two are kept deliberately
separate: tests compare them exhaustively, so a silently wrong mask -- which
trains an agent to want illegal moves, or never to find legal ones -- fails
a test instead of a training run.
A point on a half-cell boundary belongs to every cell it touches. A tile
centre touches four half-cells, so a tile is legal iff its whole 2x2 block
is. This closed-cell rule is what makes placement exactly invariant under
the 180-degree seat rotation.
TWO KINDS OF RULE, AS IN THE RUST ENGINE (arena.rs ``deploy_zone``)
CELL rules (water, no-deploy, the river band closed to troops, buildings'
own half) are grids over half-cells, and a point needs every cell it
touches. POINT rules are tested on the point itself: the bodies already on
the board, and the closed NoDeploySize rect of every alive enemy crown tower
(``DeployRules``). A troop's body is judged where the tap resolves -- its tile
centre under placement.TAP_SNAP -- and an own building's, tower's or live
bottle's tile does not refuse a troop tap, which the engine moves off it
(``PlacementOracle.own_tower_zone``), unless the move finds nowhere to go
(``GridActionParser.moved_taps_that_land``). The rect rule equals a cell rule
only while every rect edge lies on a half-cell boundary (the shipped sizes do);
testing the point keeps the mask right if a regenerated cards.json ever breaks
that.
A BUILDING IS ASKED A DIFFERENT QUESTION, AND IT IS NOT "DOES IT FIT HERE"
A building stands on a square of TILES, and a tap that does not fit is not
refused: the game moves the building to the nearest place it does fit. So
for a building card the mask answers "will a tap here build anything",
which stops depending on what is already on the board -- nothing can be in
the way, because being in the way relocates rather than refuses. Only the
cell rules remain. What a tap actually BUILDS, and where, is the engine's
``building_placement``, and the mask is not the place to ask it: an action
space over 2 304 tiles cannot say "here, but two tiles left" anyway.
Measured against the engine on 2026-09-22, tile centres, both seats, a board
with buildings and towers standing: the cell rules alone reproduce
``check_deploy`` for a Cannon on all 576 cells, where the old
body-overlap rule missed 18 of them. Before relocation shipped, those 18 were
right; a building really was refused for touching another body.
PlacementOracle
¶
Placement legality grids from a state snapshot. Engine frame internally.
own_tower_boxes(state, team)
¶
The placement box (engine frame, closed) of every ALIVE crown tower team owns.
own_building_boxes(state, team)
¶
The placement box (engine frame, closed) of every ALIVE building team owns.
own_live_bottles(state, team)
staticmethod
¶
The points of the live bottles team owns: spell objects standing out a positive
fuse (a Rage's bottle, a Lumberjack's death bottle), on the board from the cast to the
release. A zero fuse (the Goblin Curse's area that makes an area) is no bottle; the
engine reports it with no fuse left, so it is left out by its delay_ticks. Under
spells.SUMMON_FUSE_START = death_bomb_flight, which does not ship, a bottle can stand one
tick at zero, and that tick is missed.
own_tower_zone(state, team, px, py)
¶
Where a team troop tap is moved off an own crown tower (the half-open arm), an
own building (placement.TROOP_BUILDING_TAPS = as_tower_tap) or an own live bottle's tile
(placement.LIVE_BOTTLE_TAPS = client16402_relocate, state.rs
relocate_off_own_live_bottle), so no body blocks it.
The core snaps the tap to a one-tile box, floored in the PLACER's frame (placement.SNAP_EVEN_CORNER = placer_frame: in the own frame, then back) or in the arena's (absolute), and moves the troop when that box shares positive area with the placement box of an alive own crown tower or building (state.rs relocate_off_own_crown_tower). The two frames pick different tiles only for a point on a tile edge, which no point the mask asks about is. Broadcasts over px, py.
cell_grid(state, team, placement, troop_laws=True)
¶
The CELL rules. Troop territory here is only 'not the river band'; the
enemy tower rects are a point rule, applied in legal_points.
Per placement (the Rust core's Arena::deploy_zone by state.rs
deploy_rule): SPELL no cell rule; SPELL_NOT_ON_WATER water only (no
no-deploy, no territory: the king block is a legal Goblin Barrel target);
TROOP and ROLLING water, no-deploy and the river band; BUILDING water,
no-deploy and own half; TUNNEL water only, as SPELL_NOT_ON_WATER (the
no-deploy strips and the enemy half are its ground). A MIRROR has no grid of its
own: it is placed as the card it copies, which is what to ask with.
enemy_rects(state, team)
¶
The rects a team troop may not be placed in: every ALIVE enemy crown tower's.
legal_points(state, team, card, px, py)
¶
Legality of placing card at engine-frame points (px, py), any shape.
Ignores elixir and hand. A point must lie strictly inside the arena, every
half-cell it touches must pass cell_grid, and it must pass the point
rules (enemy tower rects for troops and rolling spells, footprints).
points(pitch_div)
¶
Engine-frame subtile coordinates of the candidate points, [ny, nx].
grid_key(state, team, card, pitch_div)
¶
Everything point_grid reads, as a hashable key -- or None, do not cache.
The grid is NOT a function of the card, and that is the whole point. It reads the placement class, the acting team, the pitch, which enemy crown towers are still standing, and the position and radius of every NON-troop entity (troops do not block a deploy). A building adds its own radius to each footprint, so that enters the key too; for a troop the term is zero, which is why one grid serves every troop card in a hand.
That collapses the calls that actually happen. Per env step both seats build a
mask over up to four hand slots and an observation over the OPPONENT's troop
zone -- and Blue's enemy_troop_zone is the identical grid Red's mask needs.
Measured on MockEngine: 2.66 point_grid calls per env.step before this,
and the observation's own call was 45% of the time spent building it.
Returns None for a placement whose grid this cannot key safely, so a future rule that reads something else is a cache MISS rather than a stale hit.
point_grid(state, team, card, pitch_div, moves=True)
¶
Engine-frame legality of placing card at each candidate point.
moves=False: as if no tap were moved off an own tower, building or bottle, so a
body under the tap refuses it. The parser compares the two to find the taps that are
legal only because they are moved (GridActionParser.moved_taps_that_land).
pitch_div=1: tile centres [32, 18]. pitch_div=2: half-cell centres [64, 36]. Ignores elixir and hand; those are applied per slot by the parser.
The same rules as legal_points, specialised to the two regular grids
because this is the hot path (4 mask slots + 3 obs zones per seat per env
step): a tile centre touches exactly its tile's 2x2 half-cells and a
half-cell centre exactly one, and every point rule separates into 1-D x and
y tests that broadcast. Measured 2026-09-13: routing this through the
general legal_points cost 170/251 us per call (tile/half, Knight, all
towers up, MockEngine opening board) and dropped env.step through the full
Python stack from about 1 012/s to 753/s on the Rust engine; this
specialisation measures 82/91 us on the same board.
tests/test_rust_engine.py holds this path equal to legal_points.
MEMOISED on everything it reads (grid_key). The returned array is marked
READ-ONLY, because callers share it: writing to it would change what another
seat's mask sees. Both shipped callers already copy on their way out -- the
mask reshapes into its own buffer, the observation casts to float32 -- so the
flag is a guard against a future one, not a change to either.
ActionParser
¶
Bases: ABC
Agent action -> engine command, plus the legality mask for that action space.
action_mask(state, team)
abstractmethod
¶
int8 array of shape (space.n,): 1 = the engine will accept this action now.
mask_plane_shape()
¶
Shape the mask MINUS the no-op reshapes to, or None if it does not.
obs.py hands the policy the mask twice: flat for the head, and as
mask_planes for a convolutional trunk, which is only meaningful when
the action space is a grid. A parser whose space is not one returns None
and no mask_planes key appears in the observation.
config()
¶
Constructor state, JSON-able, for ClashParallelEnv.config().
parse(action, state, team)
abstractmethod
¶
None means no-op.
GridActionParser
¶
Bases: ActionParser
Discrete(1 + HAND_SIZE * ny * nx) over a regular grid of placement points.
buildings picks what a building tap means; see BUILDING_TAP_ARMS. It changes
the ACTION SPACE and not the rules, so two parsers with different arms describe the
same battles and a policy trained under one is not comparable to a policy trained
under the other without saying so.
mask_plane_shape()
¶
[hand slot, y, x]: the mask without index 0, in the acting seat's own frame.
encode lays the space out as 1 + slot * ny * nx + y * nx + x, so
mask[1:].reshape(HAND_SIZE, ny, nx) is that same mask with no
arithmetic -- a view, not a copy.
button_of(action)
¶
The ability button action presses, or None for the no-op or a tile action.
moved_taps_that_land(state, team, slot, card, legal)
¶
Own-frame legal without the MOVED troop taps the engine would still refuse.
A troop tap on an own crown tower, an own building or an own live bottle's tile is moved
off it (PlacementOracle.own_tower_zone), so the mask offers it where a body would
otherwise refuse it. But the move searches a bounded ring of tiles for one that fits,
and when none does the tap stays where it was and the body under it refuses it (state.rs
ring_nearest_fit). Measured on a board tiled with own Cannons (2026-09-28): 156 taps a
seat offered and refused OCCUPIED, each a DeployRefused in a training run. So the engine
is asked about exactly those cells, the ones legal only because the tap is moved, as
buildable asks it where a building lands: the move is its rule, and a copy here
would drift from it.
MEMOISED on what the move reads: the team, the card's placement and kind (every troop
of one placement takes the same territory and the same footprint test, so one answer
serves a whole hand of them; the Miner's and the Heal's classes are their own), the
pitch, every non-troop body (the ring's fit is off every building's box, and off the
troop's territory, which no enemy tower changes) and the own live bottles, and the
cells asked about. Troops do not enter it. It reads the engine's CURRENT battle, as
buildable does.
buildable(state, team, card, legal)
¶
Own-frame grid of the taps this arm offers for card, asked of the engine.
Both arms need the engine, for different reasons. any_tap needs it because
relocation does NOT always rescue a tap: when nothing fits within the engine's
search the tap is refused, and the cell rules alone cannot tell, so a mask built
from them offers actions the engine turns down. Measured on a Cannon lattice
filling one half, 15 buildings was enough: the engine refused every tap and the
cell-rule mask offered all 960. taps_where_the_building_stays needs it to
know where the building would land.
Asks the engine, once per cell, where the building would land. It reads the
engine's CURRENT battle, which is the state the caller is masking for; a parser
handed a snapshot of some other battle would get a mask for the live one. The
env masks the battle it just stepped, so this holds there, and
tests/test_building_tap_arms.py grades the mask against the engine rather
than assuming it.
MEMOISED, and it has to be. Without a memo this was 58% of env.step in a
profile of 400 real steps: one engine call per offered cell, 96 000 of them,
because a hand holding a building pays it every step for both seats. The first
measurement missed that entirely, because the hand it sampled held no building.
The key is everything relocation reads and nothing else: the acting team, the
card's own size, and every body that can block a box, which is the buildings and
the ALIVE towers. Not the tick, not elixir, not troops -- a troop cannot block a
building (DeployRules), so including it would miss the memo on every step
for nothing. Buildings and towers change rarely, so this hits across steps rather
than only across the two seats of one step.
TileActionParser
¶
Bases: GridActionParser
The default: Discrete(2305), tile centres (plus n_buttons with ability_buttons).
HalfTileActionParser
¶
Bases: GridActionParser
Discrete(9217), half-cell centres. For measuring whether resolution matters.
landing_tile(arena, snap_even_corner, team, x, y)
¶
The own-frame TILE a building whose centre the engine put at engine-frame (x, y)
stands on, for team.
An odd box's centre (a 3x3 Cannon's) is a tile centre, the same tile in any frame. An
even box's (a 2x2 Tesla's) is a tile CORNER, and of the four tiles meeting there it
stands on the one its snap floored the tap to: in the placer's own frame under
placement.SNAP_EVEN_CORNER = placer_frame, in the arena's under absolute. Flooring the
corner in the own frame under absolute put every Red Tesla one tile up and right of the
tile it was tapped on, and the taps_where_the_building_stays arm offered Red 50 taps
to Blue's 142 (measured on the placement batch's build, 2026-09-28).
mask_disagreements(engine, parser, state, team)
¶
Every action where the mask and engine.check_deploy disagree, as
(action, mask value, engine status). state must be engine.state().
Exhaustive over the whole action space. This is the check to run against any new engine (the Rust core included) before training on it.
royalegym.state_mutator
¶
State mutators: how each episode begins.
A mutator returns either a MatchSetup (the engine builds a fresh battle from
it) or a Snapshot (the engine loads an exact saved state). All randomness
comes from the np.random.Generator the env passes in, which is itself seeded
from env.reset(seed=...), so a seeded reset is reproducible end to end.
Curriculum is expressed here, not in the engine: start mid-game with a tower already down, start from a scripted defensive board, replay a saved position.
A battle here is described whole (MatchSetup is the complete starting state) and then handed
to the engine, rather than edited in place, so mutators compose by choice
(WeightedStateMutator picks one per episode) rather than by chaining, and
there is no MutatorSequence. Variations on a start are subclasses of
DefaultStateMutator that fill in more of the MatchSetup.
Snapshot
dataclass
¶
An exact engine state to resume from (Engine.save_state bytes).
StateMutator
¶
Bases: ABC
How each episode starts. Write build(rng, cards), returning a MatchSetup (decks,
shuffle, forms) or a Snapshot to start mid-battle. StateSetter is the same class
under its older name.
DefaultStateMutator
¶
Bases: StateMutator
A normal battle from tick 0.
decks: fixed [blue, red] decks, each 8 card names (or catalogue ids, or a mix);
None draws a random 8-card deck per team (at most one champion, as the ladder allows).
Names are looked up in the engine's
catalogue at every build, so a deck written by name means the same cards on every
card table; an id is only a position in it.
mirror: both teams get Blue's deck in the same order (ShuffleMode.MIRRORED)
-- the setting for self-play symmetry checks.
MidGameStateMutator
¶
Bases: DefaultStateMutator
Start partway through the match with randomised elixir and tower damage.
Ticks and elixir are given as inclusive integer ranges. tower_down_prob
is the chance each princess tower starts destroyed (awarding the crown).
Tower HP ranges are fractions in PERCENT of the engine's default max HP, which
the mutator does not know -- so it is expressed as tower_hp_percent and
converted using max_tower_hp supplied by the caller.
ScriptedBoardStateMutator
¶
SnapshotStateMutator
¶
WeightedStateMutator
¶
Bases: StateMutator
Curriculum mix: pick a child mutator by weight each episode.
set_weights lets a training loop anneal the mix (e.g. from scripted drills
toward full games) without rebuilding the env.
DeckCurriculumStateMutator
¶
Bases: StateMutator
A deck curriculum: one deck you name, on one seat or both, against a pool.
Each episode
- with probability
pthe deck is played. Then with probabilitymirror_pboth seats get it in the same order (ShuffleMode.MIRRORED). Otherwise it goes toseatand the other seat draws frompool. - with probability
1 - pboth seats draw frompool, independently.
seat is "blue", "red", or "either" (a fair coin each episode). With "either"
and no mirror, the deck is on each seat in p / 2 of episodes.
pool is a list of decks, drawn uniformly (list a deck twice to weight it), or
None for a random deck of eight different cards from the catalogue. Pool draws use
shuffle; a mirror always uses ShuffleMode.MIRRORED.
Decks are card NAMES, looked up in the catalogue the env passes to build. A name
the catalogue does not have fails the first reset, for every deck in the pool and
not only the one drawn, so a typo cannot wait hours for its turn.
config() is exactly the constructor's keyword arguments, JSON-able, so
from_config(m.config()) rebuilds a mutator that draws the same episodes from the
same seed, and a config file can name the class and pass the dict as its kwargs.
set_curriculum changes any of them between episodes without rebuilding the env::
main = ["HogRider", "Musketeer", "Cannon", "Skeletons",
"Fireball", "Log", "Knight", "Zap"]
m = DeckCurriculumStateMutator(main, p=1.0, mirror_p=1.0) # mirror matches
env = ClashParallelEnv(engine, state_mutator=m)
...
m.set_curriculum(p=0.9, mirror_p=0.0, pool=[deck_a, deck_b]) # then a pool
set_curriculum changes this object only. Envs built in other processes (an
EnvFactory recipe run in a worker) hold their own copy and need the call too:
send m.config() and call set_curriculum(**config) there.
from_config(config)
classmethod
¶
The mutator config() describes. Also takes the {"class", "params"}
record ClashParallelEnv.config() keeps under "state_mutator".
set_curriculum(**changes)
¶
Change any constructor argument, by name, from the next episode on.
All or nothing: a value that is refused leaves the mutator as it was.
random_deck(rng, cards)
¶
Eight different cards of the catalogue, with at most MAX_CHAMPIONS champions.
A draw with more is drawn again, so every draw the rule allows is the deck the same seed dealt before the rule.
deck_ids(names, cards, where='deck')
¶
Card ids for a deck written as card NAMES, looked up in an engine's catalogue.
A card id is a position in one catalogue, and positions move between card tables, so a deck kept as names means the same cards on every engine that has them. A name the catalogue does not have raises ValueError naming it: a typo, or a card this engine does not simulate.
royalegym.done_condition
¶
Done conditions: state -> bool, in one of two roles.
A DoneCondition answers one question per env step: is this episode over?
Which kind of over is decided by the slot the env receives it in:
termination_cond: the MDP really ended (the value of the next state is 0): a king tower fell, the clock ran out, a curriculum objective was met.truncation_cond: we stopped watching (bootstrap from the next state): a step or tick budget was spent.
Mixing them up biases every value estimate near the cut, so the shipped
conditions declare their role by subclassing TerminationCondition or
TruncationCondition, and the env refuses a condition passed in the wrong
slot. Both are thin subclasses of DoneCondition that add nothing but the
name; AnyCondition / AllCondition stay role-neutral so they can combine
conditions in either slot.
DoneCondition
¶
Bases: ABC
When an episode ends. Write is_done(state); reset and config are optional.
The env takes one for termination (the battle is decided) and one for truncation (the
episode is cut short). TerminalCondition is the same class under its older name.
TerminationCondition
¶
Bases: DoneCondition, ABC
A DoneCondition for the termination slot: the episode's outcome is decided.
TruncationCondition
¶
GameOverCondition
¶
StepLimitCondition
¶
TickLimitCondition
¶
FirstCrownCondition
¶
Bases: TerminationCondition
Curriculum: end as soon as any crown changes hands since reset.
A termination, not a truncation: in this curriculum task the crown IS the outcome, so there is nothing to bootstrap from.
AnyCondition
¶
Bases: _CompositeCondition
OR of several conditions. Every child is checked every step (so counters advance).
AllCondition
¶
Bases: _CompositeCondition
AND of several conditions. Every child is checked every step (so counters advance).
The engines¶
RustEngine is the real one. MockEngine is a pure-Python stand-in that runs the whole API
without the Rust build, which is useful before you have built anything and useless for anything
about game fidelity.
royalegym.rust_engine
¶
RustEngine: protocol.Engine over the compiled Rust core (../RoyaleSim/crates/royalesim).
WHAT IT IS
The adapter that lets ClashParallelEnv and everything else in royalegym run
on the real engine with no change to the RL layer. The Rust side is
royalesim.Battle (../RoyaleSim/crates/royalesim/src/py.rs). A step is one
Rust call (validate, apply, tick N times with the GIL released); the state
snapshot is one JSON byte string decoded straight into protocol structs.
CONVENTIONS, AND HOW EACH IS KNOWN RATHER THAN BELIEVED
* Positions: protocol.py's engine frame IS the Rust frame (Blue defends low
y, origin at Blue's back-left corner, integer subtiles). They pass through
_engine_xy unchanged. tests/test_rust_engine.py cross-checks tower
centres, passability of every half-cell and deploy legality over a grid
against MockEngine, which reads the same arena.json independently.
* Tower slots: protocol TowerSlot is named in the OWNER's frame under a
180-degree seat rotation, so Red's own-LEFT princess is the engine's
right-lane tower. The table is not typed in: _derive_slot_of_k matches
the Rust tower centres to Arena.princess_centers / king_centers
(protocol's naming) and refuses to start if they are not a bijection.
* Card ids: position in the catalogue (Supercell internal names from
data/derived/cards.json, e.g. "Archer" not "Archers").
* Deploy reasons: Rust returns indices into DEPLOY_REASONS, a list of
DeployStatus NAMES, so no protocol number is copied into Rust.
* Troop territory: the engine and the mask run the same shipped mechanic, the
closed NoDeploySize rect of every alive enemy crown tower. The engine reads
its sizes from the cards.json in the checkout it was built in, the mask from
the one under data_dir(); construction compares
Battle.tower_no_deploy_rects() and TERRITORY_MODEL with DeployRules
and refuses on any difference (territory_differences).
* MatchSetup: protocol.validate_setup runs FIRST in reset, before the
core or this adapter touches anything, so a refused setup raises the
Protocol's ValueError (same text as MockEngine) and the running battle is
untouched. Checking only spawn_violation and the deck ids here would let
start_tick -1 or 232, shuffle -1 and tower hp 231 reach PyO3 as an
OverflowError, and a malformed tower_hp as an IndexError in the slot mapping --
neither of which MockEngine raises. The core's
own checks (py.rs Battle.reset, state.rs scenario_spawn_now) are the
authority the protocol rules were written from; tests/test_rust_engine.py and
tests/test_parity_hardening.py hold the two together by bypassing this check.
* Spells: the catalogue kind code IS the deploy rule, mapped to
Placement by _PLACEMENT_OF_KIND (0 TROOP, 1 BUILDING, 2 SPELL, 3
ROLLING, 4 SPELL_NOT_ON_WATER, 5 TUNNEL, 6 MIRROR). A unit a spell releases
(the Goblin Barrel's Goblins) is not a card; the core reports it under the releasing spell's
catalogue id. state() decodes the core's live spell objects into
BattleState.spells and the stun / knockback timers into EntityState.
STALE BUILDS ARE REFUSED
The Rust crate compiles calibration.json and arena.json in (include_str!).
A calibration edited after the last maturin develop would silently run the
old constants, so construction compares every calibration value and the
arena with the files on disk and raises on any difference. The same check
refuses a Calibration.with_override the compiled engine cannot honour.
THE CARD TABLE IS READ, NOT COMPILED IN
Every royalesim.Battle reads data/derived/cards.json from the RoyaleSim
checkout the engine was built in, when it is constructed (card.rs load_repo).
Re-running tools/extract_cards.py --vintage 2018 --out data/derived/cards.json
in that checkout changes the next RustEngine's card table at once: no rebuild, and
no stale-build refusal, because there is nothing compiled to be stale.
ROYALESIM_DATA_DIR does not move that file; it moves only what this package
reads. So build in the checkout whose data you want.
Which table an engine got is in its config(): cards_json_fnv1a64 (the hash
RoyaleSim's replay fixtures record), cards_vintage and
cards_json_hash_source (see RustEngine.card_table_stamp).
WHAT DIFFERS FROM MockEngine ON PURPOSE (engine mechanics, not adapter choices)
Cards and towers run at the engine's unified level (card_level, 9 on the
2018 rarity table) where the mock uses CSV level 1; spells travel, roll, stun and
knock back here and resolve instantly there; crowns_from_destroyed_towers=False
is refused.
RustEngine
¶
Deterministic Rust battle engine behind protocol.Engine.
set_command_delay_ticks(delay)
¶
THE COMMAND DELAY, in ticks, from the next reset on: an int for both seats, or
(Blue, Red). A play or press accepted on tick T runs on T + delay, checked again in
full then; until it runs the card stays in hand, the bar unspent, and the card or
button refuses another command (CARD_PENDING). 0, the default, runs every command at
once. The live client's is measured at 21-22 ticks (RoyaleSim df69520).
A delay above 0 needs an engine that has it and can name its refusal: refused here, not at the first refused command.
unit_hitpoints(card_id, level)
¶
Every unit a play of card_id at unified level puts on the board, as the
engine's (role, unit name, hitpoints) rows: the card's own row first (role "own",
the unit its catalogue row describes), then each unit it puts down, each at the level
it takes (roles "second_summon", "spawn", "death_spawn", "release"). The engine's
rows as they are (RoyaleSim Battle.unit_hitpoints), so a consumer can tell a
unit's hitpoints at the level it was played at, which the catalogue row gives only
at the catalogue's level (a Knight: 1766 at 11, 1938 at 12).
ValueError for an unknown card id or a level the card's ladder lacks, naming it.
NotImplementedError from an engine build before the rows (RoyaleSim 6909b6f).
Not part of protocol.Engine: MockEngine has no levels, so a consumer asks for it
with getattr and keeps its own rule where it is absent.
reset(seed, setup)
¶
protocol.validate_setup first -- nothing is read into the core, and no
adapter table is indexed, until the whole setup has passed. See the module doc.
building_placement(team, card_name, x, y)
¶
Where a building tapped at ENGINE-frame (x, y) would end up, or None.
(centre_x, centre_y, (x0, y0, x1, y1)), engine frame, subtiles. A tap whose
box does not fit is not refused: the engine moves the building to the nearest
place it does fit, so the centre this returns is often not the point tapped.
None means the tap is refused outright, which a tap outside the arena, on water,
on a no-deploy cell or outside the card's territory still is.
This is on the engine because WHERE a building lands is the engine's rule.
Working it out here would be a second copy of that rule, and the two would part.
action.GridActionParser asks it for the taps_where_the_building_stays
arm, and an engine that cannot answer refuses that arm rather than pretending.
config()
¶
Constructor state, JSON-able, for ClashParallelEnv.config().
Everything that makes two RustEngines run different battles from the same
commands: which cards are in the catalogue, the level they run at, and which
pathfinder and deploy-clamp arms were selected. SymmetricRustEngine is a
different engine for this purpose and a checkpoint that does not say so cannot
be reproduced -- and the clamp has to be named separately from the pathfinder,
because selecting one does not select the other. The card NAMES do not say
which card table they were read from, so card_table_stamp is in here too.
engine_identity()
¶
Which engine this is: the data compiled in (build_digest) and the binary
loaded in THIS process (engine_binary_sha256). A trace header records both,
taken while it records, which is the moment the docstring of
engine_binary_digest asks for.
card_table_stamp()
¶
Which card table this engine read, taken when it was constructed.
cards_json_fnv1a64: FNV-1a 64 of the cards.json bytes, 16 hex digits, the
value RoyaleSim's replay fixtures record under the same name.
cards_vintage: that file's provenance.vintage.
cards_json_hash_source: "engine" when the compiled engine reported the
hash of the table it loaded (ENGINE_CARD_HASH_NAMES), and then the vintage
is "unknown" unless cards_json_path hashes to the same value;
"disk_at_construction" when it was hashed from cards_json_path, the
file the engine reads, just before and just after the engine read it (a file
that changed in between is refused); "unavailable" when that file could
not be found, and then the other two are "" and "unknown".
debug_nudge(uid, dx, dy)
¶
TEST-ONLY: move a live entity by (dx, dy) subtiles.
SymmetricRustEngine
¶
Bases: RustEngine
RustEngine under the frame-planned pathfinder (path_search="trace_fitted_astar"),
which also selects the fixed-distance knockback, AND the own-frame deploy clamp
(ground_y_clamp="deploy_column_range_own_frame"), which it does not (see
RustEngine.__init__).
FOR ROTATION-MIRROR GATES ONLY. The shipped search is the game's own, measured on
client 16.402, and it is not seat-symmetric: its goal scan and
neighbour order run in absolute arena coordinates, so a Red unit and its
rotated Blue twin can publish different equal-cost routes (measured: 20 of 54 twin
problems on the shipped arena).
That is the real game and RustEngine reproduces it. The tests that assert a
mirrored battle stays a rotation to the end exist to catch seat bias in the ENGINE'S
OTHER SYSTEMS and in this adapter, so they run under this class, whose routes are
exact rotations by construction.
THE DEPLOY CLAMP IS THE SAME KIND OF THING AND HAD TO BE ASKED FOR SEPARATELY.
formation.GROUND_Y_CLAMP's shipped arm is measured per side and is deliberately
not the rotation of itself, so a multi-unit card deployed at mirrored points does
not land at mirrored positions -- measured through the engine, Skeletons over the
230 legal mirrored tile centres broke the rotation 22 times under the shipped arm
and 0 times under the own-frame one. Until the core exposed the key this class set
only path_search, so the rotation gates ran against the asymmetric clamp and
correctly reported an asymmetry that is real and intended.
A KEY WITH NO KEYWORD IS SELECTED THROUGH calibration_overrides: SYMMETRIC_ARMS.
The first was targeting.FIRST_TOWER_PICK (RoyaleSim 126992a). Its shipped arm reads
a summon member's lane in the arena's frame, measured that way on both seats, so a
Skeleton Army on the centre line splits one member differently for Blue and for Red.
Three rotation gates went red on it at tick 140 while this class's one-tick probe read
clean: a lane pick shows only when the members start walking. This class took the old
current_x arm until RoyaleSim 1d661b0 added client_spawn_lane_own_frame, the
client's lane rule read in the placer's own frame, which it now prefers.
symmetry_problems()
¶
Where this vehicle is NOT a rotation mirror, measured. Empty means it is.
Not raised from __init__ on purpose. Some callers want this class for its
determinism rather than its symmetry -- the spawn-ordinal parity gates, for
instance -- and refusing to construct would fail six tests whose claim has
nothing to do with seats. The check belongs where the claim is made:
rotation_divergence calls it before measuring anything, and one test
asserts it is empty so the state of the vehicle is visible on its own.
symmetry_report()
¶
symmetry_problems with what to do about it. Empty when there is nothing.
core_import_message(error)
¶
What a reader is told when the engine is not installed: the install page, whose line works before PyPI too, then the page for a source build.
core_available()
¶
Whether the battle engine (royalesim) is installed; CORE_IMPORT_ERROR says
why not when it is not.
engine_binary_digest()
cached
¶
A short hash of the compiled extension FILE that is loaded, or "unknown".
build_digest identifies the DATA an engine was built with. Nothing identified
the CODE: the extension exposes no version, and its distribution version is a
constant, so an engine rebuilt from a modified or uncommitted Rust tree was
indistinguishable from the one before it. Every guard here, and the ones in the
sibling repos, compared data and would have passed.
This does not say which commit built it, which is the engine's to publish. It says
whether the binary is the same binary, which is the part that was invisible: two
builds from different sources do not produce the same file. Read it beside
build_digest -- one moving without the other is the interesting case, and a
result recorded against a binary nobody can identify is worth less than one that
names it.
Cached: about 4 ms over 2.3 MB, paid once per process. "unknown" rather than raising when the module is not a file on disk, because a missing provenance stamp should not stop a battle.
TAKE IT IN THE PROCESS THAT PRODUCED THE RESULT. This describes the binary THIS process loaded. Fetched later from a second process it describes whatever is on disk by then, which is not the same claim and is the easier one to make by accident: a recording, a picture or a metrics row stamped after the fact says which engine exists now, not which engine made it. Record it beside the result, not beside the report.
build_digest()
¶
A short hash of the data the loaded extension was COMPILED with.
The same hash protocol.calibration_digest takes of calibration.json on
disk, over the copy compiled into the extension, plus the arena it was built
with. A checkpoint pins this next to the on-disk digest: equal means the
policy was trained on an engine built from the data that is there now, and
unequal says which way to look without needing the two files side by side.
Module-level and also reachable as RustEngine.build_digest so it can be
read without constructing an engine -- construction is exactly what refuses
on a stale build. The card table is not in it, because it is not compiled in:
RustEngine.card_table_stamp says which one an engine read.
stale_build_differences(calibration, arena_path=None)
¶
What differs between the compiled-in data and calibration / arena.json.
engine_cards_json_path()
¶
(path, how it was found) of the cards.json the compiled engine reads.
Not data_dir(): ROYALESIM_DATA_DIR moves what this package reads and never
what the engine reads. In order: the path the engine states, if it states one
(ENGINE_CARD_PATH_NAMES), found by "engine"; the build directory compiled into
the extension, "build checkout"; the sibling checkout of the documented layout,
"sibling checkout", a fallback that is right only when the engine was built there.
cards_json_stamp(path)
¶
(FNV-1a 64 of the file's bytes, its provenance.vintage), hashed once per version.
The hash is protocol.fnv1a64 over the raw bytes, the one RoyaleSim's replay
fixtures record as cards_json_fnv1a64. Keyed on size and modification time, so
a regenerated file is hashed again and an unchanged one is not.
catalogue_vintage_split(rust_cards, mock_cards, engine_vintage=None)
¶
Why the two engines are reading DIFFERENT card tables, or None.
DECIDED BY THE TABLE, NOT BY THE ROWS. The engine's cards.json names the table it is
(provenance.vintage, or engine_vintage when given). When that is the table
extracted from MockEngine's own pack (RAW_CARD_PACK_TABLE_VINTAGE), the two read
the same data, and any row difference is a defect in an engine or in this adapter,
which builds the rows: None is returned and the caller's comparison FAILS on it.
Until 2026-09-27 a row difference alone meant "different tables", so an adapter bug
in one of these fields (a radius scaled in Python) read as a table split and skipped.
The compiled engine reads data/derived/cards.json from the RoyaleSim checkout it was
built in, each time an engine is constructed; MockEngine reads the raw CSVs under
data_dir(), every time. In a public checkout that is built there, the two are
the same vintage by construction: the newer client packs are not redistributed, so
the extractor can only produce the tracked table, and any difference here is then a
real defect. On a machine that HAS a newer pack and regenerated cards.json from it,
the two are simply different tables and every cross-engine comparison is measuring
the data rather than the engines.
THE TWO HALVES ARE READ FROM DIFFERENT PLACES, which is the part that catches people
out. Pointing ROYALESIM_DATA_DIR at a 2018 data directory moves MockEngine's half
and does not move the engine's at all -- measured: with a pure 2018 data dir the
extension still reported 95 cards and Goblins at 4, because it went on reading the
cards.json of the checkout it was built in. Regenerating THAT file moves the engine's
half at once, with no rebuild. So build in the checkout whose data you want.
stale_build_differences does not cover this: it compares calibration.json and
arena.json, the two files compiled in, and the skip is what handles the catalogue.
RustEngine.card_table_stamp names the table an engine actually read.
THAT IS OBSERVED, NOT PREDICTED. A clean clone of the four repos, its data generated from the tracked 2018 tables and royalesim built in it, runs the tests this function guards: 96 passed, 0 skipped, so the two engines agree over the shared card set and the contract holds. On a machine that also holds a private client pack they skip instead, which is this function working rather than the contract going unchecked.
Returned as a ready reason string so a test can skip on it and SAY SO. A skip is not a pass: the comparison that skipped still has to run somewhere.
territory_differences(battle, rules, arena, slot_of_k)
¶
What differs between the compiled engine's troop territory and rules.
Reads the numbers the ENGINE holds (Battle.tower_no_deploy_rects, engine
tower k order, mapped to protocol slots through slot_of_k), never the
adapter's own copy: a check that compares the Python rules with themselves
cannot fail.
check_field_order(core)
¶
Refuse an engine whose positional columns do not line up with this package's.
WHY THIS EXISTS. EntityState, SpellState and ProjectileState are array_like: the
engine's state_json sends each row as a JSON ARRAY and it is decoded BY POSITION. A
type mismatch raises on its own, but two adjacent fields of one type do not -- on
2026-09-24 the engine gained target_uid and attack_phase, two ints side by
side, and a swap between sim's order and this package's would have drawn a uid as an
attack phase without a word. The day before, DEPLOY_REASONS had failed exactly this
way: a positional list grew in one repo and not the other.
THE RULE. The engine's list must be a PREFIX of ours. Ours may be longer -- a newer royalegym reading an older engine, whose missing trailing columns decode to their "not reported" defaults. The engine's may not be longer, and no shared column may be in a different place.
Returns, per export name, whether that layout was actually checked. An engine that exports no field list passes UNCHECKED, and says so, rather than passing silently.
rotation_probe(engine, tiles=PROBE_TILES)
¶
Deploy each multi-unit card at the same own-frame tile from both seats; report drift.
THE DIRECT MEASUREMENT of the property a rotation gate assumes. Both seats command the same tile in their own frame, so after one tick the two groups must be exact rotations of each other. Anything else means some system under the engine is keyed on the seat or on the absolute arena frame.
Multi-unit cards only, because a single unit lands on the tap and a formation is laid out AROUND it -- which is where a per-side offset shows up. Troops only: a building or a spell does not have a ring.
This exists because the alternative failed twice. Deciding from a list of which keys are asymmetry sources means the list is right only until the engine gains a key, and both times it gained one the gates went red pointing at the wrong thing -- at the observation builder, and at the engine. Measuring the property instead is immune to a key nobody has named yet.
troop_ruled_spells(engine, battle_args, overrides)
¶
The names of the catalogue's kind-3 spells (placement ROLLING) that the core judges
by a TROOP's rule: accepted on free ground in the caster's own half, refused OCCUPIED on
its own princess tower. Measured on a probe battle built exactly as engine's, both
seats, and cached per build, card table, catalogue and constructor arguments. Both seats
must agree, or the spell is left as it is and the every-card mask gate names it.
battle_takes(keyword)
¶
Whether the core's Battle constructor takes keyword, from its text signature
(pyo3 #[pyo3(signature = ...)]). False without a core or a signature.
tunnelling_cards(engine, battle_args, overrides)
¶
The names of the catalogue's troop and building cards (codes 0 and 1) that the core
accepts in the ENEMY half, on free land well inside the enemy's tower rects, where it
refuses every other troop and building OUT_OF_TERRITORY: the cards that travel under
ground (placement.SPAWN_PATHFIND_TERRITORY). Measured on a probe battle built exactly
as engine's, both seats, and cached per build, card table, catalogue and
constructor arguments. Both seats must agree, and the same card must be accepted on
free land in its own half, or it is left as it is and the every-card gate names it.
A card the opening deal never puts in hand is not asked, and keeps its code.
symmetric_overrides(ledger)
¶
The symmetric arm of each SYMMETRIC_ARMS key this compiled ledger ships some OTHER
arm of: the most preferred one its candidates list (the first, if it lists none).
A key the ledger lacks is left out, so an engine older than the key still constructs; one already shipping the arm is left out, so nothing is overridden for nothing; one listing none of the symmetric arms is left out, and the rotation probe reports it.
royalegym.mock_engine
¶
MockEngine: a pure-Python stand-in that satisfies protocol.Engine.
WHAT IT IS FOR Building and testing the whole RL layer before the Rust engine exists. It is deliberately simple, but it is not a toy in the three ways that matter to the RL layer: it uses the REAL arena (water, bridges, no-deploy, king block) from data/derived/arena.json, it reads every physics constant it uses from calibration.json (or globals.csv when the registry does not have it yet), and it obeys the three invariants the real engine must obey.
WHAT IT IS NOT A fidelity claim. No unit collision, no projectile travel (ranged hits are instant), no stun/slow, spells hit once, splash is centred on the target, all cards and towers at CSV base level, overtime that ends level is decided by calibration match.OVERTIME_TIEBREAK (the same rule table the Rust engine reads; the rule itself is not yet measured on a recorded match). None of this should be used to argue about how the real game behaves.
ITS SPELL EFFECTS ARE NOT THE ENGINE'S. The Rust core (../RoyaleSim/crates/royalesim
spell.rs) flies Fireball / Arrows / Goblin Barrel to their target, lands and
rolls The Log, stuns with Zap and knocks units back. Here every spell resolves
in the tick it materialises: circle spells and the Log's strip hit once at the
tap, no stun, no knockback, and the Goblin Barrel's Goblins appear at the tap at
once. So ``BattleState.spells`` is always empty and every ``stun_ticks`` /
``knockback_ticks`` is 0. What IS held to the engine is the part the RL layer
must agree on regardless of mechanics: which cards are spells, each spell's
placement class, every deploy verdict and reason code (tests/test_rust_engine.py,
three-way at every half-cell), the catalogue row shape (count/radius/hp 0), and
that a released Goblin reports its barrel's card id. Do not read a spell trade
off a MockEngine battle.
THE INVARIANTS, AS THEY APPLY HERE
1. No floats: integer subtiles, integer elixir with an exact rational scale,
truncation-toward-zero division (_tdiv) so that a Red unit's step is
exactly the negation of the mirrored Blue unit's step. Python's //
floors, and floor(-a/b) != -floor(a/b): using it here breaks the mirror
test, which is how that test was checked (see tests/test_env_mock_engine.py).
2. Simultaneity: every phase reads start-of-phase state and writes into
buffers (targets, proposed moves, damage, deaths, spawns) applied in one
pass. Every tie-break is in the ACTING team's own frame, so nothing
prefers Blue.
3. No baked numbers: see MockEngine.__init__.
Pcg32
¶
Bases: Struct
Python port of royalesim::Rng (lib.rs). Same stream for the same seed.
Cross-language equality has NOT been verified against the compiled crate; a parity test for it belongs here once royalesim exposes Rng.
MockEngine
¶
Deterministic, seeded, integer-only. Satisfies protocol.Engine.
config()
¶
Constructor state, JSON-able, for ClashParallelEnv.config().
The card NAMES, because a catalogue subset is what makes one run's card ids
mean something different from another's, and card_level for the same
reason the Rust adapter reports it. footprint_model because it decides
where a building may stand (FOOTPRINT_MODEL). card_table_stamp because
the names do not say which numbers were read for them.
card_table_stamp()
¶
Which card table this engine read, as RustEngine.card_table_stamp says it.
cards_vintage: the raw client pack the stats come from (RAW_CARD_PACK).
cards_loaded_fnv1a64: FNV-1a 64 (protocol.fnv1a64) of every template
built from that pack -- units, spells and cards, msgpack in catalogue order --
so it moves with any number read from the CSVs and not with a column nobody
reads. There is no cards_json_fnv1a64: this engine reads no cards.json.
reset(seed, setup)
¶
Validate EVERYTHING (protocol.setup_violation), then build the new battle
off to the side, then commit it. A refused setup raises ValueError and the
running battle is untouched, as in the Rust core's Battle.reset (which
builds a local BattleState and assigns it last).
The order matters. Validating decks and spawns, assigning self._s and
only then refusing a destroyed king mid-construction leaves a refused reset
with tick 0, the new elixir and hands and zeroed Red towers in place.
load_state(blob)
¶
Restore a snapshot, matching its cards to this catalogue BY NAME.
Refuses (ValueError, running battle untouched) when a hand, a queue, a
pending deploy or a non-tower board entity names a card this catalogue does
not have -- the Rust catalogue_violation rule and messages. A snapshot
from a permuted superset catalogue is re-indexed and plays the same cards.
Decoding the blob straight into self._s reads its ids through whatever
catalogue loaded them: a Knight snapshot plays as a Valkyrie, and a one-card
catalogue crashes.
tiebreak_drain_step(lowest)
¶
The hp every standing crown tower loses this tick, from the lowest crown tower hp of all six read before it (state.rs tiebreak_drain_step).
snapshot_catalogue_violation(s, catalogue)
¶
Why snapshot s cannot run behind a catalogue holding catalogue names.
The Rust py.rs::catalogue_violation rule, in its order and with its messages:
every card a battle can still produce -- a hand slot, the cycle queue, a pending
deploy, a non-tower entity on the board -- must be in the catalogue. Crown
towers are never catalogue cards. Module-level so a test can plant it away.
Recording and watching¶
royalegym.replay
¶
Battle traces: record, save, load, and VERIFY by re-simulation.
A trace is two things at once
- a replay in the sense calibration.json means -- seed + initial setup + the
command log + periodic state hashes -- which
verify_tracere-simulates and checks hash by hash; and - enough per-frame entity data (positions, hp, radius) for
render.pyto draw the battle as a self-contained HTML page without an engine (python -m royalegym.render trace.msgpack -o battle.html).
Encoding is msgspec msgpack (.msgpack, compact) or JSON (.json, for
eyeballing). Per-entity rows use array_like structs, so a frame is a list of
short arrays; the header names the columns so the viewer reads them by name
(render.py refuses a trace that lacks a column it draws).
State hashes are hex strings, because a u64 does not survive a JavaScript Number.
SPELLS (2026-09-13). A frame also records the engine's live spell objects
(BattleState.spells) and, through EntityState, each entity's stun and
knockback timers; the header names the spell columns in spell_fields. Both are
trailing fields with defaults, so a version-1 trace recorded before they existed still
decodes (with no spells and zero timers) and the version number did not move. A
MockEngine trace has no spell objects: its spells resolve inside a tick.
ReplayRecorder
¶
Attach to an env (recorder=). Keeps the most recent trace in trace.
frame_every_tick=True records a frame per engine tick (smooth viewing,
and the tightest determinism check); False records one per decision.
on_finish is called with each completed trace, e.g. to save it.
SavingReplayRecorder
¶
Bases: ReplayRecorder
A recorder that WRITES completed traces, so a rollout worker can produce them.
WHY THIS EXISTS RATHER THAN on_finish
ReplayRecorder already takes on_finish, and for a script that is the better
tool. It is unreachable from a training run: the recorder is an EnvFactory
COMPONENT, given as an importable class plus JSON kwargs, and a callable is not
something JSON can express. So a worker that gets a pickled recipe can be handed
this class and a directory, and cannot be handed a function. Every argument here is
a string, an int or a bool for that reason.
Asked for by train, who had built and measured the rest: a 3-minute battle is 3,601
frames and 1.6 MB on disk, and their viewer plays one to a real viewer at wall clock
in 57 MB. Watching a live run cannot work -- `time/collection` is 4.35 s of a 62.94 s
iteration, and those 4.35 seconds hold about 82,000 engine ticks, so it is a firehose
between silences rather than slow gameplay. Replaying a saved trace is the form that
works, and the training run producing them was the one missing piece.
WHAT IT COSTS THE RUN
One write every every completed episodes, and nothing else. Recording itself is
whatever frame_every_tick already costs; this adds serialisation of a finished
trace, not per-tick work.
THREE THINGS IT DOES THAT A NAIVE VERSION WOULD NOT, each because a rollout worker is
not a script:
* The PID is in the name. Workers are separate PROCESSES with separate counters,
so two of them would otherwise write the same 000050 file and one would win.
* The write is atomic, via a temporary file and os.replace. A viewer reading
the directory while a worker writes would otherwise find a half-written trace, and
msgpack does not fail loudly on one.
* The directory is created and tested at CONSTRUCTION, so a bad path fails when
the config is loaded rather than after the first episode of a five-hour run. A
recorder that silently saves nothing is the failure train hit from the other side:
a publisher that reported "3601 frames, 0 sent" and looked like it worked.
save_trace(trace, path)
¶
Write a recorded battle to path (msgpack, or JSON when it ends in .json) and
return the path. royaleviser PATH opens it.
load_trace(path)
¶
Read a battle save_trace wrote (msgpack, or JSON when it ends in .json).
verify_trace(trace, engine)
¶
Re-simulate trace on engine and list every divergence (empty = identical).
Checks the catalogue first, then the deploy statuses of every step, every recorded frame hash, and the final hash. The first divergence is usually the informative one; later ones are consequences.
WHEN THE REPLAY DIVERGES AND THE ENGINE IS ANOTHER BUILD, the first line says so, naming both. It is said only on a divergence: a different binary that replays exactly is ordinary (a CI runner compiles its own), and the note is there so that a list of hash lines is not read as a changed battle when it may be a changed build.
THE CARDS A SETUP DEALS ARE COMPARED BY NAME first. A setup names its cards by POSITION (a deck is a list of catalogue ids), and the catalogue renumbers whenever a card becomes loadable, so the same setup on an engine with a different table deals different cards: the first deploy diverges and the hash list says nothing about why. So when a dealt id names another card on this engine, or none, the result is one line naming it, and nothing is replayed. Only the ids the setup uses are compared: an id no deck or spawn holds changes nothing a replay reads. A trace that starts from a snapshot is not compared, since the snapshot holds the engine's own card indices.
royalegym.viser
¶
The engine side of RoyaleViser: publish battle state over UDP only while a viewer watches.
The viewer's rule: a separate process, never in the tick loop, zero cost when nobody is watching. This module is the whole of what royalegym knows about the viewer; it imports nothing from RoyaleViser (dependency direction stays RoyaleLearn -> RoyaleGym -> RoyaleSim) and nothing graphical.
from royalegym.env import ClashParallelEnv, ClashSelfPlayVecEnv
from royalegym.viser import ViserPublisher
env = ClashParallelEnv(viser=ViserPublisher()) # 127.0.0.1:9870
# or, for self-play, with the ROYALEVISER environment variable set to 127.0.0.1:9870:
vec = ClashSelfPlayVecEnv(8) # binds ONE, watches game 0
# then, in another process: python -m royaleviser --stream 127.0.0.1:9870
WHO READS ROYALEVISER
ViserPublisher.from_env(), and only ClashSelfPlayVecEnv calls it.
ClashParallelEnv publishes when it is HANDED a publisher and reads no
environment variable of its own: a viewer watches ONE battle and has one fixed
port, so N envs each building their own from the variable is OSError 10048,
which is what self-play with more than one game used to raise. The decision
belongs to whoever knows how many battles there are.
PROTOCOL
The viewer sends the heartbeat datagram HELLO to (host, port) once a second while it is
open. publish costs one clock read while nobody has said hello in ATTACH_TIMEOUT_S:
the socket is polled for heartbeats at most once a second (a non-blocking recvfrom).
While attached, every call encodes one frame with msgspec msgpack and sends ONE datagram
to the last heartbeat's address. A datagram over MAX_DATAGRAM bytes (UDP over IPv4) is
resent with every unit's path emptied, then counted in dropped if still too big.
Measured 2026-09-20 (MockEngine): 2.6 KB per frame with 12 units + 6 towers, 11 KB with 60
entities, far under the limit.
WIRE FORM
frame_dict is royaleviser.model.Frame as a plain dict with that dataclass's field
names, so royaleviser.model.decode_frame reads the datagram without a converter and
royaleviser.sources.TraceSource draws a trace with the same unit/spell rows. The
field list is repeated here on purpose (../RoyaleViser/tests/test_sources.py round-trips a
published frame through the viewer's decoder and its contract check, so a drift fails
there). Positions stay in engine subtiles; units_per_tile says how many per tile.
ViserPublisher
¶
Sends frames to an attached viewer; detached, publish is one clock read.
port 0 binds an ephemeral port (tests); address is the bound (host, port).
publish(state, cards, arena) builds the frame dict, adds the spawn/death lines it
derives from the previous published state (only while attached, so a viewer that
attaches mid-battle does not see every unit "spawn"), and sends it.
publish_dict sends a ready dict (royaleviser.sources.Publisher's path).
from_env()
classmethod
¶
A publisher for ROYALEVISER=host:port, or None when the variable is unset/empty.
ROYALEVISER=1 (or true, on, yes) means the default address, HOST:PORT.
ROYALEVISER_RUN names the run these frames come from. The ports are fixed, so
two runs on one machine reach the same viewer and it cannot tell whose frames it
is drawing beside whose learning panel; the name in every frame is what lets it
say. Unset is the empty string, and then no run key is sent at all.
publish(state, cards, arena, decks=None, events=(), meta=None, forms=None)
¶
One datagram for state if a viewer is attached. Returns whether one was sent.
publish_dict(d)
¶
Send one frame dict to the attached viewer (the size policy lives here).
names_of(cards)
¶
card id -> name for one catalogue; unknown ids print as #<id> (never silent).
unit_dict(e, name_of)
¶
royaleviser.model.Unit as a dict. Path is not in an engine state.
footprint is the engine's own box for a building or tower, passed through as it
came: None for a troop and for an engine that reports no box, so the viewer draws
its marked stand-in rather than a box made up here.
NOT REPORTED IS None, NEVER A VALUE. target, direction and state were
hard-coded None until the engine exported them (2026-09-24). An engine that still
exports nothing decodes to EntityState's "did not say" defaults, and each maps back
to None here: no target, a zero facing, and attack_phase -1. A zero facing in
particular must not reach the viewer as [0, 0], because the viewer normalises the
direction's length and a zero vector has none.
status is the buff list as the engine names it, names whole: the viewer splits
sim's "|"-joined names itself.
projectile_dict(p, name_of)
¶
One projectile in flight, for the viewer. ENGINE frame, subtiles, like units.
name is what fired it: the card, "tower" for a crown tower (firer -1), and None
when the engine does not know (firer -2, a projectile restored from a snapshot older
than the field). Sim keeps -2 apart from -1 so that "unknown" never reads as "a tower
fired this", and mapping every negative id to "tower" would undo exactly that.
player_dict(p, name_of, deck, forms=None)
¶
royaleviser.model.Player: the engine knows everything but the cycle beyond next_card.
An empty hand slot (EMPTY_CARD) is the empty string: known to be empty, not unknown.
The special forms, by NAME: evo is the engine's rows with the card's name for its
id; abilities is each button's [name, available, spent, cost], the viewer's four
columns. The name is the row's card where the engine states it (a champion's button
does), else the k-th hero entry of the deck, else "" where the deck's forms are unknown.
ability_cooldowns is its own key, parallel to abilities: each button's
cooldown ticks left (0 when ready, and 0 while a champion's ability runs, the button
then dark), -1 where the engine's row does not say. Only when some row says, so a frame
from an engine without the column is what it always was.
frame_dict(state, name_of, units_per_tile, decks=None, events=(), meta=None, forms=None)
¶
The viewer's Frame for one BattleState (see WIRE FORM).
play_event(tick, team, name, x, y, units_per_tile)
¶
One events-panel line for an accepted deploy: "t120 Blue plays Knight (3.5, 14.5)".
Self-play bookkeeping¶
royalegym.selfplay
¶
Self-play: opponents, an opponent pool, sampling and Elo bookkeeping.
OPPONENTS
An Opponent maps (observation, action mask, rng) -> action. The rng is
the env's own generator, so a seeded env with a stochastic opponent is still
reproducible (gymnasium's step-determinism check relies on this).
POOL SAMPLING
uniform every snapshot equally likely -- broad, protects against forgetting
latest always the newest -- pure self-play, fastest to cycle into rock-
paper-scissors loops
pfsp prioritised fictitious self-play (AlphaStar): weight each snapshot
by f(P[learner beats it]); hard f(p)=(1-p)^power focuses on
opponents the learner still loses to, variance f(p)=p(1-p)
focuses on even matchups.
Win probabilities use a Beta(1,1) prior, so an unplayed snapshot reads 0.5
instead of dividing by zero or being ignored forever.
ELO Standard logistic Elo with base-10 / 400 scale. It is bookkeeping for humans and for PFSP diagnostics, not a training signal. Elo assumes transitive strength; a Clash meta is not transitive, so read it with that in mind.
NoopOpponent
¶
Never plays a card. The easiest possible scripted opponent.
RandomLegalOpponent
¶
Uniform over legal actions, taking no-op with probability noop_prob.
Pure uniform over 2305 actions picks no-op ~0% of the time and spams cards the instant they are affordable, which is a strange opponent to learn against.
CallableOpponent
¶
Wrap a frozen policy function fn(obs, mask) -> action.
OpponentPool
¶
Snapshot registry with sampling strategies and Elo/head-to-head tracking.
The learner is an ordinary entry (conventionally id "learner") so its Elo
is tracked on the same scale as the snapshots it plays.
record_result(a, b, score_a)
¶
Record one game. score_a is 1 (a won), 0.5 (draw) or 0 (b won).
Returns the new (elo_a, elo_b). The update is zero-sum: the pool's total
Elo is conserved, checked by
tests/test_selfplay_pool.py::test_the_elo_update_is_zero_sum. It was not
checked by anything for the life of this class, while this line said it was --
which is worse than saying nothing, because it stops the next reader looking.
win_rate(a, b)
¶
P[a beats b] with a Beta(1,1) prior.