Observation spec¶
This page is for contributors writing or reading observation code. It lists every
channel and every vector slot a royalegym observation builder writes, the range each
one takes, and whether it is fair (something a person watching the match could
write down) or a reveal, read out of the half of the state a player cannot see.
The code is royalegym/obs.py. This file is the reference. vector_layout(),
vector_offsets(), spatial_channels() and ObsBuilder.channel_names() are the same
thing in a form a program can read, and the test suite asserts the two agree.
Throughout, n is the number of cards in the engine's catalogue
(len(engine.cards())). Everything is in the acting seat's own frame: Red's board
is rotated 180°, Red's towers are "own". A mirrored battle gives the other seat a
bit-identical observation, which is why one policy can play both seats.
1. The fair / reveal split¶
A builder takes a Reveal, a frozen dataclass whose five fields all default to
False:
from royalegym import Reveal, SpatialObsBuilder
SpatialObsBuilder() # fair
SpatialObsBuilder(reveal=Reveal(enemy_hand=True)) # cheating, and says so: 4(n+1) more slots
An enabled field adds channels or slots; it is never present-but-zero. A fair
observation and a cheating one do not have the same width, so a checkpoint cannot
quietly be trained with one and evaluated with the other. Fair fields come first and
keep their indices, so turning a reveal on never moves a fair feature.
ClashParallelEnv.config() records the Reveal for the checkpoint.
Reveal field |
what it opens | cost |
|---|---|---|
enemy_elixir |
the opponent's bar, read from the state | no new slot: it swaps the source of the slot that already holds the count (§4) |
enemy_hand |
the opponent's four hand slots | +4(n+1) vector slots |
enemy_next_card |
the opponent's cycle position 5 | +(n+1) vector slots |
enemy_deck |
as much of the opponent's deck as the state holds | +n vector slots |
enemy_spell_aim |
where the opponent's live spells will land | +1 spatial channel; +2 appended columns on every entity-list spells row |
own_spell_aim is not gated. The player chose where to throw their own spell.
2. SpatialObsBuilder: spatial, float32 [C, 32, 18]¶
C = 20 fair, +1 with Reveal.enemy_spell_aim. Entities are rasterised into the tile
containing their centre in the own frame. Every plane is clipped to [0, 64]
(SPATIAL_CLIP).
| # | channel | range | meaning | fair? |
|---|---|---|---|---|
| 0 | own_ground_troops |
0..64 | count of own non-flying troops in the tile | fair |
| 1 | own_air_troops |
0..64 | count of own flying troops | fair |
| 2 | own_buildings |
0..64 | count of own buildings (crown towers excluded) | fair |
| 3 | own_towers |
0..64 | count of own crown towers by centre | fair |
| 4 | own_hp |
0..64 | sum of own entity hp / 1000 | fair |
| 5 | enemy_ground_troops |
0..64 | as 0, for the opponent | fair |
| 6 | enemy_air_troops |
0..64 | as 1 | fair |
| 7 | enemy_buildings |
0..64 | as 2 | fair |
| 8 | enemy_towers |
0..64 | as 3 | fair |
| 9 | enemy_hp |
0..64 | as 4 | fair |
| 10 | own_deploying |
0..64 | own entities still inside their deploy timer | fair |
| 11 | enemy_deploying |
0..64 | as 10 | fair |
| 12 | water |
0..1 | fraction of the tile's four half-cells that are water (static) | fair |
| 13 | no_deploy |
0..1 | fraction flagged no-deploy (static) | fair |
| 14 | enemy_troop_zone |
0 or 1 | where the OPPONENT could put a troop right now | fair |
| 15 | own_spells |
0..64 | own live spell objects by current centre | fair |
| 16 | enemy_spells |
0..64 | the opponent's live spell objects by current centre | fair |
| 17 | own_spell_aim |
0..64 | own live spells by aim point (landing point; a Log's roll end) | fair |
| 18 | own_stunned |
0..64 | own entities with stun_ticks > 0 |
fair |
| 19 | enemy_stunned |
0..64 | as 18 | fair |
| 20 | enemy_spell_aim |
0..64 | the opponent's live spells by aim point | reveal |
One known gap in the "fair" column: a unit invisible to its enemy, such as a Royal Ghost, is
still shown to the enemy seat. It is counted in enemy_ground_troops (5) and enemy_hp (9)
where it stands, and the entity-list builder writes it a full row. No builder reads a unit's
status flags. Hiding or marking such a unit waits on a measurement of what a player sees in the
real game. It is not fixed.
Channels 0–11 are integer sums converted to float32 exactly once, so no plane can
depend on the engine's entity list order (see obs.py, NO FLOAT MAY DEPEND ON ENTITY
LIST ORDER).
ObsBuilder.spatial_layout() returns (name, is_static) per plane. Static means a
function of the arena alone (water and no_deploy, and nothing else), so a consumer
that stores observations can hold those two once instead of once per transition. Read
the declaration rather than deciding by sampling. own_towers and enemy_towers are
constant on any sample in which no tower falls, because a crown tower never moves. A
consumer that concluded "static" from that would never see a tower destroyed. A test
asserts exactly that.
What is deliberately not here¶
own_troop_zoneandown_building_zonewere deleted. The action space isDiscrete(2305)= no-op + 4 hand slots × 18 × 32 tiles, so the mask already states per-slot, per-tile legality exactly. The zones were a coarser restatement of it. They cost two of the threePlacementOracle.point_gridcalls the observation made per seat per step. Measured: 6.66 → 2.66point_gridcalls perenv.step(both seats). The measurement alternated the two trees three times in one session, because the absolute numbers move by a third with machine load while the comparison does not. The suite's own throughput report went 766 -> 953env.step/s on the Rust engine and 475 -> 580 onMockEngine(medians of three rounds).enemy_troop_zonestays. It is about the opponent's options and is in no mask.- A knockback channel. Under
knockback.DURATION_MS = 0the push is instant and the timer reads 0 between ticks, so the plane would be a constant zero no coverage guard could check.knockback_tickssurvives as a per-entity feature.
An open question: stun is nearly invisible at 500 ms decisions¶
A Zap's stun is 10 ticks and a decision is 10 ticks. Measured on the Rust engine: a
Zap cast at tick k of a decision leaves stun_ticks = k at the one observation
that follows (1, 4, 6, 10 for k = 0, 3, 5, 9) and 0 at every observation after. So
own_stunned / enemy_stunned fire for at most one step per Zap, and the entity
row's stun_ticks / 100 reads 0.01 to 0.10 for that one step.
The features stay as "is stunned now". That is what the engine reports, and it is what the seat-flip and cell-by-cell tests can check exactly. "Was stunned since the last observation" is the feature a policy could actually use, but it is a different thing: it depends on the decision rate and not only on the state. Recorded here rather than changed quietly.
3. mask_planes, int8 [4, 32, 18], both builders¶
The flat action_mask with index 0 (the no-op) removed, reshaped. The action space
is laid out as 1 + slot * ny * nx + y * nx + x, so this is a view, not a
recomputation: mask_planes[slot, y, x] == action_mask[encode(slot, x, y)].
action_mask, int8 [2305], is still there for the policy head. The info dict no
longer repeats it. ClashParallelEnv.state() leaves out both. Legality is not
state, and a centralised critic does not need 2 304 duplicated numbers per seat.
A parser whose action space is not a grid returns None from mask_plane_shape() and
the key is simply absent.
3b. card_ids, uint8 [2, 32, 18], SpatialObsBuilder only (behind a flag, default off)¶
AGREED WITH learn BEFORE IT IS WIRED, because their stem does an embedding lookup on these and most of the ways to get it wrong fail silently rather than loudly.
spatial cannot tell a Giant from a Knight standing on the same tile: it carries no
per-unit card identity. That is a ceiling on play quality and it is invisible in any
throughput comparison, which is why D2 was nearly decided on the wrong axis. These planes
close it without a new trunk.
| key | card_ids, its OWN key, not a slice of spatial |
| shape | [2, 32, 18] — plane 0 own, plane 1 enemy, in the seat's own frame |
| dtype | uint8, declared Box(0, vocab - 1, dtype=np.uint8) |
| vocab | num_cards + 2, from the LOADED card table: one index per card, plus empty ground and crown towers |
| switch | SpatialObsBuilder(card_identity=True), shipped 2026-09-24; OFF by default |
Why its own key and not a channel of spatial. spatial is declared
Box(0.0, SPATIAL_CLIP, float32), and a learner's codec picks exact-versus-scaled-half
storage from the declared bounds. A card id inside a float box is stored as a scaled half,
comes back as 6.997 instead of 7, and .long() reads card 6. Nothing fails; the policy
learns from a catalogue that is subtly wrong. mask_planes is already a separate
integer-valued key, so this follows a path the codec has met.
The vocabulary.
| index | meaning |
|---|---|
| 0 | empty tile |
| 1 | a crown tower, which is not a card |
2 + card_id |
a unit of that card |
Index 0 is reserved so that "nothing here" and "the first card in the catalogue" are not
the same number. Crown towers get index 1 for exactly the same reason one step along: they
are the only entity class carrying no card (card_id == EMPTY_CARD == -1), and leaving
their tiles at 0 would make 0 mean both "empty ground" and "a tower stands here". It also
carries real information rather than a constant, because a destroyed tower's tile becomes
genuinely empty. Units a spell releases are NOT in this class; the engine reports them
under the releasing spell's catalogue id, so they are ordinary cards.
The vocabulary size is not a constant. It is a fact about the table the engine loaded.
This paragraph said "100 cards on the 15.535 table" until 2026-09-24, when the engine loaded
101 -- the census moves as cards become loadable, which is the paragraph's own point made at
its own expense. It is published as the card_ids Box's upper bound, so a network sizes
nn.Embedding(observation_space["card_ids"].high.max() + 1, C) at construction. That also
puts it in RoyaleLearn's EnvSpec with no separate channel: rollout/envspec.py records
every observation key's shape, dtype, low and high from the observation space itself. A
literal would be wrong on the next census, and an embedding sized by guessing truncates
silently.
Catalogue ids are POSITIONAL, and that is the trap this ships with a guard against.
Making one more card loadable renumbers every later id, so an embedding table indexed by
those ids would have the same checkpoint reading a different game afterwards with nothing
failing. So the planes ship with an explicit card_names list recorded in the run
identity: a renumber is then a REFUSAL rather than a quiet reinterpretation. If the
vocabulary ever exceeds 255, construction refuses and names the storage decision instead of
widening the dtype silently.
Ties are broken by lowest uid, and this is not cosmetic. BattleState.entities is
NOT in uid order — a live battle gives [0, 2, 4, 1, 3, 5] — so "whichever comes first"
would make the plane a function of iteration order, and two runs of one seed could differ.
uid is unique for a whole battle and never reused.
Spells in flight do not appear here, and that is structural rather than a choice:
BattleState.spells is a separate list from entities, so a live spell has no entity and
no tile. This matches the same object being invisible to a learner's committed-elixir
accounting, which is the one outcome that creates no new discrepancy between the two.
Default OFF until train's policy-head arm has run. Flipping an observation shape under
a paired comparison invalidates both arms and looks like a result. BUILDING it was never
gated, only turning it on: with the switch off the observation is byte-for-byte the shipped
one -- no new key, the same vector width and offsets, the same config().
What the switch turns on, as shipped. One switch, because it is one change to what a
network sees:
- card_ids as above (obs.card_id_planes).
- enemy_last_card in the vector, one-hot [n+1], the last card the enemy played (D2's
decision put it with the planes). It sits at the END of the fair block, so no existing fair
offset moves, and it is read from the same cycle memory enemy_possible_hand uses.
- config() records card_identity: true and card_names, the catalogue in order; a builder
constructed with card_names REFUSES to bind to an engine whose names differ, naming the
first id that moved.
The memory learns a play by diffing two hand snapshots, one per decision. That is exact because a seat plays at most one card per decision; a caller that steps twice without observing in between can let the second card -- the one that replaced the first -- enter and leave inside one gap, where no diff can see it. The env never does that; a hand-written harness can.
These planes enter at the STEM, and the trunk is unchanged (learn). card_ids is a
separate key, embedded and concatenated into the channels, so adding it widens one
convolution and adds one embedding table whose input width is a constructor argument. That
is the fact that makes landing them after the Rust port cheap rather than a rework, and it
is why the port targets today's shape: the equality harness holds both implementations to
each other, so a new plane is a change both sides make and the harness catches the
divergence.
FLIPPING THE FLAG IS A TWO-REPO CHANGE, NOT A CONFIG EDIT (learn). A new observation
KEY is the right storage decision and it is not free on the learner's side, where a new
float channel inside spatial would have cost nothing. It touches four things there: the
codec learns a third key, since it reads spatial and vector and derives mask_planes
from the mask and has no general any-key path; the buffer's row layout and row_bytes
change, and that is the number the shared-memory rectangle is sized from; ObsBatch gains
a field, and so does every site that constructs one; and the stem does the lookup and
concatenates. So the flag is off by default on BOTH sides, and turning it on is a
coordinated change rather than a switch. The cost is worth paying: the alternative was a
card id stored as a scaled half, which fails silently, and silent is worse than work.
3c. Queued: positional channels (NOT built, and frozen until train's run ends)¶
Recorded so the reasoning survives, including the part of it that was WRONG, because the wrong version is the one a reader is likely to re-derive.
The dead claim. Counting distinct per-tile feature vectors in spatial gives 11 to 18
distinct vectors across 576 tiles, with the most common covering 43% of the board. Measured
on build d872d792711934c2, blue seat, at steps 0, 20 and 60 of a played battle with 6, 12
and 15 entities on the board; the 43% held at all three. It is
tempting to conclude that a pointer head, which scores a tile by an inner product with that
tile's features, therefore has ~18 logits available and cannot separate 250 tiles by any
weights. That conclusion is false. The head's inner product is not with this array: the
learner's trunk concatenates two coordinate planes, normalised y and x, into the stem's
input, so the feature map the head sees distinguishes two empty tiles by construction. The
measurement is real and the inference on top of it is not.
What is true, in learn's formulation: what distinguishes two empty tiles is their coordinates plus whatever falls inside the receptive field, which at four residual blocks of 3x3 convolutions is roughly 9 to 11 tiles. So a tile further than that from every entity is described by its position and the static planes alone. That predicts a policy can learn "deploy at this coordinate" and cannot learn "deploy 14 tiles from that Giant" without more depth. It is weaker than the dead claim and, unlike it, checkable.
The queued channels. Distance to each tower, and lane. Worth a plane each because a 3x3 convolution stack has to spend DEPTH computing something an input plane could simply state, and depth is what this network does not have much of.
Not now, and the reason is the same as the card planes'. The observation shape is frozen until train's several-hundred-iteration run finishes: changing what the network sees mid-comparison invalidates both arms and looks like a result. These land behind the card-identity planes.
These do NOT overlap with card_ids. The card planes differentiate about 19 OCCUPIED
tiles of 576; the tiles a positional channel helps are the empty ones. Two different gaps.
Both this session and learn had been counting them as one.
4. The flat vector, float32 [12n + 43] (fair)¶
All slots are clipped to [0, 1]. The offsets in the table are for n = 16, which is
MockEngine's default catalogue and what the test suite runs on. It is not the
full card list, which is larger and gives a wider vector. Read offsets from
vector_offsets(n, reveal). Never copy a number out of this table into code.
| slots (n=16) | field | size | range | meaning | fair? |
|---|---|---|---|---|---|
| 0 | own_elixir |
1 | 0..1 | own elixir / MAX_MANA | fair |
| 1 | enemy_elixir |
1 | 0..1 | opponent's elixir / MAX_MANA (counted, see below) | fair (source switches under Reveal.enemy_elixir) |
| 2–69 | own_hand_cards |
4(n+1) | 0/1 | hand slot card one-hot; index n = empty slot | fair |
| 70–73 | own_hand_cost |
4 | 0..1 | hand slot elixir cost / MAX_MANA | fair |
| 74–77 | own_hand_affordable |
4 | 0/1 | affordable right now: the bar less own_pending_cost pays it, and it has no play waiting |
fair |
| 78–81 | own_hand_pending |
4 | 0/1 | this slot's card has a play waiting to run (command delay, RoyaleSim df69520); the seat's own taps only | fair |
| 82 | own_pending_cost |
1 | 0..1 | elixir the own waiting commands hold / MAX_MANA | fair |
| 83–99 | own_next_card |
n+1 | 0/1 | cycle position 5 | fair |
| 100–150 | own_cycle_6_8 |
3(n+1) | 0/1 | cycle positions 6, 7, 8; index n = not deduced yet | fair |
| 151–166 | own_deck |
n | 0/1 | own deck multi-hot, as deduced so far | fair |
| 167–183 | own_last_card |
n+1 | 0/1 | last card own played; index n = none yet | fair |
| 184 | own_ticks_since_play |
1 | 0..1 | ticks since own last play / 600, clipped | fair |
| 185 | own_elixir_leaked |
1 | 0..1 | own elixir lost to the cap so far / 20, clipped | fair |
| 186–201 | enemy_cards_seen |
n | 0/1 | cards the opponent has played at least once | fair |
| 202–217 | enemy_possible_hand |
n | 0/1 | cards that could be in the opponent's hand now | fair |
| 218 | enemy_plays |
1 | 0..1 | opponent's plays this match / 40, clipped | fair |
| 219–221 | own_tower_hp |
3 | 0..1 | tower hp / max, [king, left, right] |
fair |
| 222–224 | enemy_tower_hp |
3 | 0..1 | the same, in the opponent's own-frame slots | fair |
| 225–226 | crowns |
2 | 0..1 | own crowns / 3, enemy crowns / 3 | fair |
| 227–228 | king_active |
2 | 0/1 | own king active, enemy king active | fair |
| 229–231 | clock |
3 | 0..1 | regulation left / regulation, in overtime, overtime left / overtime | fair |
| 232–234 | elixir_rate |
3 | 0/1 | one-hot over 1x, 2x, 3x (triple elixir late in overtime, RoyaleSim df69520 on) | fair |
| appended | enemy_hand_cards |
4(n+1) | 0/1 | the opponent's hand | reveal (enemy_hand) |
| appended | enemy_next_card |
n+1 | 0/1 | the opponent's cycle position 5 | reveal (enemy_next_card) |
| appended | enemy_deck |
n | 0/1 | the opponent's deck | reveal (enemy_deck) |
600, 20 and 40 are presentation constants (PLAY_GAP_TICKS, LEAK_SCALE,
PLAYS_SCALE) that decide where a feature saturates. They are not game numbers and
are not read from calibration.
12n + 43, not 12n + 36¶
The specification this rewrite was built to called the width 12n + 36. It is 12n + 43, and the extra slots are real rather than accidents. The arithmetic, term by term:
| block | width |
|---|---|
| the previous layout | 5n + 30 |
own_deck |
+ n |
own_cycle_6_8 |
+ 3n + 3 |
own_last_card |
+ n + 1 |
enemy_cards_seen |
+ n |
enemy_possible_hand |
+ n |
own_ticks_since_play, own_elixir_leaked, enemy_plays |
+ 3 |
elixir_rate's 3x slot (RoyaleSim df69520's triple elixir) |
+ 1 |
own_hand_pending, own_pending_cost (RoyaleSim df69520's command delay) |
+ 5 |
| total | 12n + 43 |
The previous layout's 5n + 30 includes the one enemy-elixir slot, which is kept and now holds the count, so it is not double-counted. A test asserts 12n + 43 directly.
The counted enemy elixir¶
enemy_elixir is maintained by the builder, not read from the state, and this is the
one fair feature that needs saying carefully.
- It is seeded once, at the start of the match, from the opponent's bar. The starting amount is public, and a curriculum start that hands one side extra elixir is equally public. After that the state's copy is never read again.
- From there it is the engine's own arithmetic (
protocol.ElixirLaw, derived fromcalibration.jsonandglobals.csv): pay for each play seen, regenerate over the ticks that passed at the rate the clock says, clamp at the cap. All in integer "fine" units (one elixir islcm(regen 1x, regen 2x)of them), because milli-elixir cannot carry the law. A tick is worth 17.857… milli at the shipped numbers, and a milli-space sum drifts inside one match. - A play is a public event: a unit appears, a spell is cast. The builder reads it from the opponent's hand changing between two observed states, which names the same event and names the card exactly.
The result is bit-exact against the bar the engine keeps, which is why it belongs in
the fair set rather than being an estimate. Reveal.enemy_elixir swaps in the value
read from the state. A test plays a battle out and asserts the two agree at every
step, from both seats, on a busy game and on a quiet one (the quiet game is what
exercises the cap). Verified exact, per step, on: the opening, a start_tick that
crosses the 2x threshold, a start already in overtime, asymmetric starting elixir,
decision_ms of 50 (one tick per step) and 3000 (sixty), a randomised mid-game
curriculum, and a game quiet enough to sit at the cap.
MatchMemory.exact says when it cannot be. The same law runs on the player's own
bar, which is visible, so the count is checked every step against a number the memory
is not allowed to guess at. The moment the two disagree, exact goes False and stays
False for the match. Two things make that happen: a deck that repeats a card (a play
that swaps a card for itself changes no hand slot, so it is unseen; a real deck is
eight distinct cards), and an engine whose elixir law is not the one in
calibration.json. The builder then resyncs the own bar from the observed value so it
stops drifting. It leaves the enemy count alone, because the only way to repair it
would be to read it. A training run that wants the guarantee can assert exact.
The cycle features¶
A card played goes to the back of an 8-card cycle. Hand is positions 1–4,
next_card is 5, and 6–8 are behind it.
own_cycle_6_8starts unknown (all three one-hots point at index n) and learns one position per play, so the whole cycle is visible after three plays.own_deckfills in the same way.enemy_possible_handis the same rule applied to the opponent: the last four cards they played are exactly the four behind their hand, so those are 0 and everything else is 1. "Everything else" is the whole catalogue until eight distinct cards have been seen, at which point their deck is known and the answer narrows to it.
The builder is stateful¶
These features make the builder carry a MatchMemory per seat. Two guards keep one
episode out of the next. ObsBuilder.reset(state), which the env calls, re-seeds both
seats. And MatchMemory.observe re-seeds whenever the clock moves backwards,
which can only be a new battle. observe is also idempotent by tick, so building the
same state twice cannot drift. That is what lets the seat-flip and list-order gates
build the same state dozens of times.
Tests replay a seeded episode twice in the same env and require the two observation sequences to be identical, with a plant that removes both guards and shows the difference.
The same fields without an engine¶
Most of the vector is what a player remembers, not what the board shows. Four published names let code outside the env compute those fields with no engine running: for recorded matches, for a bot playing through a client, for any tool that has a log of plays. The env computes its own vector through the same four, so the numbers match by construction.
MatchMemory.start(tick, own_elixir_milli, enemy_elixir_milli, own_hand, next_card)begins a match from what a player sees at its start.MatchMemory.advance(tick, regular_ticks, overtime, own_plays, enemy_plays)moves it totickthrough the plays since the last call, each a(tick, card)pair. A play inside the interval splits the regeneration at its tick. The env'sobservecalls the same method, dating every play at the previous observation, because the engine pays a command before the step's first tick. Ability presses (a hero's, a champion's) go in as the optionalown_pressesandfoe_presses, each a(tick, elixir)pair: paid from the bar like a play, and no play.MatchClock.at(tick)is the clock of a match still running attick, from the rules.MatchClock.of(state)is the clock an engine reports.fair_fields(memory, clock, hand, next_card, own_elixir_milli, cards, max_mana)returns every fair field except the four the board decides (BOARD_FIELDS: tower hitpoints, crowns, awake kings), by name, asFAIR_FIELDSlists them.build_vectorcalls it for those fields.
The caller supplies what a player sees anyway: the own hand, the next card and the own bar. Rules a caller has to get right:
- A play at tick p is paid before tick p runs. An observation at tick T sees the plays made before T.
own_ticks_since_playcounts from the observation that first showed the play, not from the play. That is what the env has always recorded.MatchMemory.unaffordablecounts plays the counted bar could not pay, (own, enemy). It stays at zero on an engine's own log. Anywhere else it means a missing play or a different elixir law.
tests/test_fair_fields.py plays battles on MockEngine and RustEngine, dates every
accepted play, and compares every field with the env's vector at every step, for both
seats. It also holds MatchClock.at to the engine's clock at every step. The battles
cover a full turn of both queues, spells in flight, the switch to 2x, the exact tick
regulation ends, ticks past 4800, gaps of 1 to 13 ticks, and a seat that leaks at a
full bar. Plays dated one tick late must fail, and so must an overtime rule one tick
late. When fair_fields was split out of build_vector, every vector the builders
wrote over those battles was compared byte for byte before and after, on both engines.
The whole surface such a caller may rely on, because it is more than the four names
above. From royalegym.obs: MatchMemory(num_cards, law) with bind(cards), start,
advance and show_own_hand(hand, next_card), and the attributes tick, own_fine,
foe_fine and unaffordable; MatchClock, fair_fields, FAIR_FIELDS and
BOARD_FIELDS. From royalegym.protocol: ElixirLaw.load(calibration) and
ElixirLaw.to_milli(fine), default_calibration(), CardInfo, DECK_SIZE and
HAND_SIZE. tests/test_fair_fields.py names every one of them, so renaming or removing
any fails this repo's own suite, not only a caller's.
5. EntityListObsBuilder¶
For attention / transformer policies. entities [N, 18 + n + 1], spells [M, 14 + n],
plus the same vector, action_mask and mask_planes. Column names are in
ENTITY_FEATURE_NAMES and SPELL_FEATURE_NAMES; the full per-column meaning is in
the class docstring.
Rows are sorted by a key made of every field the row is built from, so two rows that tie are identical and the order is seat-invariant whatever order the engine listed them in.
The only reveal here is the aim point. Columns 9 and 10 are the aim of the viewer's
own spells and are 0 on an enemy row; Reveal.enemy_spell_aim appends two more
columns for the opponent's, so the space really does change width.
The sort key puts the aim last, and that is not cosmetic. Row order is observable
(it decides whose delay and hit count appear first), and the key must still name every
field a row is built from, so the aim cannot leave it. With the aim ranked early, two
enemy spells alike in everything visible came out in an order set by where they were
going. Two states differing only in two hidden aim points gave delay columns
[0.03, 0.07] and [0.07, 0.03]. With the aim last, hidden data can only order rows
whose every visible field ties, and those rows write the same numbers. A test and a
plant hold this.
6. Ground truth¶
The observation is built from BattleState, which is whatever the engine reports
(royalegym/protocol.py). Where a channel is described as matching the live game, the
comparison is against traces recorded by RoyaleLive, the client instrument that
records ground-truth traces from the real game.
MockEngine resolves spells inside a tick and models no status effects, so on it the
spell and stun channels and the spells array are always zero. The Rust engine's
tests are where those channels are exercised on real battles.