The environments¶
This is where you live. Everything you will actually write when you make a bot, you write against RoyaleGym. It takes the engine underneath and turns a battle into the shape a learning library expects: you get an observation, you pick a move, you get a reward.
It speaks the two standard Python interfaces for this, Gymnasium and PettingZoo, so the code looks like every tutorial you have already read.

That is a whole battle between two players picking at random from the moves that are legal. It is the program further down this page, watched in the viewer.
The five pieces¶
A match is made of five decisions. Each one is a small Python class, and each one already has a version that works, so you can ignore four of them and still have a bot.
| The decision | The class you subclass | What ships in the box |
|---|---|---|
| What your bot sees | ObsBuilder |
SpatialObsBuilder: a picture of the board plus a list of numbers. Also EntityListObsBuilder |
| What its moves mean | ActionParser |
TileActionParser: 2305 moves, one per card-and-tile pair, plus waiting. Also HalfTileActionParser at finer resolution |
| What it is rewarded for | RewardFunction |
default_reward(): winning 1.0, crowns 0.2, tower damage 0.1, elixir trades 0.02, added up |
| How a match starts | StateMutator |
DefaultStateMutator: a fresh battle. Also mid-game, a board you set up by hand, a saved snapshot, or a weighted mix of those |
| When the match ends | DoneCondition |
GameOverCondition for the real ending, StepLimitCondition and TickLimitCondition for cutting an episode short |
Most people change exactly one of these: the reward. That is Writing a reward function, and it is a few lines.
The last row has a wrinkle worth knowing before it bites you. There are two ways an episode can stop, and they are not the same thing. Terminated means the match is genuinely over and the next state is worth nothing. Truncated means you cut it short and the next state still mattered. Mixing them up quietly ruins every value estimate near the cut, so the shipped conditions declare which one they are, and the environment refuses one in the wrong slot.
Two APIs, and which one you want¶
Both drive the same battle. The difference is how many seats you are filling.
Use ClashParallelEnv, the PettingZoo parallel API. You hand it a move for each
player and it hands you back an observation, a reward and a done flag for each player.
This is what you want for self-play, where one bot learns from both sides of every match.
ClashSelfPlayVecEnv runs N battles at once as 2N player slots, so a single batch carries
both sides of all of them.
Use ClashGymEnv, the plain Gymnasium API. You are one player. The other seat is filled
by an Opponent you pass in, and you never see its moves.
This is what you want when you are starting out, when you want a fixed opponent to measure against, or when you want to hand the environment straight to a library that only speaks single-agent Gymnasium.
There is a ready-made opponent in the box either way. RandomLegalOpponent picks at random from
the moves that are legal right now, and NoopOpponent never plays anything. So you have
something to train against from the first minute.
A whole battle, both seats¶
This is the complete program. Nothing but royalegym and numpy.
import numpy as np
from royalegym import (ClashParallelEnv, DefaultStateMutator, RandomLegalOpponent, RustEngine)
engine = RustEngine()
by_name = {c.name: c.card_id for c in engine.cards()}
deck = [by_name[n] for n in ("Knight", "Archer", "Giant", "Minions",
"Fireball", "Cannon", "Zap", "Musketeer")]
env = ClashParallelEnv(engine=engine,
state_mutator=DefaultStateMutator(decks=[deck, deck]))
obs, info = env.reset(seed=0)
rng, policy = np.random.default_rng(0), RandomLegalOpponent(noop_prob=0.7)
while env.agents:
actions = {a: policy.act(obs[a], obs[a]["action_mask"], rng) for a in env.agents}
obs, reward, terminated, truncated, info = env.step(actions)
s = env.battle_state
print(f"winner {s.winner} crowns {[p.crowns for p in s.players]} tick {s.tick}")
Blue is player 0 and Red is player 1. Tick 3600 is the full three minutes, and it was one crown
each then: Red took Blue's left princess tower at tick 1800 and Blue took Red's left one at tick
3580. So the match went to overtime, where the first crown wins, and nobody took one before it ran
out at tick 6000. Level on crowns, the tiebreak decided it: play stopped, the board was cleared
down to the crown towers, and from tick 6067 every tower lost the same health each tick. Blue's
weakest tower had the least left, 796 against Red's 2240, so it fell first, at tick 6123, and Red
won 2-1 (re-run 2026-10-02 on engine build 6f18fbeda772dfad, RoyaleSim 0.1.4, commit 9ee48a3,
with the 15.535 card table).
One env step is half a second of game time, which is 10 ticks. That battle was 613 steps, the full five minutes and six seconds of tiebreak, so each player made 613 decisions. It takes under a second of real time.
Name the deck, and name it card by card
Notice the deck is looked up by name and not by number. A card id is only a position in the catalogue, and the positions move between card tables, so the same number is not the same card on two machines.
If you leave the deck out entirely, each side is dealt eight random cards from whatever catalogue your machine built. The same seed then gives you a different battle from the one above. Naming eight cards makes the battle a function of the program.
One seat, and the env plays the other¶
Same deck and seed, Gymnasium shape. You are Blue. Red is the random-legal opponent.
import numpy as np
from royalegym import (ClashGymEnv, DefaultStateMutator, RandomLegalOpponent, RustEngine)
engine = RustEngine()
by_name = {c.name: c.card_id for c in engine.cards()}
deck = [by_name[n] for n in ("Knight", "Archer", "Giant", "Minions",
"Fireball", "Cannon", "Zap", "Musketeer")]
env = ClashGymEnv(agent="blue", # you are Blue
opponent=RandomLegalOpponent(noop_prob=0.7), # Red is scripted
engine=engine,
state_mutator=DefaultStateMutator(decks=[deck, deck]))
obs, info = env.reset(seed=0)
print("legal moves on the first step:", int(obs["action_mask"].sum()), "of 2305")
rng, me = np.random.default_rng(0), RandomLegalOpponent(noop_prob=0.7)
total, steps, done = 0.0, 0, False
while not done:
obs, reward, terminated, truncated, info = env.step(me.act(obs, obs["action_mask"], rng))
total += reward
steps += 1
done = terminated or truncated
print(f"steps {steps} reward {total:.3f} terminated {terminated}")
Only the wait is legal on the first step, because a match refuses every deploy for its opening seconds. Play opens after 9 steps, four and a half seconds in, and this hand then has 1318 legal moves.
Blue won. 614 steps is the three minutes of normal time, the two of overtime and the tiebreak:
each side took one princess tower in normal time, nobody scored in overtime, and the tiebreak took
Red's weakest tower first, a second crown for Blue. terminated True says the match ended for real rather than being cut short. The reward is positive because default_reward() pays 1.0 for a win and charges 1.0 for a
loss, with the crown and tower terms on top.
Swap me.act(...) for your own policy and that loop is a training loop with the learning taken
out. Run it twice and you get the same numbers, because the seed fixes everything.
The exact numbers depend on the engine build, the deck and the card table your machine built, so
treat them as "this ran", not as constants. These were run on 2026-10-02 on engine build
6f18fbeda772dfad, RoyaleSim 0.1.4, commit 9ee48a3, with the 15.535 card table.
The legality mask, which is the part people like¶
Your bot is told exactly which moves are legal before it picks one. Every observation carries an
action_mask with one entry per move. It never wastes a decision on a card it cannot afford or a
tile it is not allowed to deploy on.
Legality is a property of the pair, the card and the tile together. Elixir depends on the card. Where you may place depends on the card too: spells go anywhere, buildings never go into the enemy pocket, troops need your own territory. A mask that could only say "this card" or "that tile" separately would happily suggest putting a Knight on the enemy king. This one cannot.
What it covers: elixir, territory, water, the river band, the bodies of enemy buildings, and the no-deploy rectangle around each living enemy crown tower. Your own buildings and princess towers do not block a troop: the engine moves a troop or a Heal tapped on one off it, as the game does, and the mask offers that tap, unless there is nowhere to move it. The no-deploy block around your own king still refuses.
On the first step of a battle only the wait is legal, because a match refuses every deploy for its first 90 ticks. That is why the program above prints 1 there. After that the number moves with your hand, your elixir and your deck.
Three practical notes.
- The mask arrives as
int8, which is what PettingZoo'sparallel_api_testand Gymnasium'sDiscrete.sample(mask=...)expect.action_masks()returns it as booleans, which is what sb3-contrib's MaskablePPO calls for. mask_planesis the same mask reshaped into a grid, for a network with convolution layers.- If you send a move the engine refuses anyway, it becomes a wait and is reported in
info["deploy_status"]. Nothing explodes.
The mask is worked out from the board on the Python side, completely separately from the engine. The test suite then compares it against the engine's own ruling for every card in the default catalogue, for both seats and every move, on four boards: the opening, one with buildings down, one with a princess tower gone, and one with both. So a wrong mask for any card on those boards fails a test instead of quietly poisoning a training run.
What your bot can and cannot see¶
By default your bot sees what a person watching the match could write down, with one known gap: a unit invisible to its enemy, such as a Royal Ghost, is still shown to the enemy seat. Hiding or marking it waits on a measurement of what a player sees in the real game. It is not fixed.
Your bot does not see the opponent's hand. It does get the opponent's elixir, as a count kept from the plays it watched, the way a player counts it in their head.
You can turn a hidden thing on with a Reveal, for a curriculum or for debugging. Doing so makes
the observation wider rather than filling in blanks, and ClashParallelEnv.config() records
that you did it. So a checkpoint always says whether the bot was allowed to cheat. Every channel
and every slot is in
docs/observation-spec.md.
Never write down an observation width
The size of the observation depends on how many cards your machine's catalogue holds, and that changes. Read the shapes off a running environment instead of typing them into your code. Everything in the project does it that way for the same reason.
When you would touch it¶
- You want a different reward. This is the common one. See Writing a reward function.
- You want your bot to see something extra, or see it differently.
- You want moves at a different resolution, or a different action space entirely.
- You want matches to start somewhere other than 0-0, for instance from a damaged mid-game board or from a saved snapshot, to build a curriculum.
- You want to record a battle and re-run it later, or check that it replays identically.
When you would not¶
| You want to | Go here instead |
|---|---|
| fix a card that behaves wrong | The engine |
| add a mechanic the battle does not model | The engine |
| run PPO, keep checkpoints, rate policies against each other | The learner |
| look at what your bot did | The viewer |
Speed, and the honest caveat¶
The durable figure is a ratio: the Rust engine runs 1.16 to 1.38 times the pure-Python stand-in, measured by alternating the two inside one process so that whatever the machine is doing it does to both. An env step is one decision for each player.
Absolute rates on the maintainer's laptop have ranged from 859 to 1812 env steps per second, depending on the hour and what else was running. At 10 ticks a step, that is roughly 8,600 to 18,100 three-minute battles an hour in one process. Take any single figure as an illustration rather than a target.
Stepped directly, the same engine does tens of thousands of ticks a second. The gap is Python: on every single step it builds both players' observations and both players' legal-move lists.
That gap is the project's first open item, and the fix is to move the default observation into the engine. Until that lands, most of the time it takes to play a battle through an environment goes into building observations and legal-move lists rather than into the battle itself. Nobody is pretending otherwise.
That is also why the Rust engine is only 1.16 to 1.38 times as fast as MockEngine, the
pure-Python stand-in, and not the landslide you might expect. The Python around both engines is
most of the cost.
Two more things that are easy to miss¶
Working from source, you do not need the Rust engine to start. royalegym imports without
it. RustEngine() then raises an error that names the install page and the build command, and
MockEngine runs the whole API in the meantime. MockEngine reads the 2018 card tables in a
RoyaleSim checkout, so it only works from source, not in a pip install. It is a stand-in and not a second simulator: spells resolve instantly, there are no
stuns or knockbacks, and cards run at base level. Fine for writing code, wrong for judging a bot.
The learner trains on these environments, as of 2026-09-22, and no bot has
been trained with it yet. The environments also expose action_masks() in the form
MaskablePPO expects, so you can point an existing library at them instead if you prefer.
Where the detail is¶
- RoyaleGym's README for the whole picture.
docs/architecture.mdfor the layers, the engine contract, the action space and the module map.docs/observation-spec.mdfor every channel and every vector slot.docs/background.mdfor what is publicly known about the game's rules.
Ready to build one? Your first bot