The learner¶
RoyaleLearn closed its training loop on 2026-09-22, and nobody has trained a bot with it yet. Both halves matter, so here they are in order.
The command runs end to end: rollouts, gradient steps, a checkpoint with its manifest, a
snapshot in the opponent pool, and a resume that reloads it in a fresh process. It needs
torch, which the plain install does not pull in. From the Royale folder:
The commands on this page run from the RoyaleLearn folder, like the last line here.
Every run so far has been a short test. An iteration is one round of playing battles and then learning from them. As of 2026-09-22 the longest run on the real engine had ten iterations, each a quarter of the usual size, and the smoke test that proves the loop closes has three. So nothing is known about how long a useful run takes, what it costs, or whether the bot it produces is any good.
Around the loop is the boring half you would otherwise write yourself, and it has been tested far longer than the loop has: the settings file and its sanity checks, the run identity that a resume is checked against, the networks, the code that turns an observation into numbers, the buffer that holds experience, the workers that collect battles, the ladder that rates one policy against another, the metrics output and the checkpoint store.
You can also use a library you already know
You can point an existing library at the environments instead of using
RoyaleLearn. They expose action_masks() in the form sb3-contrib's MaskablePPO expects.
No page here walks through that, so you would be working from that library's own docs.
What it is for¶
One job: turn battles into a bot that is actually better than the one before it.
-
It plays itself
N battles run at once as 2N player slots, so one bot learns from both sides of every match in a single batch. Both seats feed the same policy.
-
It keeps a ladder of its old selves
Past versions of your bot are kept frozen in a pool. A new version is rated against them, and only joins the pool when its win rate clears a gate. That is how you tell real improvement from a bot that only beats its own latest quirk.
-
A resumed run is the same run
A checkpoint holds the policy, the critic, the optimizer, the pool and the seeds. Reload it and the curve carries on where it was, rather than restarting somewhere nearby.
-
You can watch it happen
Set one environment variable and the viewer attaches to the training run from another window, with your loss and your ladder rating in a panel beside the board.
What runs today¶
The package imports without torch. That matters more than it sounds: the settings and the run identity are useful on a machine where you have not installed a gigabyte of deep-learning libraries, and torch is only pulled in when you ask for a name that actually needs it.
Under that, everything runs: the engine, the environments, the legal-move mask, seeding from end to end, the opponent-pool bookkeeping, and the viewer stream.
The commands¶
config works without torch; the other three need the torch extra.
From a fresh install, 2026-09-27
A fresh install on a 4-CPU Linux machine with no GPU ran these from four fresh clones.
config and doctor worked, and the smoke config trained to its checkpoint. train and
bench with the laptop profile stopped at their first collection, because the mask offered
Heal on tiles the engine refused. RoyaleGym b0948de fixed that. Those two have not been
re-run from a fresh install since.
config writes a config you can edit, doctor runs the first-run checks, and bench measures
this machine's throughput. Run those three first. The last line is the real training run: the
laptop profile on the Rust engine, with a limit of 100,000,000 timesteps.
doctor is the one worth knowing about in advance. It builds one environment, prints the engine
build fingerprint and the observation shapes, and checks every action against the engine at one
state, the one its sampled play reaches. It sees only the cards in hand then, so it can pass
while the mask is wrong for another card, as it did for Heal until 2026-09-27. It also works out how much
memory the run will need. It refuses a run that is over the
memory budget in your config, and warns when a run needs more than is free right now. bench
measures your own machine instead of quoting somebody else's.
There is also examples/train_1v1.py, which is about fifteen lines: load the laptop config,
change a few fields, run it. Weights & Biases, an online dashboard for training numbers, is off
in it. To use it, set USE_WANDB = True at the top, install the wandb extra and sign in to a
W&B account.
Changing the reward¶
This is the thing most people will want, so here is where it lives. The reward function is in
royalelearn/rewards.py and is assembled in default_potential_reward(). You change it by
writing a RewardFunction subclass, which is RoyaleGym's base class, and naming it in the
config's env block. Naming it in the config means it gets recorded in the checkpoint, so months
later the file still says what the bot was trained to want. examples/custom_reward.py shows
it: it adds one term to the shipped reward and names the result in the config.
Writing a reward function is the page for this, and it works today against the environments.
Do not re-tune the shipped weights
The shipped reward terms are built as potential differences, which is a shape that cannot change which strategy is best. It can only change how fast the bot finds it. That property is worth keeping.
If you nudge a weight every time you see a behaviour you dislike, the weight is standing in for a term that is missing. Add the term.
When you would touch it¶
- You want to train a bot properly, with self-play, a ladder and checkpoints.
- You want a different rating scheme, a different way of sampling opponents from the pool, or your metrics somewhere other than where they go by default. Each of those is a base class with a default, so you replace one without forking the training loop.
- You want to read why it is built the way it is. That is the point of
docs/design.md, which carries the reasoning rather than just the plan.
When you would not¶
| You want to | Go here instead |
|---|---|
| train with a library you already know, such as MaskablePPO | The environments |
| change what the bot wants | Writing a reward function |
| change what it sees or what its moves mean | The environments |
| fix how the battle behaves | The engine |
| watch a run | The viewer |
Two design decisions you may disagree with¶
They are written down so you can argue with them in the Discord rather than guess at them.
There is deliberately no quick version first. No throwaway trainer, no baseline learner to tide people over. The layers underneath were finished to a standard, and a half-finished harness would be the thing everyone used forever.
Nothing types a number that the environment already knows. Every shape, every width, every field position is read off a running environment at startup. The card catalogue on the maintainer's machine once went from 65 to 95 cards in a morning and no code noticed, which is exactly the point. If you ever see a page or a config with an observation width written into it, that page is already wrong.
Where the detail is¶
- RoyaleLearn's README for what is built and what is open.
docs/design.mdfor the pieces, the metric, and the conventions the harness is held to.docs/harness-spec.mdfor the specification itself.