This comes out of a management simulation with a headless core, a four to five year campaign, and 1,624 automated tests. Fifteen of those tests do something the other 1,609 cannot: they play the game.
The gap they close is the one I did not see coming. I had a suite that was comprehensively green while the game was quietly unplayable, and nothing in it was wrong. Every rule did what it said. The rules together did not make a game.
MonoBehaviour.Update, this article is an
argument for separating it before it is an argument for testing it.
Why a unit test cannot see balance
A unit test asks whether a rule behaves correctly. Balance is not a property of any rule. It is a property of all of them running together for four years, which means it lives in a place no unit test looks.
The comment at the top of my playability fixture is the shortest version of the problem I have managed to write:
So the fixture plays it. Not a sampled approximation of play, not a statistical model of play — an actual campaign, inside the test, one simulated day at a time, with a scripted player making the decisions a person would be making.
The bot is the floor, not the ceiling
The first instinct is to make the bot good. That instinct is wrong, and getting it wrong wastes the whole exercise, because a bot that plays well only ever tells you that an expert can win. Nobody was asking that.
Mine is deliberately ordinary, and it says so:
Its whole decision loop is ten plain steps in a fixed order, re-run every simulated day. No search, no evaluation function, nothing that a first-time player could not do by reading the screen in front of them:
That ordinariness is what makes the assertion mean something. When this player goes bankrupt, the balance is wrong, not the player, and there is no argument to be had about whether a better strategy existed. The floor test then reads exactly as it sounds:
Assert a band, not a number
Here is the mistake that took me longest to find, and it is the one I would fix first in somebody else's project: a balance test with only a lower bound rots silently.
Every patch makes the game a little more generous. A new research node, a rebalanced price, a bug fix that happened to remove a cost. Each change individually is defensible, the floor test keeps passing more and more comfortably, and nothing ever goes red. You end up with a game nobody can lose and a suite that is proud of it.
The fix is to write the ceiling down as well. The comment on mine states the failure mode in both directions, which is the part worth copying even if you take nothing else from this article:
In practice that is four upper bounds sitting next to the lower ones:
Read the assertions as design decisions, because that is what they are
“Four years must not be enough to finish the technology tree” is not a technical statement. It is a decision about what the game is, written somewhere it will be enforced, which is more than most design documents manage.
This turns out to be the quiet benefit of the whole approach. The balance tests became the only place where my intentions about difficulty are written down in a form that argues back.
The control case: proving that playing matters
Every test so far can pass in a game where your decisions are decorative. A game that plays itself is still solvent, still shipping, still green. To catch that you need the experiment that every science class teaches and almost no test suite contains: a control.
Two campaigns, same seed. In one the bot plays properly. In the other, somebody ships a single thing and then does nothing at all for four years:
The + 10.0 matters more than the direction. “Better than doing
nothing” is a bar an idle game clears. A wide margin is the claim that
attention is rewarded, and it is the difference between a simulation and a toy.
The same shape generalises, and once you see it you will write four more. Does raising
money actually matter? Compare a funded run against an unfunded one. Is the architecture
you unlock in year three genuinely better than the one it was built from? Run both and
assert the new one wins. I have one named
AHouseFamilyBeatsTheDenseBaselineItWasBuiltFrom, which is a sentence about
design wearing a test's clothes.
The question every failing balance test asks
Now the uncomfortable part, and the reason I wanted to write this article rather than the tidy version of it.
When a balance test goes red you have two moves. You can change the game, or you can make the bot smarter. Both produce a green suite and both look about the same in the diff, and only one of them is honest — except that it is not always the same one, which is what makes this genuinely hard rather than merely tempting.
I hit this when the service side of my simulation started falling over. The fix was to let the bot move a slider it had been ignoring. Here is what I wrote next to it, because I did not trust myself to remember why it was allowed:
That last line is the rule I now apply, and it is a test you can actually run on yourself: would a first-time player, looking at that screen, do this? If yes, teaching the bot is modelling reality and the test gets more honest, not less. If the control only exists in the fixture, or the bot is being handed information no player can get, or the numbers are being nudged until it passes — you are not testing the game any more, you are negotiating with it.
Two details from that comment worth stealing
“The cluster did a hundred and seventy per cent of its own work.” That is a real bug, and no unit test could have seen it: every individual allocation was correct, the halves simply did not have to add up. It took a bot playing for five years to surface it.
“Which this operator did, for five years, without noticing.” Once the bug was fixed, a new legitimate failure became reachable — a company can now starve its own customers. Fixing a simulation bug does not reduce the test surface. It usually widens it.
A failing balance test has to describe the game, not the number
A unit test that fails tells you a function is wrong and you go and read the function. A balance test that fails tells you the company went bankrupt in year three, and that is almost useless on its own, because the cause is somewhere in 1,095 simulated days.
So every assertion in these tests carries a snapshot of the world with it:
The difference this makes is the difference between a half hour of bisecting and reading one line. A failure that says share 94%, research 31/31, cash enormous is a game that got too easy. A failure that says models 1, cash negative, rounds 0 is a player who never got off the ground. Same red test, opposite problems, and the string tells you which one before you open anything.
The cheapest upgrade in this whole article
If you already have balance tests and want one improvement tonight, it is this. Put the five or six numbers that describe your game state into every assertion message in the fixture. It costs four lines and it changes what a red build means.
The bot has to be stateless, and that is not obvious
One constraint catches everybody once, including me, and it only appears when you combine balance tests with save tests — which you should, because a save that quietly changes the campaign is a balance bug that travels.
The test is: play four years, save, load, play a fifth year, and assert the loaded run matches the one that never saved. The moment you write it, the bot acquires a rule it did not have before. Anything the bot remembers in a field of its own is state that the save file does not contain, so the reloaded run will diverge — and the test will blame the simulation for a difference the fixture created.
So the rule is: every decision the bot makes must be derived from the simulation state, never from something the bot has been carrying. It is the same discipline that makes a campaign replayable in the first place, which is the subject of the companion article on deterministic randomness — and it is worth noticing that the constraint arrives from a direction nobody warns you about.
What this costs, honestly
These are not cheap tests, and pretending otherwise would be the kind of advice I write this site to avoid.
| Cost | What it actually looks like |
|---|---|
| Runtime | Each of these plays 1,460 to 1,826 simulated days, several of them twice for a comparison. They are the slowest tests in the suite by a wide margin. |
| Maintenance | Real. By my own count in the code, the fixture has had to be taught to use a control the game grew three separate times — the parameter ceiling, the product line, and the support desk. Every new player-facing system is a possible new bot behaviour. |
| Judgement | The one that does not reduce. Every red test asks whether you are fixing the game or teaching the bot, and no amount of tooling answers it for you. |
Against that: fifteen tests cover the question “is this still a game”, they run on every change, and they found a cluster that was doing a hundred and seventy per cent of its own work. I would pay the runtime again.
Where to start, if you have none of this
In the order I would write them again:
- The floor. An ordinary bot plays a full campaign and does not go bankrupt. One test, and it will fail the first time you run it.
- The control. The same seed played passively, asserting the active run wins by a wide margin. This is the one that proves your game is a game.
- The ceiling. Add upper bounds to the floor test until it describes a band. Write the design decision into the assertion message.
- The save. Play, save, load, keep playing, assert the two runs agree. Expect to make the bot stateless to get there.
- The context strings. Four lines, and every future failure explains itself.
The first two are a single evening and they are where nearly all of the value is. The rest is refinement.
Questions people ask about this
Can you unit test game balance?
Not with ordinary unit tests. A unit test asks whether a rule behaves correctly, and balance is not a property of any one rule — it is a property of all of them running together for a long time. What works is a scripted player with a deliberately ordinary strategy that plays a full campaign inside the test, after which you assert on the state it reached. A simulation can be perfectly consistent at the level of every rule and completely unplayable as a game.
How good should the bot be?
Deliberately mediocre. It exists to establish a floor, not a ceiling: no lookahead, no clever timing, no information a first-time player would not have. Then the assertion reads as if that player cannot survive, the balance is wrong rather than the player. A bot tuned to play optimally only tells you an expert can win, which was never the question.
What should a balance test assert?
A band, with both sides. One side says the game is survivable; the other says it is not trivial — the player should not be comfortably ahead of the field, should not own three quarters of the market, should not have finished the technology tree. A test with only a lower bound goes quietly green as the game drifts easier, which is the most common way balance rots without anyone noticing.
When a balance test fails, do you fix the game or fix the bot?
The real question, and the diff looks identical either way. Teaching the bot is legitimate when the game has grown a control a human would obviously use: you are modelling the player who can see the dial. It is dishonest when the control exists only in the test, when the bot gets information a player cannot, or when it is tuned until the number passes. The usable test is whether a first-time player looking at that screen would do the same thing.
How do you prove that playing well actually matters?
With a control case, and most projects do not have one. Run two campaigns from the same seed: one played properly, one that ships a single thing and then does nothing for four years. Assert the active run wins by a wide margin. Without it, every other assertion can pass in a game where the decisions are decorative.
Does this replace human playtesting?
No, and it is not trying to. It covers exactly one question — whether the numbers add up to a game that can be won and can be lost — and it covers it on every single change, which no human playtester can. Everything about whether the game is interesting, whether the screens make sense, whether a decision feels like a decision, needs people. What this buys is that people stop spending their session discovering that year three is impossible.
PlayabilityTests.cs, linked from this article's sources. The comments quoted
here are the real ones, not cleaned up for the article.