Technology · Game design

The bot plays it
four hundred times,
so you do not have to.

A green suite proves your rules are consistent. It says nothing about whether the game is winnable, whether standing still loses, or whether any of your decisions matter. Those need a different kind of test, and an uncomfortable question every time one fails.

By Marcin Firmuga·2026-10-05·10 min read·Technology

This comes out of a management simulation with a headless core, a four to five year campaign, and 1,624 automated tests. Fifteen of those tests do something the other 1,609 cannot: they play the game.

The gap they close is the one I did not see coming. I had a suite that was comprehensively green while the game was quietly unplayable, and nothing in it was wrong. Every rule did what it said. The rules together did not make a game.

What this needs from your project. A simulation core that can run without a scene, a renderer or a frame loop — something you can step forward in a plain C# test. If your game logic only runs inside MonoBehaviour.Update, this article is an argument for separating it before it is an argument for testing it.

Why a unit test cannot see balance

A unit test asks whether a rule behaves correctly. Balance is not a property of any rule. It is a property of all of them running together for four years, which means it lives in a place no unit test looks.

The comment at the top of my playability fixture is the shortest version of the problem I have managed to write:

/// Is the game playable. /// /// Every other test asks whether a rule behaves correctly. These ask whether the rules /// together make a game somebody can win, and whether standing still loses. A simulation /// can be perfectly consistent and completely unplayable, and the only way to find out /// is to play it.

So the fixture plays it. Not a sampled approximation of play, not a statistical model of play — an actual campaign, inside the test, one simulated day at a time, with a scripted player making the decisions a person would be making.

The bot is the floor, not the ceiling

The first instinct is to make the bot good. That instinct is wrong, and getting it wrong wastes the whole exercise, because a bot that plays well only ever tells you that an expert can win. Nobody was asking that.

Mine is deliberately ordinary, and it says so:

/// A deliberately ordinary strategy. No lookahead, no intel, no clever timing. It exists /// to establish the floor: whatever a thoughtful human does should beat this, and this /// should not go bankrupt.

Its whole decision loop is ten plain steps in a fixed order, re-run every simulated day. No search, no evaluation function, nothing that a first-time player could not do by reading the screen in front of them:

private void Decide() { BalanceTheCluster(); ShipAnythingFinished(); SizeTheFleet(); RaiseWhenOffered(); PushTheTree(); AdoptBetterArchitectures(); BuyDataWhenAffordable(); KeepOptimising(); StartARunWhenIdle(); StaffTheDesk(); }

That ordinariness is what makes the assertion mean something. When this player goes bankrupt, the balance is wrong, not the player, and there is no argument to be had about whether a better strategy existed. The floor test then reads exactly as it sounds:

[Test] public void AnOrdinaryCompetentPlayerSurvivesFourYears() { var simulation = NewGame(); var operatorBot = new ScriptedOperator(simulation); operatorBot.Run(1460); Assert.That(simulation.State.IsBankrupt, Is.False, $"Went under on {simulation.State.Date} with {operatorBot.ModelsShipped} model(s) shipped."); Assert.That(operatorBot.ModelsShipped, Is.GreaterThanOrEqualTo(3), "Four years should fit several model generations."); Assert.That(simulation.State.BestCapability, Is.GreaterThan(30.0), $"Capability stalled at {simulation.State.BestCapability:0.0}."); }

Assert a band, not a number

Here is the mistake that took me longest to find, and it is the one I would fix first in somebody else's project: a balance test with only a lower bound rots silently.

Every patch makes the game a little more generous. A new research node, a rebalanced price, a bug fix that happened to remove a cost. Each change individually is defensible, the floor test keeps passing more and more comfortably, and nothing ever goes red. You end up with a game nobody can lose and a suite that is proud of it.

The fix is to write the ceiling down as well. The comment on mine states the failure mode in both directions, which is the part worth copying even if you take nothing else from this article:

/// The difficulty band. An ordinary player should finish four years in the race and not /// in front of it: still behind the frontier, still short of dominating the market, still /// solvent. If this test starts passing trivially the game has gone soft; if it starts /// failing on the low side the game has gone unfair.

In practice that is four upper bounds sitting next to the lower ones:

Assert.That(simulation.State.IsBankrupt, Is.False, context); Assert.That(gap, Is.GreaterThan(-8.0), $"An ordinary player should not be comfortably ahead of the whole field. {context}"); Assert.That(report.MarketShare, Is.LessThan(0.75), $"An ordinary player should not own the market. {context}"); Assert.That(simulation.State.UnlockedResearch.Count, Is.LessThan(ResearchTree.All.Count), $"Four years must not be enough to finish the technology tree. {context}"); Assert.That(simulation.State.HasResearch(ResearchNodeId.ArtificialSuperintelligence), Is.False, $"The end game must stay out of reach in the first four years. {context}");

Read the assertions as design decisions, because that is what they are

“Four years must not be enough to finish the technology tree” is not a technical statement. It is a decision about what the game is, written somewhere it will be enforced, which is more than most design documents manage.

This turns out to be the quiet benefit of the whole approach. The balance tests became the only place where my intentions about difficulty are written down in a form that argues back.

The control case: proving that playing matters

Every test so far can pass in a game where your decisions are decorative. A game that plays itself is still solvent, still shipping, still green. To catch that you need the experiment that every science class teaches and almost no test suite contains: a control.

Two campaigns, same seed. In one the bot plays properly. In the other, somebody ships a single thing and then does nothing at all for four years:

[Test] public void ShippingOnceAndDoingNothingElseIsPunished() { // The control case. Same start, same seed, no maintenance after the first model. var passive = NewGame(); passive.SetRentedAccelerators(500); passive.TryStartTraining(new ModelBlueprint("Only model", ArchitectureId.DenseTransformer, 20, 400, DatasetSource.WebCrawl), out _); passive.Advance(40); passive.TryReleaseModel(0, 1.0, out _); passive.Advance(1400); var active = NewGame(); new ScriptedOperator(active).Run(1440); Assert.That(active.State.BestCapability, Is.GreaterThan(passive.State.BestCapability + 10.0), "Playing well has to beat playing once by a wide margin."); Assert.That(active.State.LifetimeRevenueUsd, Is.GreaterThan(passive.State.LifetimeRevenueUsd)); }

The + 10.0 matters more than the direction. “Better than doing nothing” is a bar an idle game clears. A wide margin is the claim that attention is rewarded, and it is the difference between a simulation and a toy.

The same shape generalises, and once you see it you will write four more. Does raising money actually matter? Compare a funded run against an unfunded one. Is the architecture you unlock in year three genuinely better than the one it was built from? Run both and assert the new one wins. I have one named AHouseFamilyBeatsTheDenseBaselineItWasBuiltFrom, which is a sentence about design wearing a test's clothes.

The question every failing balance test asks

Now the uncomfortable part, and the reason I wanted to write this article rather than the tidy version of it.

When a balance test goes red you have two moves. You can change the game, or you can make the bot smarter. Both produce a green suite and both look about the same in the diff, and only one of them is honest — except that it is not always the same one, which is what makes this genuinely hard rather than merely tempting.

I hit this when the service side of my simulation started falling over. The fix was to let the bot move a slider it had been ignoring. Here is what I wrote next to it, because I did not trust myself to remember why it was allowed:

/// **This is an operator change, not a softened assertion, and the distinction matters /// because the same move can be either.** Serving used to take the whole fleet whenever /// no training run was in flight, while research went on taking its share regardless, so /// the cluster did a hundred and seventy per cent of its own work. Making the halves add /// to one is a fix; it also means a company can now genuinely starve its customers by /// researching, which this operator did, for five years, without noticing. /// /// A player watching the service dial go red gives the customers more of the fleet. That /// is one slider on the COMPUTE screen and it is exactly as available to them as it is /// here. Model the player who can see the control, not the one who cannot.

That last line is the rule I now apply, and it is a test you can actually run on yourself: would a first-time player, looking at that screen, do this? If yes, teaching the bot is modelling reality and the test gets more honest, not less. If the control only exists in the fixture, or the bot is being handed information no player can get, or the numbers are being nudged until it passes — you are not testing the game any more, you are negotiating with it.

Two details from that comment worth stealing

“The cluster did a hundred and seventy per cent of its own work.” That is a real bug, and no unit test could have seen it: every individual allocation was correct, the halves simply did not have to add up. It took a bot playing for five years to surface it.

“Which this operator did, for five years, without noticing.” Once the bug was fixed, a new legitimate failure became reachable — a company can now starve its own customers. Fixing a simulation bug does not reduce the test surface. It usually widens it.

A failing balance test has to describe the game, not the number

A unit test that fails tells you a function is wrong and you go and read the function. A balance test that fails tells you the company went bankrupt in year three, and that is almost useless on its own, because the cause is somewhere in 1,095 simulated days.

So every assertion in these tests carries a snapshot of the world with it:

var context = $"cap {simulation.State.BestCapability:F1}, frontier {report.FrontierCapability:F1}, " + $"share {report.MarketShare:P1}, cash {simulation.State.CashUsd:N0}, " + $"models {bot.ModelsShipped}, rounds {bot.RoundsRaised}, " + $"research {simulation.State.UnlockedResearch.Count}/{ResearchTree.All.Count}";

The difference this makes is the difference between a half hour of bisecting and reading one line. A failure that says share 94%, research 31/31, cash enormous is a game that got too easy. A failure that says models 1, cash negative, rounds 0 is a player who never got off the ground. Same red test, opposite problems, and the string tells you which one before you open anything.

The cheapest upgrade in this whole article

If you already have balance tests and want one improvement tonight, it is this. Put the five or six numbers that describe your game state into every assertion message in the fixture. It costs four lines and it changes what a red build means.

The bot has to be stateless, and that is not obvious

One constraint catches everybody once, including me, and it only appears when you combine balance tests with save tests — which you should, because a save that quietly changes the campaign is a balance bug that travels.

The test is: play four years, save, load, play a fifth year, and assert the loaded run matches the one that never saved. The moment you write it, the bot acquires a rule it did not have before. Anything the bot remembers in a field of its own is state that the save file does not contain, so the reloaded run will diverge — and the test will blame the simulation for a difference the fixture created.

// **Read from the desk, never from the last report.** `AYearFourSaveRunsIdentically` // rebuilds the operator after loading, so anything this decision remembers between days // is a difference between the run that saved and the run that loaded, and the guard reads // that as the simulation diverging. The queue is saved; a field on the operator is not.

So the rule is: every decision the bot makes must be derived from the simulation state, never from something the bot has been carrying. It is the same discipline that makes a campaign replayable in the first place, which is the subject of the companion article on deterministic randomness — and it is worth noticing that the constraint arrives from a direction nobody warns you about.

What this costs, honestly

These are not cheap tests, and pretending otherwise would be the kind of advice I write this site to avoid.

CostWhat it actually looks like
Runtime Each of these plays 1,460 to 1,826 simulated days, several of them twice for a comparison. They are the slowest tests in the suite by a wide margin.
Maintenance Real. By my own count in the code, the fixture has had to be taught to use a control the game grew three separate times — the parameter ceiling, the product line, and the support desk. Every new player-facing system is a possible new bot behaviour.
Judgement The one that does not reduce. Every red test asks whether you are fixing the game or teaching the bot, and no amount of tooling answers it for you.

Against that: fifteen tests cover the question “is this still a game”, they run on every change, and they found a cluster that was doing a hundred and seventy per cent of its own work. I would pay the runtime again.

Where to start, if you have none of this

In the order I would write them again:

The first two are a single evening and they are where nearly all of the value is. The rest is refinement.

Questions people ask about this

Can you unit test game balance?

Not with ordinary unit tests. A unit test asks whether a rule behaves correctly, and balance is not a property of any one rule — it is a property of all of them running together for a long time. What works is a scripted player with a deliberately ordinary strategy that plays a full campaign inside the test, after which you assert on the state it reached. A simulation can be perfectly consistent at the level of every rule and completely unplayable as a game.

How good should the bot be?

Deliberately mediocre. It exists to establish a floor, not a ceiling: no lookahead, no clever timing, no information a first-time player would not have. Then the assertion reads as if that player cannot survive, the balance is wrong rather than the player. A bot tuned to play optimally only tells you an expert can win, which was never the question.

What should a balance test assert?

A band, with both sides. One side says the game is survivable; the other says it is not trivial — the player should not be comfortably ahead of the field, should not own three quarters of the market, should not have finished the technology tree. A test with only a lower bound goes quietly green as the game drifts easier, which is the most common way balance rots without anyone noticing.

When a balance test fails, do you fix the game or fix the bot?

The real question, and the diff looks identical either way. Teaching the bot is legitimate when the game has grown a control a human would obviously use: you are modelling the player who can see the dial. It is dishonest when the control exists only in the test, when the bot gets information a player cannot, or when it is tuned until the number passes. The usable test is whether a first-time player looking at that screen would do the same thing.

How do you prove that playing well actually matters?

With a control case, and most projects do not have one. Run two campaigns from the same seed: one played properly, one that ships a single thing and then does nothing for four years. Assert the active run wins by a wide margin. Without it, every other assertion can pass in a game where the decisions are decorative.

Does this replace human playtesting?

No, and it is not trying to. It covers exactly one question — whether the numbers add up to a game that can be won and can be lost — and it covers it on every single change, which no human playtester can. Everything about whether the game is interesting, whether the screens make sense, whether a decision feels like a decision, needs people. What this buys is that people stop spending their session discovering that year three is impossible.

All of it is readable. The scripted operator, the difficulty band, the control case and the save comparison are in the public repository for Scaling Laws, in PlayabilityTests.cs, linked from this article's sources. The comments quoted here are the real ones, not cleaned up for the article.
More from this series: 1,541 tests that never start the game · designing a tycoon economy as laws · why solo projects die at the same four features.
MF

Marcin Firmuga

Solo developer · HCK_Labs · building in public

I write about what I actually shipped, with real numbers and real code, including the parts that did nothing. More: my story.