Technology · Unity

1,541 of my
1,624 game tests
never start the game.

This is not a guide to the Unity Test Framework. It is about the decision made long before the first test exists, which quietly determines whether you will be able to write the second thousand or not.

By Marcin Firmuga·2026-10-05·10 min read·Technology

Every tutorial about testing in Unity shows you how to write a test. Almost none of them tell you why your game will fight you when you try, and that part is architecture, not tooling.

Here is the shape of a suite that did not fight me, measured rather than remembered:

WhereFilesTestsNeeds the game running
Edit Mode1791,541No
Play Mode1883Yes

Ninety five per cent of them run with no scene, no MonoBehaviour and no frame. That ratio is not discipline and it is not a preference. It is a direct consequence of one rule, and the rule is the whole article.

Scope. This is written from a management simulation, where state compounds over four to five in-game years. If your game is a short authored experience with hand-placed content, read the last section first — the honest answer there is different.

One object may change the state. That is the rule.

The thing that makes a game testable is not a framework, a mocking library or dependency injection. It is that there is exactly one place where the world can change, and everything — every button, every tick, every system — goes through it.

In my project that object carries this comment, and it is the most load-bearing paragraph in the codebase:

/// The ONE thing allowed to change CompanyState. Every player action and the whole daily /// tick go through here, which is why a test can drive four years of company history /// without a scene, a MonoBehaviour or a frame. /// /// The day, in order: deliveries land, the run consumes compute, the market splits demand, /// the bills come out, the gates get re-checked. Nothing in that order is negotiable, /// because a cluster that arrives in the morning should serve tokens the same afternoon.

Read the causal claim in the middle: which is why a test can drive four years. The testability is not a feature anyone built. It is what you get for free once nothing can mutate the world behind the gateway's back.

The second paragraph is doing quieter work. A fixed, documented order of operations means a test can assert on the state after any given day and get the same answer every time, for the same reason the campaign is replayable at all — the subject of the article on deterministic randomness. Ordering and determinism are the same problem wearing different hats.

The measurement that tells you if you have it

There is a one-line check for whether your game logic is testable, and it is more honest than any amount of intention. Count how many of your simulation files import the engine:

grep -rl "using UnityEngine" --include=*.cs Assets/YourGame/Scripts/Simulation | wc -l

In mine the answer is 0, across 98 files. Not because importing the engine is forbidden by some rule I enforce, but because once the simulation only ever manipulates its own state through one gateway, there is never a reason to reach for a Vector3, a Debug.Log or a Time.deltaTime. The engine turns out to be a thing the logic never needed.

If your number is large, that is the work. Every using UnityEngine in a rules file is a small bet that this rule will only ever run inside a game, and it is the bet that later makes a four-year campaign impossible to test in under a second.

The two engine types that sneak in and cost the most

Time.deltaTime turns frame rate into a simulation input, so the same campaign on a faster machine is a different campaign. A simulation should advance in units it owns — a day, a tick, a turn — and let the presentation layer decide how fast to call it.

Vector3 and friends look harmless and drag the whole engine assembly in behind them. If your logic genuinely needs coordinates, a four-line struct of your own keeps the boundary clean and costs nothing.

Assembly definitions are the part that enforces it

A rule nobody can break by accident is worth more than a rule everybody agrees with. Assembly definitions are how Unity gives you that, and they are the step most tutorials skip because the toy example does not need them.

Four files do the job in my project:

AssemblyWhat it is for
ScalingLaws.RuntimeThe game itself: simulation, data, persistence, UI.
ScalingLaws.EditorEditor-only tooling, which must never reach a player build.
ScalingLaws.Tests.EditModeThe 1,541. No scene, no frame.
ScalingLaws.Tests.PlayModeThe 83 that genuinely need the engine.

The test assembly definition is worth reading line by line, because three of its fields are the ones people get wrong:

{ "name": "ScalingLaws.Tests.EditMode", "references": [ "ScalingLaws.Runtime", "ScalingLaws.Editor" ], "includePlatforms": [ "Editor" ], "autoReferenced": false, "defineConstraints": [ "UNITY_INCLUDE_TESTS" ], "optionalUnityReferences": [ "TestAssemblies" ] }

The references list is the quiet benefit. Because the test assembly declares exactly what it can see, an architectural mistake — a simulation file reaching into UI, say — stops being a code review question and becomes a compile error.

What the other eighty three tests are for

It would be a tidier article if the answer were “nothing, put it all in Edit Mode”. That is wrong, and the eighty three are where a specific class of bug lives.

Play Mode tests should cover anything whose subject is the engine: scene loading, prefab wiring, UI that only exists once a document has been rendered, coroutines, physics. But the category people skip is the important one — the tests that prove the headless core is actually connected to what the player sees.

Why this matters more than it sounds

A simulation with 1,541 passing tests can still ship a screen that opens empty. The logic is proven; the wiring is not. Every bug of that shape I have had looks the same from the inside: one piece of state, two places that set it, and only one of them tells the screen.

Edit Mode cannot see this by construction, because the screen does not exist there. If you write only three Play Mode tests, make them “every screen opens and shows something on day one”, “a change in the simulation reaches the screen”, and “a click on the screen reaches the simulation”.

What 1,541 tests actually turned out to be

“Sixteen hundred tests” sounds like a number designed to impress, and on its own it would be. It is more useful broken into the four jobs they do, because the fourth one is the one almost nobody has:

KindThe question it answers
Rule tests Does this one calculation do what it says? The ordinary majority, and the cheapest to write.
Consistency tests Is the content coherent? That every research prerequisite exists, that the tree has no cycles, that every gated thing has exactly one gate. These read catalogues rather than code, and they catch the mistakes you make at two in the morning in a data file.
Persistence tests Does a save survive a round trip, and does an old save still load? Each one is cheap; together they are why a player's campaign from three months ago still opens.
Playability tests Is this still a game? A scripted bot plays four full years and the assertions describe a difficulty band. Fifteen tests, covered in their own article, and they found bugs no unit test could see.

The consistency layer is the one I would add first to someone else's project. A test that walks your content and asserts it is well-formed costs an hour and never stops paying, because content is where a solo developer makes the most mistakes and has the least review.

What it costs, and when it is not worth it

I would be writing an advertisement rather than a guide if I stopped here.

The real cost is not writing the tests. It is that the architecture has to be decided early and is expensive to retrofit. Separating a simulation from the engine after the fact means finding every place where a rule reads Time.deltaTime or pokes a MonoBehaviour, and in a project of any size that is weeks, not an afternoon. I have written about the one time I had to do it, and it was not fun.

The second cost is that a headless core is a slightly less convenient way to write a game. You cannot drag a value onto a component in the inspector and have the simulation see it. Everything is explicit, which is better at month twelve and more typing in week one.

The honest case against all of this

If your game is short, authored and hand-placed — a narrative game, a puzzle game with designed levels, a platformer where the content is the scenes — most of this buys you much less. There is no four-year campaign to simulate and no compounding economy to protect. The same hours spent on levels will produce a better game.

The test that decides it: can a human play your whole game often enough to notice a regression? If yes, play it. If a full playthrough takes four in-game years and nobody will ever do that by hand after every change, the suite is not a quality ritual — it is the only instrument you have for observing what you built.

The order I would do it in again

Questions people ask about this

Can you test Unity game logic without entering Play Mode?

Yes, and most of it should be. Unity's documentation separates Edit Mode tests, which run in the editor with no game loop, from Play Mode tests, which start the game; Edit Mode uses the plain NUnit [Test] attribute and runs substantially faster. In my shipped simulation the split is 1,541 Edit Mode against 83 Play Mode, so about ninety five per cent of the suite never starts the game. The limiting factor is not the framework — it is whether your logic can run without a scene at all.

What makes game logic testable in the first place?

One rule, decided early: exactly one object may change game state, and every player action and every tick goes through it. Once that holds, a test can build the state, call the gateway and step forward without a scene, a MonoBehaviour or a frame, because nothing else can have mutated anything behind its back. The 98 simulation files that import UnityEngine zero times are a consequence of that rule, not a separate goal.

What should stay in Play Mode tests?

Anything whose subject genuinely is the engine: scene loading, prefab wiring, UI that only exists once a document is rendered, coroutines, physics. Plus the category people skip — the integration checks proving the headless core is actually connected to what the player sees. That omission is why a game can have thousands of green tests and a screen that opens empty.

Do I need assembly definitions to test Unity code?

For anything beyond a toy, yes. The test assembly sets includePlatforms to Editor only, adds TestAssemblies to its optional Unity references, and constrains itself with UNITY_INCLUDE_TESTS so the compiler leaves it out of a player build. They also enforce the architecture: a reference you did not declare becomes a compile error instead of a quiet dependency.

How long do 1,624 tests take to run?

The Edit Mode majority are fast because there is no game loop to wait for; the expensive ones are the handful that simulate four or five in-game years, and those dominate the total. The practical arrangement is to run everything on a commit and keep the long playability tests out of the loop you run every thirty seconds while editing a rule.

Is it worth testing a solo game project?

It depends on the game. For a simulation or strategy game where numbers compound over hours, the tests are the only way to observe what you are building, because nobody plays four in-game years by hand after every change. For a short authored experience the same effort buys much less, and the honest answer is to spend it on the levels.

The numbers in this article are counted, not remembered. The assembly definitions, the simulation folder and the test suite are in the public repository for Scaling Laws, linked from the sources. The grep in the second section is the one I ran to write it.
More from this series: designing a tycoon economy as laws · why solo projects die at the same four features · balance tests that prove a game is winnable.
MF

Marcin Firmuga

Solo developer · HCK_Labs · building in public

I write about what I actually shipped, with real numbers and real code, including the parts that did nothing. More: my story.