JAMIE BENNETTmrtz.dev

Blog

A thousand robot playtesters

You cannot tune a scoring system by feel. I have tried, more than once.

The problem is that you are the worst possible test subject for your own game. You know where the trick is. You know which option is secretly the strong one, because you made it that way on a Tuesday and forgot. Every round you play confirms a thing you already believe.

On a recent project we needed the scoring to be fair. Not roughly fair. Fair in a way you could defend out loud to somebody who had just lost.

So we wrote the players

Scripted agents at defined skill levels. A careful one. A reckless one. One who reacts late to everything. One who has clearly read the manual. Then we set them loose to play thousands of measured games, and recorded everything.

That turns a matter of opinion into a table you can read. Nobody has to win an argument about whether the middle difficulty feels harsh. You go and look.

Two rules that had to hold

Every scoring change had to satisfy two invariants before it was allowed to ship. A better player has to score better. And two equally good players have to come out even, whatever the scenario throws at them.

Both sound obvious. Both are startlingly easy to break. A scenario that punishes early mistakes harder than late ones quietly rewards whoever happened to draw the gentler opening, and you will never notice by hand, because by hand you play maybe forty rounds and the effect needs four hundred to show up.

What the robots catch

Mostly they catch rare things. An edge case that appears in one round out of four hundred turns up a dozen times in five thousand, with a log attached, at three in the morning while everyone is asleep.

They also catch drift. You adjust one number to fix a complaint, and something on the far side of the game quietly stops being fair. Running the whole fleet again after every change turns that from a lurking problem into a failed check.

What they cannot catch

Agents cannot tell you whether the game is any good. They have no boredom. They will play the dullest possible version ten thousand times and report that it is beautifully balanced, which it will be, in the way that a spreadsheet is balanced.

For that you still need people, ideally people who do not work for you and owe you nothing.

So the split we landed on is this. The robots decide whether it is fair. The humans decide whether it is fun. Both of those have to be true, and only one of them can be measured at scale.

Back to the blog

Written by Jamie Bennett in Vancouver.