← Back to news

Playgroup Experiments: testing a fix for going first

· By Maran
Playgroup Experiments: testing a fix for going first

Hi everyone,

In July we went through 480,000 tracked Commander games to see whether going first actually matters. It does, by more than we expected. A lot of you replied to that post with possible fixes, we went over them all and picked one.

Playgroup Experiments is a new opt-in mode. Your table plays with one small rules change, we record the games, and after a few thousand of them we can say what that change did. The first experiment is running now, and we want to walk through how we picked it, including the idea we started with and threw away. That part explains why the one we landed on looks so heavy-handed.

What we found

Four player games with a single winner, 257,261 of them. If seats were fair, every seat would win a quarter of the time.

Win rate by seat, against a fair 25%
Seat 129.2%
Seat 225.7%
Seat 323.6%
Seat 421.5%
20%25% fair30%
Teal runs left of the fair line where a seat is winning less than its share.

Seat 1 wins about 36% more often than seat 4. With a quarter of a million games behind it, that isn't luck.

Two questions always come up. Does the gap close out in long grindy games? No, it gets wider; past fourteen rounds seat 4 drops further behind. Seat 4 is 3.5 points below fair, but seat 1 is 4.2 points above it, and seat 2 is already close enough to fair to leave alone. The four numbers have to add up to 100%, since somebody wins every game. You can't hand everyone a bonus. Getting to fair means seat 1 comes down 4.2, seat 2 down 0.7, seat 3 up 1.4 and seat 4 up 3.5, which adds to zero. Whatever we give one seat comes out of another.

Pricing a rules change

This is where we got stuck. Saying seat 4 needs 3.5 points is easy. Working out how much scry that is, or how much life, or how many treasure tokens, is a lot harder. Luckily we are sitting on heaps of data and we managed to find something worth using inside of it. The tracker and Playgroup Live records every mulligan, so we can look up what starting a card down actually costs.

Seat 4 win rate by mulligans taken
Kept 7
21.5%
One mulligan
22.4%
Two mulligans
19.1%
The first London mulligan is free, and slightly better than free: you looked at a bad hand, dug for a better one, and put your worst card on the bottom. The second one costs about 3.3 points.

The other seats show the same pattern, a little more gently.

So we have a conversion rate. One card in your opening hand is worth roughly 3 percentage points of win rate.

Which puts a number on the problem itself. The 7.7 point gap between seat 1 and seat 4 is about two and a half cards wide.

The scry idea, and why we dropped it

Our first plan was a mild one: scry 1 for seat 3, scry 2 for seat 4. Nobody loses anything, and it sits naturally in the game.

Priced against the ruler, it falls apart. Scry doesn't add a card, it improves the one you were already going to draw, and the usual estimate puts scry 1 at about a fifth of a card. A fifth of a card is around 0.6 points, against the 3.5 that seat 4 is missing. Our fix covered about a sixth of the problem.

Small effects are very expensive to measure. Sit four evenly matched players down for 100 games and you won't see 25/25/25/25, you'll see something like 28/23/26/23 from luck alone. That drift is around 4 points at 100 games, and it shrinks as the games pile up, but slowly.

How far a seat can drift on luck alone
1004.3 pts
4002.2 pts
1,6001.1 pts
6,4000.5 pts
Games playedDrift
Each row plays four times as many games as the one above it, and gets half the drift.

Halving the drift takes four times the games, not twice. An effect half the size costs four times as much to detect, and an effect a fifth the size costs twenty five times as much.

For the scry plan that works out at roughly 115,000 four player games played with the rule switched on. We would have run it for years and ended with a result nobody could read, because "we saw no change" would cover both "the rule does nothing" and "the rule does something too small for us to see".

What we're running instead

The player who goes first skips their draw on their first turn. The player who goes last draws twice on theirs.

We know how the first half of that reads, it feels very punishing, and we went back and forth on it before committing.

It helped that we weren't inventing anything. In two player Magic the player going first already skips their first draw step, and Commander simply never carried that rule over to multiplayer. It also aims at the seat that's out of line, which is seat 1, and a fix built only out of bonuses can't do that. And it's large enough to see: one card off seat 1 and one card onto seat 4 is close to a two card swing against a two and a half card problem, which brings the sample we need down from 115,000 games to about 3,000.

Here's what we expect to happen.

Today
Seat 129.2%
Seat 225.7%
Seat 323.6%
Seat 421.5%
Seat 1 to seat 47.7 pts
What we expect
Seat 126.5%
Seat 225.6%
Seat 323.5%
Seat 424.4%
Seat 1 to seat 42.2 pts
About 70% of the problem, out of two lines of rules text. These are predictions, not results.

There's an obvious flaw in it and we'd rather point at it ourselves than have you find it. Seat 3 gets nothing, so seat 4 goes straight past it, and if this lands where we expect, seat 3 ends up the worst seat at the table. We're doing it that way deliberately. Seat 3 is about half a card short, which is roughly what a scry 2 is worth, and changing two things at once would leave us unable to say which one did the work. Scry was the wrong size for a 3.5 point hole. It's close to the right size for a 1.4 point one, so it's first in line for the next experiment.

What counts as success

We won't be claiming that every seat sits at exactly 25%, that's impossible to do. What we want to achieve is have a real impact on the win chance for both player 1 and player 4.

Set before collection started
4.5 points or less

A few limits worth stating while we're at it:

  • The three points per card figure comes from mulligan data, and players choose when to mulligan. It's the best measuring stick we have, not a law of nature. If it's off, our predictions are off with it.
  • This is a four player result. Three and five player tables have their own gaps and will need their own fixes.
  • 3,000 games gets us "the gap clearly shrank". It doesn't get us "the gap is exactly halved", which would take around 8,000. We'll decide later whether that's worth the wait.
  • We're not petitioning anyone or campaigning to change Commander. We just thought it was strange that nobody had tested this, and we happen to have the data to do it.

Turning it on

Flip Experiments on in the lobby before you start, or in the pre game screen in the tracker.

The switch locks before the die roll, so you're deciding whether to play with the experiment before anyone knows who goes first. If you could flip it after seeing the seats, whoever landed in seat 4 would push for it and whoever landed in seat 1 wouldn't, and the data would be useless inside a week.

Please do actually skip the draw. This is the one thing that can quietly wreck the whole experiment. Drawing an extra card is easy to remember, and skipping your draw is easy to forget, particularly when it's your own draw. If seat 4 always takes the bonus and seat 1 sometimes forgets the penalty, the answer we get won't just be fuzzier, it'll be wrong, and we'd build the next experiment on top of it. Playgroup Live handles both halves for you. The tracker asks the two affected players to confirm, since it can't see your hands.

Your stats are safe

Experiment games are badged wherever they appear, so you'll always know which nights were played under a house rule. They count toward your Elo and stats as normal. We went back and forth on that one and settled it deliberately: games that count for nothing get played like they count for nothing, and then the data isn't worth having.

Watch it fill in

Everything the trial collects is public while it runs, at https://playgroup.gg/commander/turn-order/experiment. You will see how many games are in, the seats' win rates once there are 300 of them (always with their uncertainty, because at 300 games a seat is known only to about five points), and the headline gap from 1,000. The result itself is read once, at 3,000 games; the page freezes that block and never moves it again. Before that, nothing on it is a result, and we wrote the threshold down before the first game so nobody, us included, gets to call it early.

What you get for playing

Every finished qualifying game counts toward Lab Partner. The Lab Partner card sleeve unlocks at 5 games for everyone at the table, and if you do not have a supporter plan you get a free month of the Supporter Pack at 20. Games count normally for ELO and your stats; they are just badged. Nothing is tied to who wins or where you sit.

One last thing

If the fix works, should it be in the game at all? A rule that makes the table fairer while making the best seat worse might be one nobody actually wants to play with. We genuinely don't know which way we'd vote.

Turn it on, play a few games, and tell us what it felt like. Come find us on Discord: https://discord.gg/wAG8aApDU5

Johannes & Marran