Future Future League

I like games. I like complicated games. I have a min/max pathology. Optimizing a play sequence is often more fun than actually winning for me.

I also like dragons, wizards, and rogues. I’m a sucker for a good heroic quest.

Magic: The Gathering (MtG) has long been my game of choice. It’s the original trading card game: players build custom decks from their collections and pit them against one another.

It’s also absurdly complex. Hidden information, random draws, thirty years of accumulated design - tens of thousands of cards, each capable of interacting with thousands of others in ways nobody planned for.

Wizards of the Coast doesn’t try to eliminate that complexity, it tries to bound it. Formats like Standard limit competitive play to the most recent sets, narrowing the legal pool to a few thousand cards and a metagame of a few hundred that actually see play.

That solves the problem for players. It creates a different one for game development.

They can’t ask whether a new card is balanced today and forever. A plurality of game interactions and deck designs exceed the theoretical and require practical observation (and even then often require intervention by banning cards in Formats). They have to ask whether it’s balanced in a version of Standard that won’t exist for another year or years.

So Wizards built the Future Future League (FFL): an internal playtest league where R&D plays tomorrow’s Standard using cards that haven’t been printed yet.

The name is literal, and it’s a correction. There used to be a plain Future League looking roughly six months out. As head designer Mark Rosewater put it, the Future League was disbanded because “six months only allowed us to recognize the problems, not fix them” - enough lead time to spot the disaster, not enough to redesign the card, retest it, and still make the printing schedule. So they pushed the horizon out another six months and doubled the word (a naming convention I respect enormously).

For a long time I read that as a story about foresight. Build a room, move into next year early, find out what breaks. I’ve since decided that’s the least interesting part of it.

Playtesting finds a card is broken. Then what?

Someone changes its rules, or cuts it from the set, before the file goes to the printer. That happens because three things are true at once:

  1. the play-testers sit inside development and can edit the card file directly
  2. there’s a print date forcing a decision by a fixed day, and
  3. nobody has to be convinced that a broken competitive format is a problem because the league is functioning play. Not a simulation of the format - the format, played by people trying to win with cards that don’t exist yet.

Take away any one of those and the league becomes an expensive way of generating opinions.

R&D without the mechanism to edit is just window dressing.

Brainstorm card art “Brainstorm” illustrated by Christopher Rush

Once you’re grading R&D on outcomes rather than opinions, the examples you’re evaluating with hit different.

The sandbox is the visible half, and the cheap half. What costs something is the wiring back - who is obliged to act on the finding, and by when.

I’m not neutral about any of this.

Zero, our R&D team at Metalab, is more or less the FFL idea with a payroll: we run our work under next year’s assumptions, and we are not protected from delivery pressure. Never have been, and I’d turn it down if it were offered.

Findings from a team with nothing at stake mostly describe a team with nothing to listen to.

The wiring back is where I’d bet most R&D attempts fail.

Without stakes no one is obliged to absorb what the R&D project taught them. That leaves persuasion - a talk, a Slack post, a deck - to do the convincing, and persuasion loses to deadline pressure roughly always (I’ve been on both sides of this, including the side quietly ignoring the finding).

So the version worth building is structural, not social.