Skip to content
Back to blog

Team Up Blog

Fighting Game Ladders: Rate the Set, Not the Game

| Team Up | 12 min read

In the FGC the unit of a result is a set, not a game — and rating games individually double-counts one observation, makes your ladder's volatility depend on the tournament format, and breaks character stats. Here's how to build a ladder that fits how the FGC actually plays.

fighting-gameselotournamentsleaderboardsguide

Nearly every ranking tool built for competitive gaming assumes a match is a match. One result goes in, ratings move, done. That assumption fits team games — a CS2 map, a League game, a Rocket League match — and it fits them so well that the assumption is usually invisible.

It does not fit the FGC. The unit of a result here is a set: first to two, or first to three, played back to back by the same two people. And the difference isn't cosmetic. Rating the individual games inside a set produces a ladder that is wrong in three specific, predictable ways — and once you see them, most of the other FGC-specific design questions answer themselves.

Games Inside a Set Are Not Independent

Elo assumes each result you feed it is an independent observation. Games inside a set aren't, and everyone in the FGC already knows why — you just don't normally say it in these terms.

By game two, both players have information they didn't have in game one. They've seen the gameplan, the setups, the escape habits. In games whose ruleset lets the loser switch character, the matchup itself can change between games. The FGC has a whole vocabulary for this — adaptation, the runback, getting figured out — and it exists precisely because a set is a conversation, not five samples of the same conversation.

That produces three concrete failures if you rate games.

It counts one observation as three to five. Two players play a single Bo5. That has produced one piece of evidence about who is better on the day. Feed the games in individually and the ratings move as though five independent contests occurred, so a single evening of one matchup moves both players up to five times as far as it should. Over a season this is how a ladder ends up dominated by whoever played the most long sets rather than whoever won them.

It makes your rating volatility a function of tournament format. This is the one that catches people out. If ratings move per game, a Bo5 bracket can shift a player's rating up to 5/3 as much as a Bo3 bracket against the same opponent — not because the result meant more, but because the organiser chose a longer format. Your top 8 (usually Bo5) then swings ratings harder than your entire pools bracket (usually Bo3). Rate sets and the format stops leaking into the numbers: one set, one result, whatever its length.

It rewards the wrong kind of dominance. A 3-0 and a 3-2 move ratings very differently under per-game rating, which feels correct and isn't. Both mean A beat B today. If you want to record that one was a beatdown, that's a stat, not a rating input — the same argument the attribution post makes about performance metrics in team games, and it applies here for the same reason.

So: one set in, one rating update out. Everything below follows from that.

Longer Sets Are Better Measurements — and You Can Quantify It

The reason the FGC runs sets at all is that a single game is a noisy read on who's better. That intuition is exactly right, and it's worth having the actual numbers, because they set your expectations for the ladder.

If a player wins any individual game with probability p, their probability of taking the set is:

Per-game win rate Bo3 Bo5 Bo7
50% 50.0% 50.0% 50.0%
55% 57.5% 59.3% 60.8%
60% 64.8% 68.3% 71.0%
65% 71.8% 76.5% 80.0%
70% 78.4% 83.7% 87.4%

A small per-game edge becomes a much larger per-set edge, which is the whole point of playing a set: the format amplifies real skill differences and suppresses variance. Note also how little Bo7 adds over Bo5 — at a 60% game rate you go from 68.3% to 71.0% for two extra potential games. That's why Bo5 is the standard for finals and Bo7 is rare: the measurement is already nearly as good as it's going to get.

There's a consequence for your ladder that surprises people the first time. Elo ratings encode win probability, so if set-level win rates are further from 50%, the rating gaps that describe your player base get wider. A player who wins 60% of games against someone is about 70 Elo above them at game level — but 68.3% of sets, which is about 133 Elo. Roughly double the spread, for the same players.

So when you switch a ladder from per-game to per-set, expect the top and bottom of the board to pull apart. That isn't inflation and it doesn't need correcting. It's the ladder finally describing set outcomes, which are the outcomes your events are actually decided by.

The practical knock-on is K-factor. You're now getting one observation per set instead of three to five, so each one has to carry more weight — but nothing like five times more, because those games were never five independent observations to begin with. If you were running something like K=24 per game, K=32 per set is a reasonable starting point: it's more per observation, and far less per evening.

Which Sets Count: The Casuals Problem

Here's where FGC advice has to diverge sharply from what's right for team games — and it's the one piece of guidance I'd give backwards from my own bracket tools post.

For a team-game league, the answer to "which matches count?" is all of them. Casual play is where match volume comes from, and volume is what makes a rating mean anything.

In the FGC, "casuals" is a term of art, and it means something specific: friendlies. Warming up. Trying the character you picked up last week. Deliberately practising a matchup you're bad at, by losing to it repeatedly on purpose. Feeding friendlies into a ladder poisons it, because half of them aren't attempts to win. A top player labbing a new character will lose sets they'd never lose in bracket, and the rating has no way to know the difference.

This is a genuine problem, because it means your obvious source of volume is off the table. An FGC ladder that only counts bracket sets is looking at maybe three to eight sets per player per week, and a lot of players enter fortnightly at best.

Two things fix it, and you want both:

  • A stakes channel outside the bracket. Sets that both players agree in advance are ranked — challenge someone, play the set, report it. This is what /challenge is for, and it's the single highest-leverage thing an FGC ladder can add, because it converts the enormous amount of play that already happens between weeklies into countable results. The key is that both players opt in beforehand, which is what separates it from friendlies.
  • Count every bracket set, including pools. Not just top 8. Pools sets are real sets played for real stakes, and they're most of the bracket's volume — a 32-entrant double elim produces around 60 sets, and the top 8 is only a handful of them.

Say the rule out loud in your server, once: bracket sets and challenge sets count; friendlies never do. Ambiguity here is what makes people stop running friendlies at all, which is much worse for a scene than a slightly thinner ladder.

Characters Are a Per-Game Attribute on a Per-Set Result

This one is a modelling trap, and it's specific enough to the FGC that nothing else on this blog runs into it.

Tracking characters is obviously desirable — matchup data, who plays what, how the tier list looks in your scene. But once the rating unit is the set and the ruleset permits switching after a loss, a single set can contain two or three characters per side. The set has one result. The characters have one game each.

Which means: don't build per-character ratings unless a set is played start to finish in one matchup. If you attribute a 3-2 set win to whichever character happened to be on screen for game five, you have assigned a five-game result to a character that played one game of it — and the rating you get out is confidently wrong rather than just noisy.

What works instead:

  • One rating per player. That's what the ladder is for, and it's the thing your seeding needs.
  • Characters as custom stats, recorded per game, displayed but not scored. Games played, games won, per matchup. That's the data you actually want when someone asks how the matchup goes in your scene, and it doesn't corrupt anything by existing.
  • Per-character ratings only for single-character locked formats, if your scene runs them. Then a set genuinely is one matchup and the attribution is clean.

The general principle is worth stating because it generalises: an attribute recorded at a finer granularity than the result can be a stat, but it cannot be a rating input.

The Structure That Fits: Weeklies On Top of a Ladder

Put together, an FGC scene wants two layers, and it's the same relationship as a tournament circuit and a ranking anywhere else — the ladder is the memory, the bracket is the event.

Underneath: a ladder that rates sets. Fed by bracket results and by opt-in challenge sets, running continuously, no start or end.

On top: the weekly. Double elimination, because everyone gets a losers run and it's what the FGC expects — and because it roughly doubles the information a bracket produces compared with single elim. Seed it from the ladder rather than from last week's placement or from memory: the ladder has watched every set anyone played since the last event, which is a far better basis than a one-or-two-set sample. Note that double elim needs a power-of-two field, so 8, 16 or 32; single elim will pad and resolve byes for you if your turnout is awkward.

And feed the results back down. A bracket set is a set. If your weekly's results don't reach the ladder, you've split your data between two systems that each have half of it.

The payoff shows up in the complaint every FGC bracket generates: I only lost to the eventual winner. Under placement points that's a raw deal — you finish where you finish. Under a rating it's handled automatically and correctly: losing a set to someone 200 points above you barely moves your number, because that outcome was expected. Over a season, the person who consistently goes out to the top seed lands where they should, and the person farming the bottom of the bracket stops climbing. That's the thing a ladder gives a scene that a stack of brackets can't.

Setting It Up

Concretely, in order:

  1. Record one result per set. Bo3 or Bo5, one entry, winner and loser.
  2. Track games and characters as stats, not as rating inputs.
  3. Set K per set, not per game — around K=32 is a sensible start for a weekly scene.
  4. Write down which sets count — bracket and challenge, never friendlies — and say it once in public.
  5. Seed your weekly from the ladder, and feed the bracket's sets back into it.
  6. Expect the rating spread to be wide. Set-level ratings separate players roughly twice as far as game-level ones. That's correct.

Frequently Asked Questions

Should a 3-0 count more than a 3-2?

Not in the rating. Both mean the set was won, and treating a 3-0 as stronger evidence reintroduces the per-game problem in a different form — a 3-0 against someone who was clearly getting figured out is not obviously more informative than a 3-2 that went the distance. Record game counts as a stat if you want the dominance story; it's genuinely interesting data, it just isn't a rating input.

What about a set that gets played out across two different days?

Rare, but the answer is the same as for any set: one result when it finishes. If it never finishes, don't record it. A forfeited or dropped set is a rules question, not a rating question, and pretending a partial set is a result is how you end up with disputes about a number.

How do I rate a team game or crew battle?

Crew battles are a genuinely different unit again — a sequence of individual sets with elimination carryover — and the honest answer is that they don't map onto a 1v1 ladder cleanly. Either record the individual sets inside them as normal sets (usually the right call, since they're real 1v1 sets for stakes) or treat the crew battle as a team result on a separate team ladder. Don't do both, or the same play counts twice.

My scene only runs a weekly. Is a ladder worth it?

If a weekly is genuinely all the competitive play that happens, a bracket tool plus placement points is a complete system and a ladder adds machinery for little gain. But in most scenes it isn't all that happens — there's a constant stream of sets in the server between events, and none of it is counted. A ladder is mostly a way to make that play matter, which tends to produce more of it.

Doesn't rating sets make the ladder converge slowly?

Slower in results-per-week, yes — you get one number per set instead of three to five. But those extra numbers were never independent evidence, so you weren't converging as fast as it looked. The real fix for convergence is volume of sets, which is exactly what an opt-in challenge flow is for.


Running a fighting game scene on Discord? Team Up rates one result per set, runs single and double elimination brackets that seed from your ladder's live Elo, and tracks per-game character stats alongside the rating without letting them distort it. Start with the leaderboard setup guide, or how to run a Discord tournament if the weekly comes first.