Skip to content
Back to blog

Team Up Blog

Why the Match That Would Settle It Never Gets Played

| Team Up | 11 min read

On a challenge ladder the participants pick their own opponents, and Elo prices a lopsided match so badly that one of any two players always wants to refuse. The standings look fine and the comparisons that matter never happen. Here's the arithmetic, and the rules that fix it.

eloleaderboardsleague-managementrating-theoryguide

Here is a ladder that looks healthy. Forty active players. Two hundred matches recorded this season. The top of the table is stable, the ratings are spread sensibly, nobody is cheating.

And the first and second place players have never played each other. Not once. Neither has the second and the fourth. Ask around and you will find that number two is quietly certain they are better than number one, and that nobody can prove otherwise, because the match that would prove it keeps not happening.

Nothing went wrong. No one did anything dishonest. This is the normal, expected behaviour of a ladder where the participants choose their own opponents — and it is worth understanding before you launch one, because the fix is a rule you write on day one, not a setting you turn on afterwards.

Elo Adjusts for Who You Played, Not for Who You Picked

The league-tables post makes the case for ratings over points on the grounds that Elo is schedule-independent: it adjusts for opponent strength, so an unbalanced fixture list does not distort it the way a points table does. That is true, and it is the single strongest argument for running a ladder.

It also has a boundary condition that nobody states, because in most formats it never binds. Elo adjusts for who you played. It assumes you did not get to pick them.

In a queue, you don't. In a scheduled league, you don't. In a bracket, you don't. In all three the schedule is imposed on you, and the schedule-independence property does exactly what it says.

On a challenge ladder you pick. And the moment picking is allowed, the arithmetic that makes Elo fair starts working against the thing you wanted it to measure.

Every Challenge Is a Bet, and the Ladder Prices It

Run the numbers through the live rating engine at the default K of 32. Here is what a challenge is worth to the higher-rated player, by rating gap:

Rating gap Their win chance If they win If they lose Wins needed to undo one loss
0 50.0% +16.0 −16.0 1.0
100 64.0% +11.5 −20.5 1.8
200 76.0% +7.7 −24.3 3.2
300 84.9% +4.8 −27.2 5.6
400 90.9% +2.9 −29.1 10.0
800 99.0% +2.9 −29.1 10.0

Read the last column as what it is: a price. Challenging someone 200 points below you means putting three good results at risk to gain one. At 400 points it is ten. A month of careful laddering can be erased by one bad night against someone you were expected to beat.

Two details worth pulling out of that table.

It stops getting worse past 400 points. The engine caps how lopsided a result is allowed to be — the effective gap is clamped to the influence range, 400 by default — so a 400-point mismatch and an 800-point one are priced identically. That is a deliberate mercy: without it, beating someone 800 points below you would be worth a rounding error and losing to them would be catastrophic. It does mean that once a challenge is clearly lopsided, making it more lopsided changes nothing.

The same table read from below is its exact mirror. For the lower-rated player, a 400-point challenge is +29.1 to win and −2.9 to lose. Free option. All upside.

So Exactly One of Any Two Players Wants to Refuse

Put those two halves together and the failure mode writes itself.

For any pair of players separated by a meaningful gap, the match is a good bet for one of them and a bad bet for the other, and the size of the asymmetry grows with the gap. The person who wants the match is the person the ladder says shouldn't win it. The person who could grant it has ten wins' worth of rating at risk and three points to gain.

So the challenge goes unanswered. Not out of cowardice, and usually not out of any conscious decision — it just never quite happens, in the way that things which are individually unappealing never quite happen. Both players are behaving completely reasonably. The ladder is what created the incentive.

The result is a ladder where the matches that do get played are the ones both parties found tolerable — which means matches between people close together, which means the ordering within each cluster is well tested and the ordering between clusters rests on almost nothing.

The Matches Nobody Wants Are the Ones Worth Most

This is the part that turns an annoyance into a real measurement problem.

How much a match teaches you about who is better depends on how uncertain its outcome was. A coin flip is maximally informative; a foregone conclusion tells you nothing you didn't already believe. That intuition is exactly right, and it is quantifiable — the information a result carries about the rating scales with the variance of the outcome, which peaks at an even match and collapses as the gap widens.

Rating gap Information, relative to an even match Matches needed to equal one even match
0 1.00× 1
200 0.73× 1.4
300 0.51× 2
400 0.33× 3
600 0.12× 8
800 0.04× 26

Twenty-six matches against someone 800 points below you carry as much information as one match against your equal.

So the incentive and the information point in exactly opposite directions. The matches the ladder makes attractive are the ones it learns nothing from, and the matches it learns everything from are the ones it has priced out of existence.

Two Unbeaten Players Never Separate

Here is the same problem at its sharpest, computed by running the real engine.

Take two players who both start at 1600. Both play the same eight opponents — a field spread from 1500 down to 1150 — and both go 8-0. Sixteen matches. Every one of them a win.

After all sixteen, the two players are 1.2 rating points apart, and which of them is ahead depends only on the order they happened to play the field in. Run it with the orders reversed and the sign flips.

That is not a bug. It is the ladder correctly reporting that it has no idea. Nothing in those sixteen matches distinguishes the two, because they beat the same people and beating people you were supposed to beat is worth almost nothing.

Now let them play each other once. A single even match moves them 32 points apart — more than twenty-five times the separation their sixteen matches against the field produced, from one game.

That is the whole argument for making the top of your ladder play itself. One match between the two candidates is worth more than an entire season of both of them farming the field, and it is precisely the match that neither of them has any incentive to arrange.

This Is Not Farming, and Detection Will Not Catch It

It is tempting to file this under manipulation, and it is worth being clear that it is not.

Match farming is somebody deliberately grinding easy opponents to inflate a number, and it has fingerprints: collapsed opponent diversity, implausible match velocity, one pair dominating a player's gains. Those are detectable, the bot watches for them, and a moderator can act.

What this post describes has no fingerprints at all. Every player has a normal spread of opponents. Nobody is playing at 3am. Nobody's gains are concentrated anywhere. The distortion is not in any individual's behaviour — it is in the set of matches that were never played, and an absence does not trip a detector.

Which is why the answer is structural. You cannot monitor your way out of this one.

Fix It With Rules, Not Settings

Four rule shapes, in rough order of how much they buy you. Any one of them beats none. Write whichever you pick into your league rules in a single sentence, publicly, before the season starts.

A range you must accept within. "You must accept a challenge from anyone within five places of you." This is the strongest of the four because it targets the problem directly: it makes the near-miss comparisons — the ones with the most information in them — compulsory, while leaving genuinely lopsided matches optional. Ladder-climbing formats have used it for decades.

A decline cooldown. "You cannot decline the same opponent twice in a row." Cheap to administer, and it kills the specific pattern where one person is permanently unavailable to one other person.

An activity floor. "One accepted challenge a week to stay ranked." This converts declining from free into expensive without anyone having to accuse anybody of anything, which is its real virtue — it is a rule about participation, not about character, so enforcing it costs no goodwill.

A must-accept window. "Once a month, the top four play a round robin." Heavier, but it is the only one that guarantees the comparison at the top actually gets made. Many ladders run this as a monthly event and seed it from the standings.

The product settings support these rather than replace them:

  • Minimum matches to be ranked and an inactivity threshold (both under /leaderboard_config ranking) are what make an activity rule enforceable — a player who stops accepting drops off the board on their own, with no moderator required.
  • Anti-farming alerts (/leaderboard_config farming) are the backstop for the dishonest version, not this one. Leave them on anyway.
  • Maximum rating advantage is worth knowing about: set it and a favourite gains nothing for beating someone beyond that gap. It removes the last reason to hunt easy matches, at the cost of making those matches feel pointless to the stronger player — which is either the point or a problem, depending on your community.

Setting It Up

  1. Decide what an entrant is. Players challenge with /challenge, which takes one opponent for a 1v1 or up to four for a free-for-all. Team ladders build rosters with /team_admin, arrange between captains, and file with /record_match team_alias.
  2. Write the acceptance rule down before the first match, and put it where people will see it. This is the whole intervention.
  3. Set a minimum-matches bar so a five-match hot streak doesn't top the table, and an inactivity threshold so a good rating cannot be banked.
  4. Rate the set, not the game, if your matches are played as best-of series — see rate the set, not the game for why.
  5. Publish head-to-head records, not just the table. It is the most-read page a ladder has, and it is how everyone notices an ordering that no result has ever tested.
  6. Run something periodic at the top — a monthly round robin, a seeded bracket — and feed its results back into the same ladder.

Frequently Asked Questions

Doesn't Elo already handle this? I thought it adjusted for opponent strength.

It does, and the ratings you end up with are not biased — a player who only beats weaker opponents converges to roughly the right number eventually. The problem is not accuracy, it is resolution and speed. Selection strips out the informative matches, so the ladder takes far longer to learn anything, and the specific comparisons everyone cares about are the ones it learns least about.

Should a declined challenge be recorded as a loss?

No, and the bot does not offer it, deliberately. A decline is usually just a bad evening, and scoring it makes the ladder punish having a life. Route it through an activity requirement instead — count what was played rather than penalising what was refused — which gets you the same behaviour with none of the arguments.

How many matches before a challenge-ladder rating is worth showing?

More than you would need in a queue, because the opponents were chosen rather than drawn. Ten is a reasonable floor for showing a player at all. But volume is only half of it: twenty matches against three opponents says less than ten against ten, so watch opponent variety at least as closely as match count.

Won't a must-accept rule just make people stop playing?

It can, if the range is too wide. "You must accept from anyone" is a rule that asks your best player to spend their evenings in matches they cannot win anything from, and they will simply leave. Keep the compulsory range narrow — the neighbours, not the whole table — because the near-miss matches are where the information is anyway. Beyond that range, optional is correct.

We're small. Everyone plays everyone already. Do I need any of this?

Not yet, and that is worth saying plainly: on a ladder of eight regulars who all play each other, selection has nowhere to hide and none of these rules earn their overhead. The problem appears at the scale where you can no longer keep the whole table in your head — somewhere around twenty-five or thirty active entrants — which is also, unhelpfully, the point at which adding rules feels like bureaucracy. Add the activity floor first; it is the one that costs nothing.


Running an open ladder on Discord? Team Up handles direct challenges, team-versus-team ladders, head-to-head records and the activity and anti-farming settings above. See the challenge ladder setup guide for the full configuration, or the Elo calculator to price a specific matchup on your own settings before you write your rules.