Somewhere in the second hour of a 10-man night, someone says the balance is broken. Usually after a 13-2. Sometimes with a screenshot.
It is worth checking, because the claim is testable and it is almost always wrong. Balanced team formation does one thing — find the split of ten players whose two total ratings are closest — and it succeeds at that on the overwhelming majority of nights, by a margin that is not close.
What it cannot do is anything about the ten people who joined. And that turns out to be where the entire experience of a 10-man night is decided.
The Balancer Almost Always Hits Its Target
Ten players can be split into two fives 126 distinct ways. That is a small enough number to search exhaustively, which is what the balanced-teams mode does — it does not approximate, it checks every split and takes the best one.
Here is what that produces. Twenty thousand simulated lobbies per row, ten players each, drawn from a pool whose ratings are normally distributed around 1200 with the spread in the first column:
| Pool spread (sd) | Median best sum gap | 90th percentile |
|---|---|---|
| 50 | 1 point | 4 |
| 100 | 2 | 9 |
| 150 | 4 | 13 |
| 200 | 5 | 18 |
| 300 | 7 | 26 |
Even for a very wide community — a 300-point standard deviation means your regulars run from roughly 900 to 1500 — the median night is split to within seven rating points across five players a side. Put differently, the share of nights where the best available split lands within 20 total points is 98.6% for a narrow pool, 93.0% for a middling one, and 84.7% even for the widest pool tested.
There is essentially no room for the balancer to be the problem. With ten players and 126 candidate splits, some split is always close, and the algorithm always finds it.
The Sums Balanced. The Matchups Didn't.
So what did people feel?
Total team rating is not what anybody experiences. What they experience is the specific person in front of them. So take the same simulated lobbies, sort each team by rating, and compare best against best, second against second, and so on — then record the widest of those five gaps. It is a proxy rather than a literal claim about who faces whom, but it is a much closer proxy for the felt experience than a team sum is:
| Pool spread (sd) | Best sum gap (median) | Widest rank-for-rank gap (median) | 90th percentile |
|---|---|---|---|
| 50 | 1 | 44 | 75 |
| 100 | 2 | 87 | 151 |
| 150 | 4 | 131 | 226 |
| 200 | 5 | 175 | 301 |
| 300 | 7 | 262 | 450 |
The left column barely moves. The right column scales almost linearly with the pool's spread.
A single night from the simulation makes it concrete. Ten players show up:
1312, 1294, 1267, 1098, 980, 941, 880, 852, 767, 689
The best split available is:
Team A 1267, 1098, 941, 880, 852 sum 5038
Team B 1312, 1294, 980, 767, 689 sum 5042
Four rating points between the two teams. That is as balanced as a lobby gets. And inside it, Team B's top two are 1312 and 1294 against Team A's 1267 and 1098 — a 196-point gap in the second slot — while Team A's bottom three are all comfortably ahead of Team B's bottom two.
Both teams have a genuine claim to be the favourite depending on which part of the game matters, and at least two people are going to have a bad evening. Nothing was miscalculated. There was no better split. The balancer did its job perfectly and the night is still lopsided, because a balanced sum is not a balanced game when the ten inputs are that far apart.
The Only Lever Is the Pool
Once you accept that, the list of things you can actually change gets short, and it stops including any setting on the queue.
Max Elo Difference does what it says, and that is the problem. There is a real setting — a cap on the rating gap allowed when forming a match — and on paper it is exactly the fix. In practice it only works if your pool is already narrow. The share of random ten-player lobbies whose full spread fits under a given cap:
| Pool spread (sd) | Cap 300 | Cap 400 | Cap 600 |
|---|---|---|---|
| 50 | 99.7% | 100% | 100% |
| 100 | 45.6% | 86.5% | 99.7% |
| 200 | 1.3% | 6.7% | 45.6% |
| 300 | 0.0% | 0.7% | 6.6% |
On a wide pool, a tight cap does not produce a narrower lobby. It produces no lobby. In a live queue that shows up as waiting rather than refusal — people sit in the queue while the bot holds out for a compatible ten — which is the same problem wearing a friendlier face. Set it generously or not at all unless your community is genuinely tight.
A rating floor makes a second, narrower pool. A minimum rating to join splits your community into a general queue and a higher one, and each of those has a smaller spread than the whole. This works, and the cost is the obvious one: two queues out of one population fill more slowly than one did, and the sparser of the two will sometimes not fill at all. That is the same trade divisions make, at a smaller scale and with less ceremony, and the arithmetic is worth reading there before you commit.
Who you invite is the real lever, and it is unglamorous. The single most effective thing a 10-man night has going for it is that the same twenty-five people keep showing up. Communities that grow their regular pool horizontally — more people at a similar level — get better nights. Communities that grow it vertically — anyone who asks — get a wider spread and worse nights, and then go looking for a setting to fix it. There isn't one. There is only who is in the lobby.
This is also why the same bot, with identical settings, produces great nights in one server and complaints in another. The configuration was never the variable.
A Night With No Roster Can Only Rate People
Two smaller consequences of teams that dissolve at the end of every match, both of which catch people out.
There is nothing durable for a team rating to attach to. The five players who just won are not a thing that will exist again. So a 10-man ladder rates players, which means it inherits the attribution problem in full — the result is a team result and the rating change lands on individuals who did not contribute equally to it. The mitigation is volume and re-drawing: over enough nights with enough different teammates, the noise averages out. Over four nights it does not, so do not read too much into a ladder's first month.
Head-to-head records from a queue are close to meaningless. In a balanced 10-man queue you are a given player's teammate about as often as their opponent, and every result is filtered through eight other people. A per-pair record built from that is mostly a record of who happened to get drawn together. Head-to-head is a genuinely great feature — on a challenge ladder, where the two people chose each other and nobody else was involved. In a queue, treat it as trivia.
And You Still Need Ten Of Them At Once
The other constraint is attendance, and it is the one that actually ends 10-man nights.
A queue needs its full lobby simultaneously. Eight regulars who are each free on a different evening produce zero matches, and a pool that looks healthy on a member count can be structurally unable to fill. This is why announced nights outperform always-open queues below a certain size: an announcement synchronises people, and a standing queue relies on them colliding.
If your format also has role or class quotas on top of the headcount, the arithmetic gets considerably worse and is worth understanding properly — the role queue post works through what a hard quota does to fill rates, and the short version is that the scarcest role gates the whole lobby regardless of how many people are queued.
Setting It Up
- Create the queue with
/queue_versus headlessat your team size, pointed at a leaderboard using Player (Specific Format). Usehostedinstead if your nights are announced rather than continuous. - Leave Team Formation on Balanced Teams. It is doing its job. Switch to captains because your community enjoys the draft, not because you think it will produce fairer teams — it will not.
- Leave Max Elo Difference unset until you have measured your pool, and then set it generously. A cap that stops lobbies from forming is worse than a lopsided lobby.
- Set a minimum-matches bar before a player is ranked, so a visitor's single lucky night does not appear above your regulars.
- Attach a map pool and let a vote or a captains ban resolve it — and make the pool comfortably larger than the ban count, or the same map survives every week.
- Decide who reports the result before your first night. "Everyone can" and "nobody did" produce the same ladder.
Frequently Asked Questions
Our teams still feel uneven. Is the algorithm actually picking the best split?
It searches all 126 possible splits and takes the closest one, so yes. If a night felt uneven despite that, look at the lobby rather than the split: check the gap between your highest and lowest-rated player that evening. If it is over about 400 points, the split was optimal and the night was still going to be lopsided, and no setting available to you would have changed that.
Would bigger teams help, or smaller?
Smaller teams make individual mismatches more visible, because there are fewer teammates to absorb one. Bigger teams hide them better but make each result a weaker signal about any individual. Neither changes the underlying spread. If you are choosing a team size, choose the one your game is designed around and your pool can fill.
Should we use captains instead?
Captains are a different product, not a better balancer. A draft is entertaining, it gives people a role before the match starts, and plenty of communities prefer it. But captains pick on reputation, reputation lags the ratings by about a season, and the resulting teams are measurably less even than the algorithmic split. Pick captains for the experience and accept the cost.
How wide is my pool, in practice?
Take your ranked players' ratings and look at the gap between roughly the 15th and 85th percentile — that is about two standard deviations. If that spread is under 200 points, your nights are structurally fine and any complaints are about something else. Over 500 and you have a real problem that only a second queue or a narrower invite list will solve.
If the balancer is fine, why does everyone blame it?
Because it is the only visible mechanism. Ten people join, a machine divides them, a lopsided game happens — the machine is the obvious suspect, and "the pool you built has a 600-point spread" is a much less satisfying answer than "the balancer is bad". Showing the split's rating gap when the match forms helps more than any argument: once people can see that the teams were four points apart, the conversation moves to where it belongs.
Running 10 mans on Discord? Team Up searches every possible split for the closest one, handles map veto, and updates ratings the moment a result is reported. See the 10 mans setup guide for the full configuration, or how to set up matchmaking queues if you are starting from nothing.
