Competitive Call of Duty has a structure almost nothing else on this blog shares. A CS2 series is the same game on different maps. A fighting game set is the same matchup repeated. A GB is a Hardpoint, then a Search, then a Control — three genuinely different games, played by the same eight people, producing one result.
That difference is small enough to overlook and large enough to break a ladder. Most Call of Duty communities that add a ranking bot configure it the way they'd configure one for any other shooter, and end up with a leaderboard that disagrees with what everyone in the lobby thinks happened.
There are two decisions that fix it, and they're both slightly counterintuitive.
One Series In, One Rating Out
Start with the setting, because everything else depends on it: the queue's Match Format should be Best Of 3 or Best Of 5, not Single Match.
The default treats every map as its own rated match. In Call of Duty that's wrong in a way it isn't quite wrong anywhere else, and it's worth separating the two reasons.
The generic reason is the one the fighting game post makes at length: games inside a series aren't independent observations, so feeding them in individually moves ratings two to three times as far as one evening of evidence justifies. That argument is true here too and I won't repeat it.
The Call of Duty-specific reason is different, and it's the one that matters more. The three maps aren't three samples of the same contest — they're three different contests, and a team can legitimately be excellent at one and poor at another. Rate maps individually and a Search specialist who takes the Search and loses the series gains rating on the night they lost. Do that for a season and your leaderboard's top end fills with teams that are very good at one mode, which is not what anyone means by "the best team in the league".
The series result is the only statement your league actually makes about a team as a whole. A 3-1 and a 3-0 both mean we beat them today, and the ladder should record that once.
There's a knock-on for the numbers. You're now getting one observation per GB instead of three, so each one should carry more weight — but nowhere near three times more, for the reason above. If you were running something like K=24 per map, K=32 per series is a sensible landing spot. The K-factor post has the full curve argument if you want to tune it properly.
Should Ratings Be Split By Mode?
This is the question every Call of Duty organiser eventually asks, and it looks like the same question Rocket League communities ask about 1s, 2s and 3s. It isn't, and the difference is worth being precise about because it flips the answer.
In Rocket League, playlists are a choice. A player who mains 1s can play only 1s, forever. The populations genuinely diverge, which is why per-playlist ratings earn their keep there.
In Call of Duty, the modes are compulsory and bundled. You do not enter a GB and play only Hardpoint. Every series contains one of each, in a fixed order, by rule. So a "Hardpoint rating" describes a slice of an event nobody ever competes in on its own.
Which gives a clean rule:
Split ratings when players can choose the format. Don't split when the format is a bundle they always play in full.
For Call of Duty that means one rating, on the series. Three separate mode ladders fragment a population that's already small — a scrim scene is often twelve to twenty teams — into three boards that each take three times as long to converge and none of which describe the thing you actually compete in.
Track modes as stats instead. Mode performance is real and worth knowing; it just isn't a rating. Record map-and-mode results inside the series and you get a Hardpoint win rate per team without asking the rating system to produce a number for a contest that never happens in isolation. Same argument the attribution post makes about performance metrics: if you want to record that a team is a Search team, that's a stat, not a rating input.
The one case where splitting is right: your league runs standalone mode nights — a Search-only tournament, a Hardpoint ladder that exists on its own. Then the format is a choice again, players opt into it separately, and it deserves its own leaderboard. That's a different competition, not a slice of this one.
The Inversion: Rate the Roster, Not the Player
Here's where Call of Duty advice diverges from nearly every other page on this site.
For Valorant 10 mans, CS2 pugs, League inhouses — the standing advice is rate players. Teams re-form every night, so there's nothing durable for a team rating to attach to, and rating temporary teams produces numbers about groupings that will never exist again.
Call of Duty GB communities are usually the opposite. They're built out of named fours with stable rosters that have often existed longer than any ladder you're about to create. The team is the competitive unit — it has a name, it has a captain, people identify with it, and it is the thing your league's standings are actually about.
When that's true, rating the team is not the exotic option. It's the correct one, and it buys you something specific: the attribution problem disappears entirely. You never have to decide how much of a series win belongs to the AR player versus the sub who went 22-30 on Control, because you're not making a claim about individuals at all. The team won. The team's rating moves.
The stack for that is three pieces that are designed to fit together:
/team_admin create— build the named roster. Use/team_admin set_captainwhile you're there; captains submit match results, which puts reporting in one known person's hands per team instead of whoever remembers.- Team Alias (Specific Format) at 4v4 — the rating type built for named teams created with
/team_admin, as opposed to the plain Team types which key off roster composition. /queue_team— the queue that takes predefined teams rather than individual players.
Record results with /record_match team_alias for series played outside the queue, which in most GB scenes is most of them.
When rating the roster breaks
Two failure modes, both worth pre-empting because both will happen:
Subs. A four with a stand-in is not quite the same team, and the rating won't know. In practice this is fine if subs are occasional — one game in ten barely moves a rating's meaning. It stops being fine when a team is effectively two different lineups sharing a name, at which point you either split them into two team entries or accept that the rating describes an average of both.
Roster churn between seasons. Team ratings age worse than player ratings, because a team can lose three of four players and keep its name and its number. A season rollover is the honest fix — archive the standings, start fresh — and it matters more in a team-rated league than a player-rated one.
If your scene is actually an open pickup queue where fours get drawn each night, invert all of this: rate players with Player (Specific Format) at 4v4, and the standard advice from the Valorant and CS2 pages applies unchanged. Both models can run side by side on separate leaderboards, which is what a community with a GB league and a pickup night should do.
The Map Pool Is Map-and-Mode Pairs
Small configuration point with an outsized effect on whether the veto feels right.
In most shooters the pool is a list of maps and the mode is fixed. In Call of Duty the mode is at least as much of the competitive decision as the map — banning Search on a map you're weak on is not the same decision as banning the map. So build the pool out of entries that name both halves:
Hardpoint — Skidrow
Search — Terminal
Control — Invasion
Then set Map Voting Strategy to Captains Ban (no reroll), which gives you the alternating ban both captains already expect. If your league runs a fixed rotation instead of a veto — plenty do, because the CDL map/mode set is small — a simple vote or no voting at all is fine, and the pool still earns its place by making the rotation visible in the queue message.
One caveat: the full veto strategies are available on versus and team queues. FFA and co-op queues only offer the simple vote.
Practice Should Not Touch the Ladder
The last setting, and the one communities most often discover too late.
Scrim scenes play a lot of games that aren't GBs. Warm-ups, practice against a friendly team, running a map five times to work out a rotation. If all of that feeds the ladder, the leaderboard mostly measures who was willing to warm up in public — and the teams that practise hardest get punished for it, which is a genuinely perverse incentive to build into your league.
Set Ranking & Voting Mode to one of the Unranked options on a practice queue. Results are still recorded and still appear in match history; they just don't move ratings. You get the record without the distortion.
Then say the rule out loud once, in the server: GBs count, practice doesn't. Ambiguity here is worse than either policy — it's what makes people stop reporting results at all.
Give the Ladder Somewhere to Land
Two things convert a leaderboard from a page nobody visits into something a scene pays attention to.
Tier roles. /tiers grants Discord roles by rating band, so a team's standing is visible next to their name every time they talk. In a scene organised around named teams this is most of the social payoff, and it costs one command.
Seasons with an end. An open ladder never finishes, which means nobody ever wins anything. A season rollover archives the standings and starts a fresh board. For a scrim scene, six to eight weeks is usually right — long enough for ratings to converge, short enough that a team that starts badly in week one hasn't lost the year. It's also, as above, the mechanism that stops a team rating outliving the roster that earned it.
Frequently Asked Questions
Should we rate teams or individual players?
Teams, if your rosters are stable and named — which is the usual shape of a GB community, and it makes the attribution problem disappear because you're never claiming anything about individuals. Rate players instead when fours are drawn fresh from a pickup queue each night. A community that runs both a GB league and a pickup night should run both, on separate leaderboards.
Should each map in a GB be rated separately?
No. Rate the series. Three maps in a GB are three different games producing one result, and rating them individually lets a mode specialist gain rating on a night they lost the series. Set the queue's Match Format to Best Of 3 or Best Of 5 so one rating change comes out of the whole thing.
Can we have separate Hardpoint, Search and Control ratings?
You can, but you usually shouldn't. Separate ratings make sense when players choose the format — the way Rocket League players choose a playlist. Call of Duty modes are bundled: every GB contains all three, so a mode rating describes a contest nobody enters on its own, and splitting fragments a small scene across three slow-converging boards. Track mode performance as stats instead. The exception is a genuine standalone mode competition, like a Search-only tournament, which is a different event and deserves its own leaderboard.
Can practice scrims be recorded without affecting the ladder?
Yes. Set the queue to one of the Unranked modes and results are recorded and kept in match history without moving any ratings, so a practice queue and a rated GB queue can coexist in the same server.
How do we handle a team playing with a sub?
If it's occasional, let it ride — one series in ten barely shifts what a team's rating means. If a roster is effectively two different lineups sharing a name, either register them as two teams or accept the rating describes the average. Season rollovers are the general fix for roster churn, and they matter more in a team-rated league than a player-rated one.
What about Warzone?
Warzone is a battle royale, so its result is a placement ordering rather than a head-to-head win, and almost none of this configuration transfers — you want placement-and-kill points over a block of drops instead. The battle royale setup guide covers that side, and the free-for-all scoring post covers the rating theory behind it.
Running a Call of Duty scrim scene on Discord? Team Up rates the series rather than the map, tracks named rosters as first-class teams, and keeps practice out of the ladder. See the Call of Duty setup guide for the configuration in brief, read rate the set, not the game for the series argument in depth, or how to run a competitive season for the rollover side.
