Every league owner eventually gets both of these complaints, often in the same week:
"I won six in a row and barely moved."
"I lost one game to a smurf and dropped 40 points."
Those are the same complaint. They are the two ends of a single dial, and turning it to satisfy one person guarantees the other one shows up. That dial is K-factor.
Most explanations of K stop at "it's the maximum points a match can move." True, but it doesn't help you choose a number. What you actually need to know is that K controls two things you care about — how fast a new player reaches their real rating, and how much everyone's rating wobbles forever after — and it moves them in opposite directions.
This post puts numbers on that tradeoff. Every figure below was generated by running the same rating code that runs live matches, so you can reason about your own league from them. If you want the mechanics of the formula itself first, start with the Elo Rating Calculator guide; this post picks up where its K-factor table leaves off.
What K Actually Does to the Curve
One line of the rating formula matters here:
delta = K x (actual_score - expected_score)
expected_score is a probability between 0 and 1, so (actual - expected) is bounded — it can never exceed 1. K is a pure multiplier on top of it. It doesn't change who gains points or which direction the rating moves; it only scales how far each match travels.
The consequence is worth stating plainly, because it's the thing people get wrong: K does not change where a player ends up. A genuine 75%-win-rate player converges to the same rating whether K is 16 or 64. Their rating sits where their expected score matches their real win rate, and that point is set by the win rate and the influence range — not by K.
equilibrium = opponent_rating + 400 x log10(p / (1 - p))
For a 75% player against a 1000-rated field, that's about 1191 at any K. What K changes is how long the trip takes and how much they bounce around once they arrive. Those are the four curves in this post's header image: one motion, four speeds.
The Convergence Table
Here's the part that's usually hand-waved. I simulated a player whose true skill gives them a fixed win rate against a 1000-rated field, starting at the default 1000, and measured two things across 4,000 trials per K:
- Matches to converge — the median number of matches before their rating first lands within 25 points of their true equilibrium.
- Steady-state swing — once converged, the standard deviation of their rating across their next 200 matches. This is permanent noise. It never decays.
For a genuine 75% winner (equilibrium ≈ 1191):
| K-Factor | Matches to converge | Steady-state swing | 95% of the time they're within |
|---|---|---|---|
| 16 | 78 | ±27 | ±53 |
| 24 | 49 | ±37 | ±72 |
| 32 (default) | 35 | ±45 | ±88 |
| 48 | 21 | ±58 | ±114 |
| 64 | 15 | ±70 | ±137 |
For a milder 65% winner (equilibrium ≈ 1108), convergence is faster because the trip is shorter, but the swing column is identical — 27, 37, 45, 58, 70. That's the tell: swing is a property of your K, not of the player. Every player on your leaderboard, at every skill level, carries the noise in that column for as long as your league exists.
Read the two middle columns together and the tradeoff is stark. Going from K 32 to K 64 more than halves the time to converge (35 → 15 matches) and buys it by inflating everyone's permanent noise from ±45 to ±70. At K 64, two players 100 points apart on your leaderboard are, quite often, indistinguishable in skill — the gap is noise. At K 16 the leaderboard is precise, but a genuinely improving player spends 78 matches climbing to where they already belong, which in a weekly-scrim community is most of a season.
That's the whole problem: you want a big K for new players and a small K for established ones.
The Fix: A Provisional K Curve
Which is exactly what provisional rating does. Instead of one K for everyone forever, a player's first N matches use an amplified K that fades back to your steady-state value. New players sprint to their real rating; the leaderboard stays quiet afterward.
Same simulation, same 75% player, now with a base K of 32 and provisional rating enabled:
| Configuration | Matches to converge | Steady-state swing |
|---|---|---|
| Flat K 32 | 35 | ±45 |
| Flat K 64 | 15 | ±70 |
| K 32 + ×2 over 10 matches (step) | 22 | ±45 |
| K 32 + ×2 over 10 matches (linear) | 29 | ±45 |
| K 32 + ×2 over 10 matches (smooth) | 31 | ±45 |
| K 32 + ×3 over 15 matches (linear) | 13 | ±45 |
The last row is the point. A ×3 boost over 15 placement matches converges faster than flat K 64 (13 matches vs 15) while every settled player keeps flat-K-32 stability. The swing column is unchanged across every provisional row, because the boost has expired by the time it's measured. The tradeoff the first table describes isn't a law — it's what you get for using one K for everyone.
Team Up gives you three fade curves, and they differ in how much of the boost lands early:
step— full multiplier for all N matches, then a hard cliff to normal. Fastest convergence of the three (22 matches above), and the most jarring: a player's 11th match moves half as far as their 10th for no visible reason.linear— straight line from the multiplier down to 1.0. The sane default.smooth— logarithmic ease-out. Biggest correction on match one, gentle tail. Feels best to the player; converges slowest of the three.
Configure it on your league's Rating Defaults page in the dashboard, under Provisional Rating. It's off by default, and the defaults when you enable it are ×2.0 over 10 matches on the linear curve. The preview chart on that page draws the same fade curve the live engine applies, so the "Match 1: K 64 → after 10 matches: K 32" labels are the real numbers your players will see.
One thing to know: the provisional boost is per player, based on each player's own match count. In a match between a 3-game newcomer and a 300-game veteran, the newcomer moves on a boosted K and the veteran moves on the base K. The match is deliberately not zero-sum. That's standard for provisional systems, and it's the right call — the veteran shouldn't eat a 60-point loss because their opponent happens to be new.
Picking Your Number
K is not a quality setting; it's a match-volume setting. The right question isn't "how competitive is my league," it's how many matches does a typical player log per season? Convergence costs matches, and if your players don't have the matches to spend, a low K just means your leaderboard is permanently wrong.
| Your community | Suggested base K | Reasoning |
|---|---|---|
| Daily queues, hundreds of matches per player | 16–24 | Matches are abundant, so slow convergence is free. Buy the precision. |
| Active league, a few matches per week | 32 | The default. ~35 matches to converge is a month or two of play. |
| Weekly scrims, 10–30 matches per season | 48 | A season is shorter than K 32's convergence window. You need the speed. |
| Tournament or one-off event | 48–64 | Ratings must be useful within a handful of games or they never will be. |
Then, whatever you picked, turn on provisional rating — it's close to free. It shortens onboarding for every new player without touching the stability of the leaderboard you just tuned. The suggested-K column above assumes you have not enabled it; with it on, you can generally sit one row lower (a weekly-scrim league can run K 32 with a ×3 boost instead of a flat 48) and get a more precise leaderboard for the same onboarding speed.
Set the base K from either surface:
- Discord —
/leaderboard_config basics set k_factor:48for one leaderboard, or/server_config elo set k_factor:48for the server-wide default. - Dashboard — the K-Factor field on your league's Rating Defaults page, and per-leaderboard overrides on the leaderboard's own settings.
Leaderboard settings override server defaults, so it's fine to run a precise K 16 ranked ladder and a fast K 48 event ladder in the same community.
Three Ways K Interacts With Other Settings
K doesn't act alone, and a couple of these surprise people.
Rating capping changes what K multiplies. Capping clamps the rating difference used in the expected-score calculation to your cap range before the delta is computed. That bounds how lopsided (actual - expected) can get, which means capping and K both shape the size of extreme results. If you've enabled capping and matches still feel swingy, K is the dial, not the cap.
Max advantage can zero a delta outright. If a player is further above their opponent than your max-advantage setting allows, a win pays nothing — regardless of K. Raising K will not make those matches pay out. Check whether max advantage is what's actually silencing them.
Streak Boost scales the delta, not K. It's applied after the rating delta is computed, based on the streak the player carried into the match. So it stacks on top of whatever K produced — including a boosted provisional K. A new player on a hot streak in a league with both features on can move a long way in one match. Worth a look at your numbers before enabling both.
And one non-interaction that matters: K has nothing to do with catching smurfs. A high K makes a smurf climb out of your low ranks faster, which is sometimes mistaken for detection. It isn't — the damage happens on the way up either way. That's what farming and smurf detection is for.
Changing K on a Live League
Editing K takes effect on the next match. It does not retroactively change the history that produced everyone's current rating, so for a while your leaderboard is a mix of two regimes — old spacing from the old K, new movement from the new one.
If you want the change applied consistently, replay the history: on the leaderboard's Danger Zone page, run Recompute Elo ratings. That re-derives every rating from stored match history using your current settings. Two caveats before you press it:
- Ratings will move, sometimes a lot, and players will notice. Announce it first.
- The replay can only use the match history you still have. On plans with a history limit, matches past that limit are gone, so the recompute starts from a truncated record rather than the true beginning.
For most leagues the honest move is to change K and let it apply going forward. Recompute is for when you've made a real configuration mistake and want it erased, not for routine tuning.
Try It Before You Ship It
The fastest way to build intuition is to stop reading and go move the number. Open the Elo Rating Calculator, set up a match between two ratings you actually have on your leaderboard, and step K through 16, 32, and 64. The calculator runs the same rating module as the live engine, so what it shows you is what your players will get.
Then check the preview chart on your Rating Defaults page with provisional rating switched on. Seeing the boosted curve collapse onto the flat one is the moment the tradeoff in this post stops being abstract.
FAQ
What K-factor should I use?
Start at 32. Lower it toward 16–24 if your players log hundreds of matches and you want a precise leaderboard; raise it to 48–64 if a season is only 10–30 matches per player. Then enable provisional rating regardless — it fixes the onboarding half of the problem without costing you leaderboard stability.
Does a higher K-factor make ratings more accurate?
No — it makes them faster and noisier. A player converges to the same rating at any K, because equilibrium is set by their win rate, not by K. Higher K reaches that rating in fewer matches but leaves a permanently larger swing around it: ±45 at K 32 versus ±70 at K 64, forever.
How many matches does it take for a rating to settle?
At the default K 32, a genuine 75% winner lands within 25 points of their true rating in about 35 matches. K 16 takes 78; K 64 takes 15. With a ×3 provisional boost over 15 placement matches, base K 32 gets there in about 13 — faster than a flat K 64, without the permanent noise.
Can I use different K-factors for different leaderboards?
Yes. Leaderboard settings override the server-wide default, so a precise ranked ladder and a fast-moving event ladder can run different K values in the same community. Set the default with /server_config elo set k_factor: and override per leaderboard with /leaderboard_config basics set k_factor:.
What happens if I change K-factor on an existing leaderboard?
New matches use the new K immediately; existing ratings are untouched, so the leaderboard temporarily mixes old and new spacing. To apply it to history, run Recompute Elo ratings from the leaderboard's Danger Zone — but expect visible rating movement, and note the replay is limited to the match history your plan retains.
Is provisional rating the same thing as Streak Boost?
No. Provisional rating amplifies a player's K-factor for their first N matches, then fades to normal. Streak Boost scales the finished delta based on the win or loss streak carried into the match, and applies at any match count. They stack, so enabling both can produce large single-match swings for a new player on a run.
