FrankJones wrote: ↑30 March 2026, 19:17
My apologies for the delayed reply. (Earlier in this thread I quoted the above portion and said I had more thoughts that I would post soon).
First of all, thank you very much for writing it up and laying out your thoughts so clearly. I do enjoy these nerdy discussions and I'll try to respond to most of your points.
FrankJones wrote: ↑30 March 2026, 19:17
If 2 players have the exact same results across 100 games against the same opponents, that is literally the definition of "two equally skilled players."
No, absoulutely not! 100 games are way to few to make a definitive statement about skill or Elo for that matter. Every competitive Backgammon or Can't Stop player can tell you that stretches of 100 games are usually not representative, both mathematics and simulations show that there is a high variance in win rate in only 100 games. If player A had 200 wins out of 300 games before and player B had 100 wins out of 300 games before, and then they have equal performance in 100 games, every sane person should assume that player A is much better than player B. For games like Backgammon or Can't Stop, it simply does not make sense to rely only on the last 100 games, particularly when you have more games available.
FrankJones wrote: ↑30 March 2026, 19:17
Why should anyone care that the "historical" Elo gap of 100 is mostly erased? It should be entirely erased. Because, one of two things is the case:
1) That 100 Elo gap was artificial in the first place. Maybe each player was actually a true 650 rating, but one had a small win streak and the other had a small losing streak, putting them at 600 and 700 instead of both at 650.
or
2) The lower rated 600-Elo player has improved and is now equal in true skill and Elo to the 700-Elo player.
It is extremely flawed that you think the above two are the only options. These are possible scenarios of course but you missed an important third one that happens many times in practice. The third one is: A good player has a bad (unlucky) stretch of 100 games and a bad player has a good (lucky) stretch of 100 games, making both have similar realized performance but not similar skill. Again every experienced and Backgammon or Can't Stop player can tell you that this happens many times.
All in all you are simply underestimating variance way too much.
FrankJones wrote: ↑30 March 2026, 19:17
Either way, one purpose of Elo and k-factors is to get players accurately rated, and as quickly possible, right? (Hance the higher temporary k-factors during a player's first 11 and 21 games played). So, the results you showed, in which higher k-factors closed this 100-Elo-gap more quickly - to me, that is not a flaw in k-factor; that is the k-factor working efficiently to correct inaccurate Elos.
If the only consequence of current K factors was that two players with similar performance in the last 100 games have the same Elo, I wouldn't mind. (Particularly if those last 100 games were weighted equally.) I don't think it's accurate but it would be good enough. Unfortunately you missed my main point of my old simulation post.
My main point: It is not about what happens after 100 games. It is about what happens after 15 or 30 games! My complaint was about the shape of the curve, not about the endpoint after 100 games. With K=40, the original Elo gap of 100 is halfed already after only 15 games and an Elo gap of only 20 is reached after only 30 games! That's just insane nonsense for games like Backgammon or Can't Stop. 15 or 30 matches are nothing. With K=20, the original Elo gap is halfed after 30 games and is reduced to only 20 after 60 games. This is much better already but still nonsense. The problem really is how strongly the most recent games are weighted for the Elo calculation. The last 10 games are 5 times more important than game 61 to 70 before now. And even within the last 10 games, the last 2 games are much more important than game 9 and 10 before now. This effect is modulated with the K factor, that's what I wanted to show with my simulations.
(The strong weighting of most recent games is plausible for complex games with steep learning curves (Terra Mystica, Through the Ages) but not for Lucky Numbers et al.)
FrankJones wrote: ↑30 March 2026, 19:17
I took a quick glance at lifetime Elo and Arena scores for "Earth" And "Terraforming Mars" For both games, the spread in Elo (from highest to the floor of 100) is similar to the spread in Arena points (from highest to the lowest shown, keeping in mind that the truly lowest arena scores might not be visible because the actual arena score is hidden for any player who is not yet purple elite.)
Spreads are not good indicators if the rating system works good or bad. The spread doesn't tell you if the right people are in the right positions. Plausible spreads also happen with bad rating systems.
FrankJones wrote: ↑30 March 2026, 19:17
So, despite the double k factor of Arena (40) compared to standard games (20), the arena results seem to reasonably mirror the Elo results, despite the fact that Arena scores get reset every 3 months and despite the lower overall player pool.
I will show the difference of K=40 vs. K=20 again in my next post.
FrankJones wrote: ↑30 March 2026, 19:17
One more thought. A lower k-factor could make it take longer for players who
should have a wide Elo gap to actually attain that gap. Which means, there is a longer delay in getting players accurately rated, and a higher number of games adversely affected by the fact that one or both players have an inaccurate Elo.
This is a very valid concern. It could be solved by having higher K factors in your first, say, 100 games before we end up with K=10. Let's say after the first 20 games with K=60 / K= 40, the next 30 games are at K=30, the next 50 games are at K=20 and only after 100 games we start with K=10. Today I also think that my original suggestion of K=5 was reaching too low. Maybe having K=10 as lowest possible K factor is good enough, still we could have a system with 3 categories, K=10, K=15, K=20.
In any case, you desire a quick alignment of calculated Elos with "true Elos", however, this is not possible in the first place for Lucky Numbers et al. High K factors make rating accuracy much worse for everyone for whom a lot of data available and it only makes rating accuracy a little bit better for those with small data available.
FrankJones wrote: ↑30 March 2026, 19:17
In either case, over a large enough sample, the results should end up where we expect.
No, absolutely not! This is one of the most common myth that I hear about the Elo system. In reality, Elos do not "
converge" (and the less so, the higher the K factor). I will show that in my next post. To me, convergence means that a series of values tends towards a certain value
and then stays at that or close to that value. This is intrinsically impossible with the current Elo system. In reality, the Elo values inevitably bounce around in a wide margin. It is true that the range where the Elo value bounces is related to your "true Elo" but when your Elo values bounce around by +/-100 which is usual, you have no real chance to know where your "true Elo" really is. And it is not getting better over time because Elo cares so much about most recent results. Only if the skewed sample with overweight of recent games has some sense (e.g., in Terra Mystica), then your Elo is very meaningful, but this skewed sample is just causing an inescapable Elo bouncing in high-luck games.
FrankJones wrote: ↑30 March 2026, 19:17
And, yes, higher-luck games have higher variance and therefore require more games to meet the criteria of "large enough sample". Some may view this as a problem; I do not.
Here you are contradicting yourself now. If you already agree that larger sample sizes are needed for higher-luck games, then you should favor lower K factors for those games. Only with lower K factors, the "large enough samples" make their way into the Elo calcualation. With current K factors, the realized sample sizes are way too short and heavily skewed towards most recent games. Above you said the quick Elo adjustment with high K factors are not a problem but now you say we need large sample sizes. You can't have both. I will show more on this in my next post.