dschingis27 wrote: ↑29 March 2026, 09:28
FrankJones wrote: ↑28 March 2026, 18:18
I'm not sure a lower K-factor will help. To me, 20 is already a low k-factor. Anyway, a lower k-factor would create a tighter bunching among players who perhaps should not be tightly bunched.
A lower K factor would help for sure and it would not create a tighter bunching among players. The mathematics and simulations are pretty clear. Here is a simulation I made a while ago:
https://boardgamearena.com/forum/viewto ... 38#p222238
Note that the K factor does not influence the calculated expected win probabilities itself, no matter the K factor, the Elo difference between players has the same meaning. The K factor only influences how strong a single game result changes the Elo score of one player. If single games influence the current Elo score less, then more games over longer time determine the current Elo score more equally. That doesn't mean that the Elo scores between different players cannot spread out over time, it just takes longer for them to spread.
My apologies for the delayed reply. (Earlier in this thread I quoted the above portion and said I had more thoughts that I would post soon).
So, if I understood your simulation correctly, you had the following premises:
1) Two players have Elo ratings 100 Elo apart. (For example, 600 and 700, though the actual numbers do not matter).
2) Those two players then each play 100 games against the same opponents and have the same results.
The conclusions you then drew were:
1) With a higher k-factor, that 100-Elo-Gap disappears more quickly. (That's not so much a conclusion you drew, but rather, an inarguable result demonstrated by your simulation. I'll trust that the math is correct, especially since this result is what I would expect anyway based on how Elo works.)
2) This represents a flaw in Elo and/or the k-factor in some cases.
That second conclusion is where I strongly disagree, because I think your simulation illustrates the exact opposite point.
Let's re-examine the first two premises, especially the second premise.
If 2 players have the exact same results across 100 games against the same opponents, that is literally the definition of "two equally skilled players." Two players who achieve the exact same results against the same players
should have the same Elo, and
should not have an Elo gap of 100.
Why should anyone care that the "historical" Elo gap of 100 is mostly erased? It should be entirely erased. Because, one of two things is the case:
1) That 100 Elo gap was artificial in the first place. Maybe each player was actually a true 650 rating, but one had a small win streak and the other had a small losing streak, putting them at 600 and 700 instead of both at 650.
or
2) The lower rated 600-Elo player has improved and is now equal in true skill and Elo to the 700-Elo player.
Either way, one purpose of Elo and k-factors is to get players accurately rated, and as quickly possible, right? (Hance the higher temporary k-factors during a player's first 11 and 21 games played). So, the results you showed, in which higher k-factors closed this 100-Elo-gap more quickly - to me, that is not a flaw in k-factor; that is the k-factor working efficiently to correct inaccurate Elos.
Other thoughts:
I took a quick glance at lifetime Elo and Arena scores for "Earth" And "Terraforming Mars" For both games, the spread in Elo (from highest to the floor of 100) is similar to the spread in Arena points (from highest to the lowest shown, keeping in mind that the truly lowest arena scores might not be visible because the actual arena score is hidden for any player who is not yet purple elite.)
So, despite the double k factor of Arena (40) compared to standard games (20), the arena results seem to reasonably mirror the Elo results, despite the fact that Arena scores get reset every 3 months and despite the lower overall player pool.
So, I'm not seeing a problem here either.
One more thought:
I feel as though one point being made in this thread is essentially a matter of circular reasoning. As such, I reject the argument as invalid.
The argument goes like this:
[Some player] says, "I'm an expert level 500-Elo player, but because of the k-factor or Elo system or both, I am not able to maintain an expert-level 500 Elo, thus Elo or k-factor is flawed".
Isn't that an obvious circular reasoning fallacy? One of the premises is being assumed without proof and then used as part of the conclusion.
I'll use myself as an example. "Earth" is a long slow grind to 700 Elo. Few players reach that level. I want to be one of those players. I have stagnated in the high 600s, reaching as high as 674 and currently at 664.
Now, I could say, "This Elo system is unfair. When I win, I gain 3 Elo, and when I lose, I lose a lot of Elo, so much that I cannot get back the Elo I lost even if I win 7 in a row. This is making it impossible for me to become the 700-rated master player I believe myself to be."
But that is making a lot of assumptions.
What my results demonstrate is that I have not played well enough to attain a 700 Elo rating, and if I wish to be a 700-Elo player, I need to play better. I need to avoid losses to 300-Elo players, and I need to play .500 or better against other master-level players. (I had two such chances recently and lost both; though overall, I have satisfied this part of the requirement.)
One more thought. A lower k-factor could make it take longer for players who
should have a wide Elo gap to actually attain that gap. Which means, there is a longer delay in getting players accurately rated, and a higher number of games adversely affected by the fact that one or both players have an inaccurate Elo. I'm not really thrilled about playing against a player who is actually truly master level but has an Elo of only 500 because the low k-factor is causing it to take forever for that player to reach 700. So, we play, and I lose, and in doing so, I lose more Elo than I
"should", so now my Elo is artificially lowered, causing other players I subsequently beat to lose more Elo than they "should", etc. I mean, it works both ways.
In either case, over a large enough sample, the results should end up where we expect. And, yes, higher-luck games have higher variance and therefore require more games to meet the criteria of "large enough sample". Some may view this as a problem; I do not.
Furthermore, one could make the case that games with more than "x" amount of luck are just not possible to "master" because the amount of luck means it's just not possible to win often enough to ever reach 700 Elo. Is that a problem? I do not see this as a problem. It is not necessary for every game to have an identical Elo ceiling. I like the fact that different games have different Elo ceilings (as shows by the colorful chart posted by the actuary), because that gives valuable information about the amount of luck involved in the game.
(I would further argue that these results give a better indicator of the skill-to-luck ratio than arbitrary values assigned by the game designer or developer). I've seen games here on BGA for which the developer assigned numerical values for luck and skill that I flatly disagree with.