Jellby wrote: ↑23 April 2026, 17:59
Unrelated to the quote above:
The elo model assumes that if A beats B 64% of the time and B beats C 64% of the time (difference 100), then A must beat C 76% of the time.
Similarly, if A beats B, and B beats C 95% of the time (difference 500), then A must beat C 99.7% of the time.
For a game where this is not true (for instance because there's always a 1 in 20 chance that you'll draw the "you win" card), elo cannot be an accurate model. We can discuss whether it's a good enough approximation, or whether the assumption is actually true for a game like chess...
Perhaps in that scenario, the nature of the game itself (too much luck) makes it impossible for anyone to win at a high enough win percentage to maintain and specific high Elo.
Perhaps, in such a high-luck game, the player who is "expected" to win x% of the time only achieved a high-enough Elo to create that situation by going on a bit of a lucky streak that is not sustainable.
I feel as though we are back to the same circular argument:
1) A player in a high-luck game goes on a hot streak, reaches 500 Elo.
2) That player says, "I'm a 500-Elo player because I have reached 500."
3) Because that Elo is not actually sustainable, but rather is the peak achievable Elo, that player's Elo then swings back downward.
4) That player says, "Elo for this game fails, because it does not allow me to stay at 500, therefore it does not give accurate Elos."
(Jellby, I'm not saying you are making this argument. You acknowledged there is a discussion to be had.)
The above argument has flaws. There are one or more fallacies in the argument. One or more premises are flawed or rely on assumptions. Specifically, if the system is being declared to be broken or inaccurate, then we don't
know that a player who reaches 500 is actually truly a 500 Elo player.
So, if that player then cannot sustain that Elo, we don't
know that Elo is failing. Maybe it is actually self-correcting what was an unsustainably high Elo resulting from a streak of good luck.
In other words, the proof that Elo is broken relies on an initial assumption that ... Elo is accurate. The conclusion of the proof ends up contradicting a premise. That is a fallacious argument.
There is a high-skill game here at BGA that also has a lot of luck. Despite having a lot of skill, the game can swing on one lucky outcome of an RNG card reveal. The nature of this game makes it very difficult for anyone to reach or sustain a 700 Elo. Anyone who approaches or reaches 700 likely got there by having a
lot of things go right. But that 700 Elo is temporary and not sustainable. I view that as an inherent characteristic of this specific game, not a flaw in the rating system. Overall though, the player ratings for this game give a pretty good idea of who the good players are, and give a very good idea of which players are the most accomplished.
Generally speaking, if we look at 3 statistics for a player:
1) Current Elo
2) Historically highest Elo achieved by that player
3) Rolling average Elo over the last x games [We could discuss what value x should be] ,
we can get a very good idea of a player's skill level and accomplishments. We could even use those stats for matchmaking, singly or in combination. A separate leaderboard could exist for each of those 3 categories.
Then we would have a leaderboard to reflect:
Current and recent accomplishment,
Peak historical accomplishment,
Consistency over a period of time.
Seems to me those solutions satisfy most or all of what the people in this thread are asking for.