ELO on luck based games

Discussions about BGA (all languages)
Forum rules
Warning: challenging a moderation in Forum = 10 days ban
More info & details about how to challenge a moderation: viewtopic.php?p=119756
User avatar
BarnardsStar
Posts: 544
Joined: 02 January 2021, 02:41

Re: ELO on luck based games

Post by BarnardsStar »

Fuchur wrote: 22 April 2026, 19:10 If ELO (in the sense of a linear measure of skill for a specific game, giving predictions for the average outcome) exists,
Nitpick: it’s logarithmic, not linear.

If two players with an Elo differential of 120 play, the odds of the stronger one winning are 2:1.
At a differential of 240 it’s 4:1.
At 360 it’s 8:1.
At 480? 16:1.

That’s logarithmic, not linear.
User avatar
tbhp
Posts: 716
Joined: 16 April 2025, 00:35

Re: ELO on luck based games

Post by tbhp »

BarnardsStar wrote: 23 April 2026, 01:59
Fuchur wrote: 22 April 2026, 19:10 If ELO (in the sense of a linear measure of skill for a specific game, giving predictions for the average outcome) exists,
Nitpick: it’s logarithmic, not linear.

If two players with an Elo differential of 120 play, the odds of the stronger one winning are 2:1.
At a differential of 240 it’s 4:1.
At 360 it’s 8:1.
At 480? 16:1.

That’s logarithmic, not linear.
The way Elo works is fascinating. I mean, if you look at a game like Chess, a player who is at 2000 Elo is already an excellent player who puts a lot of thought into their games, right? Yet according to this, that player would still only win around 1 in 64 times against a player like Magnus Carlsen.

How is it possible? Does Magnus Carlsen never have bad days? Does he never fumble his plans? It's crazy to think about.
Fuchur
Posts: 126
Joined: 20 May 2016, 22:45

Re: ELO on luck based games

Post by Fuchur »

Yeah, right, "linear" isn't a good choice -- i wanted to stress the point that a one-dimensional measure is enough to capture all of the skill.

If you think of a sport like biathlon you might think that -- in that case -- a (at least) two-dimensional skill-measure would be needed, one to measure the skill of skiing, the other of shooting and, without strong conditions on all competitions of biathlon, it shouldn't be possible to reduce the skill-measurement of that sport to one dimension.
BarnardsStar wrote: 23 April 2026, 01:59
Fuchur wrote: 22 April 2026, 19:10 If ELO (in the sense of a linear measure of skill for a specific game, giving predictions for the average outcome) exists,
Nitpick: it’s logarithmic, not linear.

If two players with an Elo differential of 120 play, the odds of the stronger one winning are 2:1.
At a differential of 240 it’s 4:1.
At 360 it’s 8:1.
At 480? 16:1.

That’s logarithmic, not linear.
FrankJones
Posts: 2444
Joined: 30 June 2024, 00:24

Re: ELO on luck based games

Post by FrankJones »

Fuchur wrote: 23 April 2026, 15:53

If you think of a sport like biathlon you might think that -- in that case -- a (at least) two-dimensional skill-measure would be needed, one to measure the skill of skiing, the other of shooting and, without strong conditions on all competitions of biathlon, it shouldn't be possible to reduce the skill-measurement of that sport to one dimension.

I do not understand this analogy.

Here at BGA, each game has an Elo system that, although using the same formula, works differently for each game, separately measuring the results of each individual game.
User avatar
BarnardsStar
Posts: 544
Joined: 02 January 2021, 02:41

Re: ELO on luck based games

Post by BarnardsStar »

Fuchur wrote: 23 April 2026, 15:53 If you think of a sport like biathlon you might think that -- in that case -- a (at least) two-dimensional skill-measure would be needed, one to measure the skill of skiing, the other of shooting and, without strong conditions on all competitions of biathlon, it shouldn't be possible to reduce the skill-measurement of that sport to one dimension.
By “strong conditions” you mean like always running it on the same course, always running on sunny days, that sort of thing?

And yet… in a lot of team sports it has become fashionable to rate players by WAR—wins over replacement. That is, if that player left the team and you had to replace them with a base-skilled player, a rookie or some such, how many fewer games would you expect to win in a season because you didn’t have that player?

This is, of course, a somewhat ridiculous exercise. You can’t possibly know that. And yet people do it as an exercise in reducing all of a player’s skill to one number. As I say, ridiculous, and yet, it has value. You can disgree over how much of a factor, say, batting average should play in WARs, and several definitions exist, but it does have value in being able to compare a top-notch slugger with a knuckleball pitcher. (Insert apt football [soccer] and cricket analogies here.)

Likewise, it may be slightly ridiculous to expect all players to have an Elo and all possible combinations of matchups to obey those Elos. (Classic example, three rock-paper-scissors players, each of whom favor one of the three outcomes.) This doesn’t mean it doesn’t have value, or that Elo doesn’t exist. It means it’s an approximation of something that may not be completely linear. But approximations still have value.
User avatar
Jellby
Posts: 3546
Joined: 31 December 2013, 12:22

Re: ELO on luck based games

Post by Jellby »

tbhp wrote: 23 April 2026, 02:30 Yet according to this, that player would still only win around 1 in 64 times against a player like Magnus Carlsen.

How is it possible? Does Magnus Carlsen never have bad days?
Yes, once every 64 days, more or less :D

Unrelated to the quote above:
The elo model assumes that if A beats B 64% of the time and B beats C 64% of the time (difference 100), then A must beat C 76% of the time.
Similarly, if A beats B, and B beats C 95% of the time (difference 500), then A must beat C 99.7% of the time.
For a game where this is not true (for instance because there's always a 1 in 20 chance that you'll draw the "you win" card), elo cannot be an accurate model. We can discuss whether it's a good enough approximation, or whether the assumption is actually true for a game like chess...
FrankJones
Posts: 2444
Joined: 30 June 2024, 00:24

Re: ELO on luck based games

Post by FrankJones »

Jellby wrote: 23 April 2026, 17:59

Unrelated to the quote above:
The elo model assumes that if A beats B 64% of the time and B beats C 64% of the time (difference 100), then A must beat C 76% of the time.
Similarly, if A beats B, and B beats C 95% of the time (difference 500), then A must beat C 99.7% of the time.
For a game where this is not true (for instance because there's always a 1 in 20 chance that you'll draw the "you win" card), elo cannot be an accurate model. We can discuss whether it's a good enough approximation, or whether the assumption is actually true for a game like chess...
Perhaps in that scenario, the nature of the game itself (too much luck) makes it impossible for anyone to win at a high enough win percentage to maintain and specific high Elo.

Perhaps, in such a high-luck game, the player who is "expected" to win x% of the time only achieved a high-enough Elo to create that situation by going on a bit of a lucky streak that is not sustainable.

I feel as though we are back to the same circular argument:

1) A player in a high-luck game goes on a hot streak, reaches 500 Elo.
2) That player says, "I'm a 500-Elo player because I have reached 500."
3) Because that Elo is not actually sustainable, but rather is the peak achievable Elo, that player's Elo then swings back downward.
4) That player says, "Elo for this game fails, because it does not allow me to stay at 500, therefore it does not give accurate Elos."

(Jellby, I'm not saying you are making this argument. You acknowledged there is a discussion to be had.)

The above argument has flaws. There are one or more fallacies in the argument. One or more premises are flawed or rely on assumptions. Specifically, if the system is being declared to be broken or inaccurate, then we don't know that a player who reaches 500 is actually truly a 500 Elo player.

So, if that player then cannot sustain that Elo, we don't know that Elo is failing. Maybe it is actually self-correcting what was an unsustainably high Elo resulting from a streak of good luck.

In other words, the proof that Elo is broken relies on an initial assumption that ... Elo is accurate. The conclusion of the proof ends up contradicting a premise. That is a fallacious argument.

There is a high-skill game here at BGA that also has a lot of luck. Despite having a lot of skill, the game can swing on one lucky outcome of an RNG card reveal. The nature of this game makes it very difficult for anyone to reach or sustain a 700 Elo. Anyone who approaches or reaches 700 likely got there by having a lot of things go right. But that 700 Elo is temporary and not sustainable. I view that as an inherent characteristic of this specific game, not a flaw in the rating system. Overall though, the player ratings for this game give a pretty good idea of who the good players are, and give a very good idea of which players are the most accomplished.

Generally speaking, if we look at 3 statistics for a player:
1) Current Elo
2) Historically highest Elo achieved by that player
3) Rolling average Elo over the last x games [We could discuss what value x should be] ,

we can get a very good idea of a player's skill level and accomplishments. We could even use those stats for matchmaking, singly or in combination. A separate leaderboard could exist for each of those 3 categories.

Then we would have a leaderboard to reflect:
Current and recent accomplishment,
Peak historical accomplishment,
Consistency over a period of time.

Seems to me those solutions satisfy most or all of what the people in this thread are asking for.
Ceaseless
Posts: 1300
Joined: 12 November 2022, 17:06

Re: ELO on luck based games

Post by Ceaseless »

Jellby wrote: 23 April 2026, 17:59 The elo model assumes that if A beats B 64% of the time and B beats C 64% of the time (difference 100), then A must beat C 76% of the time.
Similarly, if A beats B, and B beats C 95% of the time (difference 500), then A must beat C 99.7% of the time.
For a game where this is not true (for instance because there's always a 1 in 20 chance that you'll draw the "you win" card), elo cannot be an accurate model. We can discuss whether it's a good enough approximation, or whether the assumption is actually true for a game like chess...
Or in more subtle cases, maybe it's only off by a percent or two, and slowly gets less accurate the further you stretch the gaps.
Ceaseless
Posts: 1300
Joined: 12 November 2022, 17:06

Re: ELO on luck based games

Post by Ceaseless »

FrankJones wrote: 23 April 2026, 18:14 In other words, the proof that Elo is broken relies on an initial assumption that ... Elo is accurate. The conclusion of the proof ends up contradicting a premise. That is a fallacious argument.
No, that's called proof by contradiction. You assume a premise is true, find that assumption leads to a contradiction, which has to be false, and as a result the premise must be false. Whether it's successfully applied involves analysis of the specific argument, but it's a known structure.
FrankJones
Posts: 2444
Joined: 30 June 2024, 00:24

Re: ELO on luck based games

Post by FrankJones »

Ceaseless wrote: 23 April 2026, 18:37
FrankJones wrote: 23 April 2026, 18:14 In other words, the proof that Elo is broken relies on an initial assumption that ... Elo is accurate. The conclusion of the proof ends up contradicting a premise. That is a fallacious argument.
No, that's called proof by contradiction. You assume a premise is true, find that assumption leads to a contradiction, which has to be false, and as a result the premise must be false. Whether it's successfully applied involves analysis of the specific argument, but it's a known structure.
I'm familiar with proof by contradiction. One of the most famous such proofs is the proof showing that SQRT[2] is irrational.

Proof by contradiction is not what is happening here.

A proof by contradiction would start with an assumption, then state any additional premises (all of which must be true, and which cannot be self-contradictory), and then show that a logical conclusion follows that contradicts the original assumption, therefore showing that assumption to be false.

The argument by contradiction relies on all premises (other than the initial assumption) being true. When the contradiction arises, we reject the original assumption because we know with 100% certainty the other premises are true.

On the other hand, if an argument relies on one or more unproven or assumed premises, or had too many false premises, or has premises (other than the initial assumed premises) that contradict one another, that’s just a bad argument. It is either invalid or unsound.
If someone wants to make an attempt to argue by contradiction, that’s fine. It should look like this:

1) State an initial assumption
2) List any necessary premises, which must be true.
3) Show through valid argumentation techniques and structures that a conclusion can be reached that contradicts the initial assumption.

I definitely do not agree anyone has successfully done this, and I’m not sure it can be done. [With regard to the debate in this thread.]

The proof of the irrationality of SQRT[2] follows that structure. We assume SQRT[2] to be rational. We then list one or more premises that have been proven 100% true. We then perform valid mathematical operations, the result of which yield a conclusion that contradicts our initial assumption. Only because we know all other premises are true, and because we know the argument structure was valid and absent of any fallacies or further assumptions, only then can we reject the initial assumption. Which in the case of “SQRT[2] is rational”, we can do. Once we reach the contradiction, there is one and only one false premise: The initial assumption, so we reject it.

In this discussion / debate, I am asserting that the argument (not that anyone has actually written a formal logical argument in proper form) has more than one assumption, and therefore we cannot just pick and choose which one to reject. Perhaps “Elo is flawed” is not the problematic assumption, but rather, one of the other assumptions. If we arrive at a contradiction, we only know that at least one premise is false. That’s why it is crucial to have only one assumed premise and all other premises definitively true.
Post Reply

Return to “Discussions”