Page 9 of 22
Re: BGA has a serious bully problem
Posted: 04 February 2026, 07:32
by MarkSteere
tbhp wrote: ↑04 February 2026, 06:55
You seem to be assigning the blame...
Luck based games defy rating systems to a degree because of luck streaks. It's harder to precisely pinpoint a player's skill level.
The only thing I know of that's wrong with Elo in particular, that I learned today from ChatGPT, is that Elo assumes a much tighter distribution of outcomes than actually exists. The real distribution is flatter and effectively capped, so Elo overstates how much rating differences should matter. This introduces systematic inaccuracy.
I don't know why people are struggling with this. It makes perfect sense. The argument doesn't depend on ChatGPT 's credibility, although this is exactly the sort of topic ChatGPT excels at explaining.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 07:49
by tbhp
MarkSteere wrote: ↑04 February 2026, 07:32
tbhp wrote: ↑04 February 2026, 06:55
You seem to be assigning the blame...
Luck based games defy rating systems to a degree because of luck streaks. It's harder to precisely pinpoint a player's skill level.
The only thing I know of that's wrong with Elo in particular, that I learned today from ChatGPT, is that Elo assumes a much tighter distribution of outcomes than actually exists. The real distribution is flatter and effectively capped, so Elo overstates how much rating differences should matter. This introduces systematic inaccuracy.
I don't know why people are struggling with this. It makes perfect sense. The argument doesn't depend on ChatGPT 's credibility, although this is exactly the sort of topic ChatGPT excels at explaining.
Elo only assumes things that have factually happened. If a player wins 90% of their first 10 games, played against base level players, then elo will assume that this player is expected to win 90% of their games against said players. So, if such a player were to win a game against them, they would gain a lot of elo points.
One may say: but what if this player actually got this winrate thanks to luck? Then the points gained wouldn't be deserved, because they just beat a lucky opponent, not an actually good opponent.
Well, none of this matters, because this is based on the idea that "skill" should be the metric judged in a game or a sport, which is false. What we judge, all the time, is actually the results. It doesn't matter if a result was obtained through skill or luck. It's always the results themselves that we look at when we rank players, because that's what determines a winner at the end of the day.
In order to understand elo, one must do away with the idea that it must be representative of a certain idea of "skill". Elo doesn't know about that, it just sees a list of games with winners and losers. You can apply elo to a game that is 100% luck based, and its rankings would still be justified because what it relates is the reality of how the games went down. Not a measure of how much "skill" was involved in reaching those results.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 08:14
by MarkSteere
Well, ok. I have a lot to learn about ratings. The topic comes up sometimes. I've never paid close attention because I'm not concerned about my own ratings. I think it's a more of a concern for highly rated players, and to a lesser extent, low rated players. I'm somewhere around average, on average. My high rating in Oust is misleading. It comes from frequently beating novices. There are definitely more skilled Oust players out there. I just haven't been playing them.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 09:24
by ChiefPointThief
Meeplelowda wrote: ↑04 February 2026, 04:58
FrankJones wrote: ↑04 February 2026, 04:43
How many "high luck" games are there that even have players with a 500 Elo?
I'm going to go out on a limb and say zero. People conflate "high luck" with being non-deterministic, i.e., having random elements as part of the game mechanics. Games that are truly high luck, in the sense that they are really just an elaborate way of flipping a coin, don't have players in the 500s.
Captain flip for one.
I have a 68% win rate. The player above me has a 50% win rate which is a HUGE difference in any game let alone a game with a 4 luck rating. I’ve also maintained this over thousands of games and only play with players good and above (if you see someone lower than it was a game w/ a friend).
As for the whole 90% probability discussion there is a specific arena player that I am calculated to win against at that rate in a game w/ a luck factor of 3. But these win probabilities are off thus resulting in me losing 100pts throughout the arena season to this player even though I beat them 75%. People complain about top players being blocked but for me it is way more beneficial to block a player like this than a top player. After doing numbers I realized how flawed the system is. The system doesn’t gauge who has the best overall season only who finished w/ the best streak. Therefore it is flawed.
The win probabilities are laughable for team games.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 13:01
by Andrei3009
I think Frank and Mark are both right.
Elo was really designed for non-random games and works best there. By best, I mean the stability of the rating and the influence of winstreaks. In games with a high degree of randomness and also with a large k factor (arena), people often end up in places significantly above/below their average elo, which may confuse other people a little. When playing with a 700 elo checkers player, you know that his real level (true elo) is close to this value, however, even in such a serious game as conquering Mars, your 700 elo opponent can play much worse than those who are stable on this rating
FrankJones wrote: ↑04 February 2026, 06:35
Can anyone show me evidence of a game in which a player consistently maintains a 550 Elo but cannot consistently beat 150-Elo opponents?
If the true player ratings are 150 and 550, then by definition the stronger player should win ~90% of the time. However, as noted by Mark, in reality, the formula does not work so accurately for large differences in elo, and the weaker player wins much less, up to <1%. So for me, consistently winning with a difference of 400 elo is never or almost never losing. A game where this is not working is 7 wonders duel (base game), where I once lost to a person 355 (140 vs 495) ratings lower and knew that could happen again. This happends in this particular game because there is a strong randomness factor, and sometimes the strongest moves turn out to be the most intuitive, and even a weak player makes them possibly without even understanding why they are strong.
I haven't delved into this issue, but I know that Go uses a Glicko rating that works almost like elo, only without the strong influence of recent games. As far as I understand, it works like this: you take the last N (N is big number possibly containing all persons games) games of a person and assuming that his opponents played for their real rating, you calculate what the most probable ELo rating you should have in order to achieve those results for these N games
Re: BGA has a serious bully problem
Posted: 04 February 2026, 14:09
by Jellby
If I say that player A has a 90% chance of winning against player B, they play, and player A loses... Was I wrong? Did my prediction fail?
The only way to know is if players A and B play many times with the exact same circumstances, which is of course impossible.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 14:42
by Andrei3009
Jellby wrote: ↑04 February 2026, 14:09
If I say that player A has a 90% chance of winning against player B, they play, and player A loses... Was I wrong? Did my prediction fail?
The only way to know is if players A and B play many times
with the exact same circumstances, which is of course impossible.
If we aren't counting really life parts that usually don't matter that much (how much person slept and similar things) than every game will have exact circumstances because you playing same game.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 15:23
by Gooorn
If a player has 90%+ chance to win or lose against someone, they shoudnt play against each other. At least if they want to play for fun not for ELO or rank or thropies.
I prefer play for fun when i have 50% chance to win.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 15:25
by Gooorn
Gooorn wrote: ↑04 February 2026, 15:23
If a player has 90%+ chance to win or lose against someone, they shoudnt play against each other. At least if they want to play for fun not for ELO or rank or thropies.
I prefer play for fun when i have 50% chance to win.
Delete please. i wanted to edit not quote myself.
Re: BGA has a serious bully problem
Posted: 04 February 2026, 15:51
by stillframe
MarkSteere wrote: ↑04 February 2026, 00:30
FrankJones wrote: ↑03 February 2026, 20:22
MarkSteere wrote: ↑03 February 2026, 18:09
ChatGPT...
Elo does not perform as well for luck-based games because it was designed around a skill-dominant outcome model. When variance dominates, the Elo probability distribution systematically overstates how predictive rating differences really are.
The issue isn’t that Elo “can’t handle luck,” but that it assumes luck is small relative to skill.
Well, with all due respect to chatGPT, its answers are not always correct.
Oh definitely. I've seen ChatGPT be totally wrong. But on a pedestrian topic like Elo, on which a lot has been written, this seems fairly plausible, at least to me.
LLMs do indeed function by producing plausible sounding text. That is not the same as correct text. Sometimes it might be both, other times not, and the LLM doesn't know the difference,
because it doesn't know anything at all.
MarkSteere wrote: ↑04 February 2026, 01:03
If you're asking ChatGPT to be creative on a topic on which little or nothing has been written, that's where it falls flat. But if you're just asking it to summarize or synthesize tons of existing material, it's remarkably adept.
Disagree. Many times I have entered a straightforward factual query into search, and the 'AI response' at the top has been simply wrong, even while all the real search results immediately below give the correct answer.