BGA has a serious bully problem

Discussions about BGA (all languages)
Forum rules
Warning: challenging a moderation in Forum = 10 days ban
More info & details about how to challenge a moderation: viewtopic.php?p=119756
User avatar
MarkSteere
Posts: 265
Joined: 09 March 2021, 19:21

Re: BGA has a serious bully problem

Post by MarkSteere »

I think "luck based" is just a shorthand way of saying "includes elements of luck," like rolled dice.

Personally I don't think the inclusion of dice makes a huge difference. It just broadens the variance, so to speak. Instead of just not knowing what your opponent will do, you additionally don't know what the dice will do.

Without luck elements, there's still kind of a pseudo luck. Unless you're playing perfectly, there's guesswork. And where there's guesswork - even if it's educated, informed, intuitive guesswork - there's luck. Instead of tossing a fair coin, you may be tossing a loaded coin. But you still need some degree of luck to win. You might guess wrong. Repeatedly. I don't expect anyone to agree with this. Just my point of view.

I'm not too knowledgeable about rating systems. I uncharacteristically focused in on that one point about Elo, but I'll leave it to others to sort out what's best for BGA.
FrankJones
Posts: 2449
Joined: 30 June 2024, 00:24

Re: BGA has a serious bully problem

Post by FrankJones »

MarkSteere wrote: 04 February 2026, 16:32 It's not a question of ChatGPT's track record or its veracity. It's a question of "Does its argument make sense?" Is there really any doubt at this point that Elo was designed for pure strategy games like Chess? Or that it's ill suited for luck based games because it overestimates predictability in them?

It makes sense to me. A lot more sense than some of the wild theories posited in this discussion.

I've seen ChatGPT go wrong many times. But sometimes its analysis is spot on, like in this case.

If you want to dismiss ChatGPT out of hand because it "doesn't know anything," or "has been simply wrong," that's your prerogative. In my case, I'm willing to give it a fair shake.
Yes, there is doubt. I disagree that it is "ill-suited." I believe it works well.

No one is dismissing CHatGPT "out-of-hand". People are questioning one specific query by giving actual arguments and explanations and examples rather than just accepting the chatGPT result at face-value.
FrankJones
Posts: 2449
Joined: 30 June 2024, 00:24

Re: BGA has a serious bully problem

Post by FrankJones »

Andrei3009 wrote: 04 February 2026, 17:24

I will try to answer you questions
2) work better means less deviation from average (true elo) and more probability that given person plays on his/her current level. I think more formalistic definition could be given, but i don't think it's necessary right now.
1) Elo formula was created for games without randomness at all. All other games could be called luck based.

The more luck in game the more players deviate from their average rating. This could be naturally observed by manually checking to players elo(if they are playing frequently enough). In kingdomino for example top player elo fluctuate from 740 and 800, in azul top players fluctuate a lot less.

For me not luck based games are games with at least some 700+ players. In those games i feel in control and never lose to low ranked players.

3) Purpose of elo on most platforms is to give players chance to play with people of the similar level
Okay, now we're getting somewhere.

1) Although Elo may have been crated for games with zero luck, that doesn't mean we cannot use it for games with some luck. Is there a better alternative?

2) Games with 700-rated or even 800-rated players still have luck, and can still see wild fluctuations. Terraforming Mars, for example. Every player who has reached 800 Elo in that game has fallen back below 700 at some point. The more luck there is, the more variance there will be. That will be true regardless of what rating system is used though, wouldn't it?

3) And on BGA, it works very well. Of the 15-ish games I play, I get similarly skilled opponents when I play people with similar Elo ratings. It's not perfect, but that's true in chess also. If a player goes on a winning streak and goes from 2000 to 2130, while another player goes on a losing streak goes from 2280 to 2130, it's possible the players are not equally skilled. But based on their Elos at that moment, they might think they are equally skilled.
User avatar
tbhp
Posts: 717
Joined: 16 April 2025, 00:35

Re: BGA has a serious bully problem

Post by tbhp »

Andrei3009 wrote: 04 February 2026, 14:42
Jellby wrote: 04 February 2026, 14:09 If I say that player A has a 90% chance of winning against player B, they play, and player A loses... Was I wrong? Did my prediction fail?

The only way to know is if players A and B play many times with the exact same circumstances, which is of course impossible.
If we aren't counting really life parts that usually don't matter that much (how much person slept and similar things) than every game will have exact circumstances because you playing same game.
Exact circumstances in this context would mean that everything is exactly the same.

Think of the butterfly effect. If anything is different, that could lead to a different result of the game.

Predictions are never 100% certain, because we don't have some perfect map of the universe.
User avatar
BarnardsStar
Posts: 544
Joined: 02 January 2021, 02:41

Re: BGA has a serious bully problem

Post by BarnardsStar »

MarkSteere wrote: 04 February 2026, 06:16
BarnardsStar wrote: 04 February 2026, 05:52 Here’s what our new overlord told me. It actually gave about a page of text but here are choices excerpts including its bottom-line summary:
The system uses probability to predict match outcomes, which inherently accommodates luck … In luck-influenced games, while individual matches may be swayed by chance, the ELO system relies on cumulative results across many games. Over time, player skill becomes more pronounced, and the effects of luck dilute.

Conclusion: The ELO rating system remains effective in games where luck is a factor due to its adaptive mechanics in evaluating skill levels over various matches, its probabilistic approach to outcomes, and its ability to accumulate results over time. This allows it to account for the inherent unpredictability of luck, making it a robust tool for ranking players fairly.
Tom Lehrer once said, “Life is like a sewer: what you get out of it depends on what you put into it.” I’d say the same thing about LLMs.
"Effective" isn't the first word that comes to mind to describe Elo for luck based games. I was playing a lot of Backgammon a while back. After a long winning streak my Elo would go way up hundreds of points, and similarly for the other direction. I'd say it "kinda works ok much of the time" as a very crude measure for luck based games. Like your chance of ever beating anyone with a way higher rating is slim. It would require a long series of very opportune rolls.
You’re missing my point here, Mark. I’m not saying anything about Elo; I’m providing an argument from ChatGPT that is exactly opposite your ChatGPT-provided arguments. My actual point is that LLMs are not to be trusted.

My motto is, “never ask an LLM a question you cannot easily verify.” Your ChatGPT-provided arguments appear to make sense on the surface, because this is what LLMs excel at. There is no guarantee that the argument is actually true. Now, it may be, but you cannot trust it just because ChatGPT said so or because it seems to make sense. You have to actually verify it, and I’ve not seen you do that.
User avatar
tbhp
Posts: 717
Joined: 16 April 2025, 00:35

Re: BGA has a serious bully problem

Post by tbhp »

ChiefPointThief wrote: 04 February 2026, 09:24 As for the whole 90% probability discussion there is a specific arena player that I am calculated to win against at that rate in a game w/ a luck factor of 3. But these win probabilities are off thus resulting in me losing 100pts throughout the arena season to this player even though I beat them 75%.
Elo can't know about all the fishy factors that go on in real life that may mean a probability shouldn't be taken at face value.

If someone who is elo 100 is actually a smurf and I lose a ton of points when losing against them, it's not elo's fault for not being able to guess that the player was actually stronger than their elo suggested. The ratings are still mathematically correct considering the results of the games that both accounts have played thus far.

The problem is when people try to infer information that the ratings can't possibly tell you with certainty. It's not their job to do that.

Regarding this player that you played against, it seems like they may have had a lot of bad luck in their games, so when you play against them, you are expected to perform with the same success others did, otherwise you are going to lose a lot of points. It may seem unfair from a certain perspective, but it's actually totally in line with the spirit of games and sportsmanship. The results are what counts.

If you lose many football games during a season because of bad luck, you aren't gonna complain that your losses are not representative of your skill (I mean, you may, but it doesn't change the reality of the rankings). Lots of things in sports happen because of factors outside of our control, because the point of games is to be won. Skill is just a tool towards achieving victory.
Last edited by tbhp on 04 February 2026, 18:41, edited 4 times in total.
User avatar
Gooorn
Posts: 86
Joined: 27 January 2017, 20:10

Re: BGA has a serious bully problem

Post by Gooorn »

FrankJones wrote: 04 February 2026, 17:47
3) And on BGA, it works very well. Of the 15-ish games I play, I get similarly skilled opponents when I play people with similar Elo ratings.
Far not well in co-op and high luck factor games. Too easy to level up as beginner and too easy to fall down from expert level.

And other point that ELO system can't handle the different versions of the same game. A player can be good in one version but beginner in the other version of the same game.
Last edited by Gooorn on 04 February 2026, 18:35, edited 1 time in total.
stillframe
Posts: 42
Joined: 23 December 2025, 21:11

Re: BGA has a serious bully problem

Post by stillframe »

BarnardsStar wrote: 04 February 2026, 18:16You’re missing my point here, Mark. I’m not saying anything about Elo; I’m providing an argument from ChatGPT that is exactly opposite your ChatGPT-provided arguments. My actual point is that LLMs are not to be trusted.

My motto is, “never ask an LLM a question you cannot easily verify.” Your ChatGPT-provided arguments appear to make sense on the surface, because this is what LLMs excel at. There is no guarantee that the argument is actually true. Now, it may be, but you cannot trust it just because ChatGPT said so or because it seems to make sense. You have to actually verify it, and I’ve not seen you do that.
Another good point. It is quite easy to get an LLM to give opposing "opinions" or "arguments". They will say almost anything with the right prompting. It is a text generator, not a thought generator.

The Elo system is descriptive of results. If players A and B play a game 20 times, and A wins 18 of them, then A will have a rating 400 points higher (or however a particular implementation is scaled). The type of game is not an input. The fact that it was designed for chess has no bearing on its suitability for other games. If a game has a high luck factor, then this win rate will not occur, and nor will the rating difference.
User avatar
ChiefPointThief
Posts: 745
Joined: 14 August 2020, 22:27

Re: BGA has a serious bully problem

Post by ChiefPointThief »

FrankJones wrote: 04 February 2026, 17:16
ChiefPointThief wrote: 04 February 2026, 09:24
Meeplelowda wrote: 04 February 2026, 04:58
I'm going to go out on a limb and say zero. People conflate "high luck" with being non-deterministic, i.e., having random elements as part of the game mechanics. Games that are truly high luck, in the sense that they are really just an elaborate way of flipping a coin, don't have players in the 500s.
Captain flip for one.
I have a 68% win rate. The player above me has a 50% win rate which is a HUGE difference in any game let alone a game with a 4 luck rating. I’ve also maintained this over thousands of games and only play with players good and above (if you see someone lower than it was a game w/ a friend).

As for the whole 90% probability discussion there is a specific arena player that I am calculated to win against at that rate in a game w/ a luck factor of 3. But these win probabilities are off thus resulting in me losing 100pts throughout the arena season to this player even though I beat them 75%. People complain about top players being blocked but for me it is way more beneficial to block a player like this than a top player. After doing numbers I realized how flawed the system is. The system doesn’t gauge who has the best overall season only who finished w/ the best streak. Therefore it is flawed.

The win probabilities are laughable for team games.
Are you sure you're understanding the Elo system? Win% is not relevant. The calculation is based on the Elo difference. If you have a 68% win rate against generally low rated players, your Elo won't be as high as someone who has a 50% win rate against strong players.

If you think there is supposed to be a direct correlation between win% and Elo, then you are not understanding how Elo works.

The "luck factor" listed on this website is meaningless. If you are using that factor in any way, you are misunderstanding something fundamentally important.

If you cannot beat this player more than 75% of the time, then you are not good enough to beat that player 90% of the time, so I'm not sure why you're saying you "should" be winning 90% of the time. When you fail to beat that player 90% of the time, your Elo decreases to reflect that. It means you're not actually a true 400 Elo better than that opponent.

You're also referring to Arena points when you talk about Elo. Arena points and Elo are two different things.

You have a lot of misunderstandings and inaccuracies in this one post.
Yes I am sure I know how elo works & it is you who has misunderstandings and inaccuracies in your post 🤣. The purpose of me mentioning that I set my games to only good players and above was so that the same counter argument of “it must be strength of opponents” wasn’t spewed. You are arguing with me that someone who has a 50% win rate in captain flip is stronger than me who has a 68% win rate vs opponents good and above.

I know what the luck factor is not because of what bga list but because Ive actually played the game thousands of times. & even if I didn’t it’s not hard to look at a game that you randomly pick a tile out of a bag and keep it or flip it to a random character to know that it’s a game with a high luck factor 😜.

I’m not saying I should beat this player 90%. I’m actually saying the opposite. I MUST beat them 90% to break even. Which isnt an accurate estimation of the true win probabilities. If Vegas used this system they would be broke overnight.

The “competitive” mode arena is intentionally set at double the k factor to be more swingy. It is far from a fair measure of who actually performs the best in a given season for luck ganes. Which is why every arena season some player who plays a game with high luck factor heads to the forums to complain about this very topic.
FrankJones
Posts: 2449
Joined: 30 June 2024, 00:24

Re: BGA has a serious bully problem

Post by FrankJones »

Gooorn wrote: 04 February 2026, 18:29
FrankJones wrote: 04 February 2026, 17:47
3) And on BGA, it works very well. Of the 15-ish games I play, I get similarly skilled opponents when I play people with similar Elo ratings.
Far not well in co-op and high luck factor games. Too easy to level up as beginner and too easy to fall down from expert level.

And other point that ELO system can't handle the different versions of the same game. A player can be good in one version but beginner in the other version of the same game.
OH, I definitely agree there - Elo for coop games makes little sense. Some such games here on BGA have Elo disabled, I think? I could be mistaken.

I also agree about different versions of the game creating stations in which Elo doesn't help us know an ability or the results of a player in specific settings within a specific game. Chess is chess, for the most part, so we know what a 2400 chess Elo represents. Sure, there are chess variants, but when people refer to a chess Elo we know they are referring to standard chess.

As for your statement about high-luck games - that's a matter of opinion, not fact. I don't share that opinion. High luck games are inherently unlikely to have high-Elo players, and are inherently prone to lucky streaks. That doesn't mean the Elo system is flawed. The Elo system typically reflects these facts in the Elo results. Anyone expecting Elo to be a predictive factor in an upcoming high-luck game should understand that one or both players may not be rated to their exact skill level. Whether that's a "slaw" depends on the vantage point. I don't consider it a flaw. I accept that Elo for high-luck games is an indicator of what has occurred, not what is to be expected. But in the long run, we can see Elo trends of the best players.

For example, in open face Chinese poker, I lose a lot of games, and my Elo fluctuates, but my Elo still places me at or near the top, as it should, since I'm an expert at the game. Other people in my situation might be unhappy that their Elo in Open face Chinese poker cannot reach 500 or 700. But that's partly because people refuse to play larger sample games. If a game were 30 hands instead of 6 hands, The Elos would fluctuate less.
Post Reply

Return to “Discussions”