Board Game Arena, Catan, virtual dice and randomness - a short trial

Forum rules
Please DO NOT POST BUGS on this forum. Please report (and vote) bugs on : https://boardgamearena.com/bugs
muntzer
Posts: 23
Joined: 13 September 2021, 18:12

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by muntzer »

I feel pretty happy to respond to euklid314 - his handle has both my favorite Greek mathematician and the first three digits of pi :)
I have Euclid's Elements on my bookshelf.
My conclusion: muntzer found an unlikely event in his 20 games which has a probability of only 1.2%. Quite an unlikely event, but not shocking at all.
Sure - the number of times outcomes fell outside confidence intervals will indeed be subject to variation. Good point.

But what your analysis ignores is how far those outcomes fell outside the confidence intervals. Isn't that also relevant?

Let's say, in my trials, I had only one anomalous result - zero 7s rolled out of 100. That's only one result, so let's pack up and go home, right? Hopefully, you'll agree that this specific outcome is so far outside expectations that it stands in its own right as evidence for a broken dice implementation?

So, let me repeat - many of the outliers were in the "once per 300 or 400 or 500 games" territory. Why did I see so many of these outliers in only 20 games? And the Grand Doozy of them all - 28 nines rolled out of 102 rolls. We should expect to see that only ONCE in 230,177 games of Catan.

Is no one else even a little bit shocked by how far the outliers missed their confidence intervals?

As also noted in my previous reply (edit: actually, it's below) - in several cases, more than one number within the same game missed on the same side of the confidence interval - both outcomes higher or both outcomes lower than expected - with no other significantly skewed outcomes in the other direction.
Last edited by muntzer on 01 November 2022, 22:26, edited 1 time in total.
muntzer
Posts: 23
Joined: 13 September 2021, 18:12

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by muntzer »

Just in case some people didn't read all my posts:

Rolling 2s: Count of 9 from 102 rolls
Odds of event: 1 in 452 games

Rolling 3s: Count of 0 from 102 rolls
Odds of event: 1 in 340 games

Rolling 9s: Counts of 21 from 99 rolls, 28 from 102 rolls
Odds of events: 1 in 380 games, 1 in 230,177 games

Rolling 10s: Count of 1 from 97 rolls
Odds of event: 1 in 471 games

Rolling 11s: Count of 0 from 98 rolls
Odds of event: 1 in 271 games

230,177 games of Catan would take you almost 20 years to play - without stopping for toilet breaks. However, I saw that event within my third hour of playing BGA Catan. Again, with my objective already set in advance. Not that it should matter with an outcome so far outside expectations.
muntzer
Posts: 23
Joined: 13 September 2021, 18:12

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by muntzer »

siverure wrote: 01 November 2022, 16:37 I'm not sure about everything here, but my understanding is that you examined individual numbers over 20 games to see if anything was exceptionally above or below the expected rolls. It occurs to me that rolling one number an above average amount of times and another number a below average amount of times in the same game aren't independent events. I'm uncertain if you've accounted for this at all, especially given that the result is a little less than double what you expected. If you flip a coin 10 times and count 10 heads and 0 tails you only had one unusual set of coinflips, not two. I'm unsure how much this applies when you have eleven results but the same principle should have some effect.
Certainly a valid criticism. However, your example of coin tosses is one extreme: a sample space of two possible outcomes. So an excess of heads is necessarily reflected in a deficit of tails.

With a sample space of 11 outcomes, an unexpectedly high number of rolls for one number will make it more likely that other number counts will be lower than expected - of course. But we'd also expect the variation in the other 10 outcomes to follow the binomial distribution too. If we get ten 12s rolled instead of three 12s, the number of rolls that are not 12s is 90 instead of 97 - it's still a large number, and we are well within our rights to notice if those other 90 rolls are skewed as well.
The later post stating that similar expected numbers had massively disproportionate actual rolls in games seems to me like a flaw of a game involving randomness, not a flaw in the randomness. Again, for one number to be rolled an above average amount, another must have been rolled a below average amount.
Your point is well made, but my trial included cases where two or more different numbers in the same game were ALL BELOW or ALL ABOVE the confidence interval. That is, they both missed on the same side, which in combination is even less likely than each ocurring in isolation.

So, in Game #3, there were 6 twos rolled, 26 sevens rolled, and 16 tens rolled. All three missed the confidence interval on the SAME side of the interval, the high side. Only one number missed on the low side in thst game: there was 1 three rolled instead of the expected 5 or 6.

In Game #4, there were 9 twos rolled and 28 nines. Two outcomes way outside the confidence interval, both on the high side.

In Game #10, there was only one 10 rolled and only one 11 rolled. Two results that both missed on the low side.

Game #11, 3 fours rolled and 4 nines. Two results that both missed on the low side without any results missing on the high side.

Game #15, 3 fours rolled and 5 nines, both below the interval. (In this game there were also 7 twos rolled, two above the confidence interval, but again, this is only a few rolls out of 100.)

So, your observation about the events not being completely independent actually works against the BGA dice in these cases. We should expect fewer outcomes to fail on the same side of the confidence interval - in several cases, I saw the converse.
User avatar
SwHawk
Posts: 133
Joined: 23 August 2015, 16:45

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by SwHawk »

muntzer wrote: 01 November 2022, 21:59 So - why did you criticise me for using Excel?
It seems that I confused two very similar posts, and I apologize for it. Yet I notice you haven't answered to the questions I've asked.
muntzer wrote: 01 November 2022, 22:12 Let's say, in my trials, I had only one anomalous result - zero 7s rolled out of 100. That's only one result, so let's pack up and go home, right? Hopefully, you'll agree that this specific outcome is so far outside expectations that it stands in its own right as evidence for a broken dice implementation?

So, let me repeat - many of the outliers were in the "once per 300 or 400 or 500 games" territory. Why did I see so many of these outliers in only 20 games? And the Grand Doozy of them all - 28 nines rolled out of 102 rolls. We should expect to see that only ONCE in 230,177 games of Catan.

Is no one else even a little bit shocked by how far the outliers missed their confidence intervals?
Aren't we falling in two of the bias you've outlined in your first post, namely : Monkey brains and Confirmation bias ? Even though you claim that you started this trial wanting to see whether the implementation is broken or not, it seems you want to prove it IS broken...

Also to put some perspective regarding number, there are around 100 000 Catan games per month (give or take 10 000). So we WOULD expect that event to show up 1 game every 2 months, but real randomness is lumpy, so there could be clusters of that event occurring, and you might be unlucky enough to find yourself in that cluster.

My point being that your sample of 20 games, taken once, is, according to me, not enough to PROVE that the implementation is broken. Also, would BGA's implementation be broken, that would most certainly mean that PHP's implementation is broken, which would certainly already have been noticed given the wide range of software (including web applications) relying on that aspect of the language.
Last edited by SwHawk on 01 November 2022, 22:58, edited 1 time in total.
User avatar
euklid314
Posts: 680
Joined: 06 April 2020, 22:56

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by euklid314 »

The dependencies of number frequencies within a single game are much too intricate that I could tackle it mathematically. Let`s say that the 2s overshoot their interval of the expected [0,5] by 9 2s, this has almost no effect on the other numbers since there is low influence of say 9 2s versus 5 2s. Note that the 2s cannot even undershoot their interval.

If the 7s overshoot their interval however, they will likely overshoot it by a much larger number of rolls (since the 7-interval is much broader) and the effect on the other numbers will be much larger.

The longer I think about this problem, I find that the dependence of the 11 seperate 95%-intervals must be huge! Just consider the fact that you calculated the 11 95%-intervals separately, i.e. you always took the 100 rolls as a basis. You calculated how many 2s you expect from 100 rolls, how many 3s you expect from 100 rolls, etc. But almost all of these seperate cases that your mathematical model evaluates are invalid since the independent number will not add up to 100! If you choose some random number for the 2s within your 95%-interval, then some random number for the 3s within its 95%-interval, etc., then in the end it is extremely unlikely that all these numbers will add to 100. So almost all of the "likely" events within your model - that uses independent probabilies - are in reality impossible because they do not take care of the side condition of 100 rolls.

This might even influence the occurrence of those big outliers, but that I have not thought about thoroughly. More threads may follow. :-)
User avatar
euklid314
Posts: 680
Joined: 06 April 2020, 22:56

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by euklid314 »

muntzer wrote: 01 November 2022, 22:18 Just in case some people didn't read all my posts:

Rolling 2s: Count of 9 from 102 rolls
Odds of event: 1 in 452 games

Rolling 3s: Count of 0 from 102 rolls
Odds of event: 1 in 340 games

Rolling 9s: Counts of 21 from 99 rolls, 28 from 102 rolls
Odds of events: 1 in 380 games, 1 in 230,177 games

Rolling 10s: Count of 1 from 97 rolls
Odds of event: 1 in 471 games

Rolling 11s: Count of 0 from 98 rolls
Odds of event: 1 in 271 games

230,177 games of Catan would take you almost 20 years to play - without stopping for toilet breaks. However, I saw that event within my third hour of playing BGA Catan. Again, with my objective already set in advance. Not that it should matter with an outcome so far outside expectations.
You did play 20 games and looked at 11 events at each game. Thus you have in effect played "220 games". All the 1 in 400 events are not really worth mentioning, I think. Especially since the events are not independent and the mathematical dependence is not understood by me fully.

So we have only one very, very rare event. Could be just pure luck/conincidence, or the mathematical dependencies make it more probable than the above model does calculate. I do not know.

Of course it makes a big difference if you set the objective in advance! If you notice during the game that 9s were often and recount them because of this observation and then post it in the forum, this is not very surprising. On BGA hundreds of games are played each day and thousands of possible seldom events are possible noticable. Thus every day some player will notice some 1 in 100,000 event in Catan and might post it in the forum. Your case is different, but still might be just pure coincidence. :-)
muntzer
Posts: 23
Joined: 13 September 2021, 18:12

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by muntzer »

SwHawk wrote: 01 November 2022, 22:35
It seems that I confused two very similar posts, and I apologize for it. Yet I notice you haven't answered to the questions I've asked.
I don't have an infinite amount of time to respond in detail to all replies, so I've prioritised replies where people (i) seem to have read my posts in detail and (ii) raised objections that show an understanding of the maths involved.
muntzer wrote: 01 November 2022, 22:12
Is no one else even a little bit shocked by how far the outliers missed their confidence intervals?
Aren't we falling in two of the bias you've outlined in your first post, namely : Monkey brains and Confirmation bias ? Even though you claim that you started this trial wanting to see whether the implementation is broken or not, it seems you want to prove it IS broken...

Also to put some perspective regarding number, there are around 100 000 Catan games per month (give or take 10 000). So we WOULD expect that event to show up 1 game every 2 months, but real randomness is lumpy, so there could be clusters of that event occurring, and you might be unlucky enough to find yourself in that cluster.
I'm sorry, but honestly, these objections are silly. It's like someone studying links between smoking and cancer. If someone has a suspicion the two things are linked, does that invalidate their trials? Of course not. Instead, you inspect their methodology and analysis and judge their conclusions accordingly. My motivations are irrelevant, unless you suspect me of fiddling the numbers, or deciding to publish AFTER encountering the odd results.

As already stated, I set myself the task of analysing my next 20 games AFTER seeing the thread questioning the randomness of the dice. I set a start and end point in advance. If you don't believe me, that's your unsubstantiated opinion.

Furthermore, I was actually going to defend the BGA dice in general terms. I paused, given people were already making the same sorts of points, often in very supercilious tones. Alpha gamers, anyone?

So, instead, I decided to run a trial. No foregone conclusions. Nothing but probability theory to bolster my results, which are now being reviewed by several people.
My point being that your sample of 20 games, taken once, is, according to me, not enough to PROVE that the implementation is broken. Also, would BGA's implementation be broken, that would most certainly mean that PHP's implementation is broken, which would certainly already have been noticed given the wide range of software (including web applications) relying on that aspect of the language.
All research in the real world works by running trials. You don't study every smoker in the world, you study 100 smokers or 1000 smokers, depending on your budget.

If you find 45 out of 100 smokers subsequently develop cancer, compared to 10 non-smokers out of 100, do you rubbish the study for being too small? If the methodology is rigorous, it's still a statistically significant result. Very much so.

20 trials isn't comprehensive, but my budget - my time and energy - is limited. I can't claim proof that the dice implementation is flawed, but seeing so many anomalous results, particularly the 28 nines, seems like good evidence in that direction. The discussions are ongoing. But your objections would apply to any trial results, no matter how outlandish.

The actual question is: what are the chances you will see such a strange outcome in your next 20 games? If the odds are very, very low, that's not PROOF but it is evidence. That's how testing a hypothesis works.

It's irrelevant how many games of Catan are being played outside a given trial. The trial was deliberately sandboxed. If I had gone looking in the game database for 20 odd games, your objection would be completely warranted.

About PHP: that's a totally separate issue. As a developer, I know that any coding framework can be used well or used poorly. It's possible that the BGA use of PHP libraries is flawed in some way. It could be an environmental issue to do with the way the application is threaded or refreshed in some circumstances. Or maybe a well meaning developer has tried to make the dice rolls in Catan more even than a truly random set, to improve game play, but has inadvertently introduced a bug that has the opposite effect.

Whatever the case, I'm happy to be demonstrated to have made a mistake. What I can't accept is people making a priori judgements that "the dice can't be flawed, so I'm going to ignore any evidence to the contrary".

Sometimes five sevens come up in a row. Sure. But if it happens much more often than it should, there's clear evidence there COULD be a flaw in the BGA implementation. To say otherwise is to shut your eyes and close your mind.
User avatar
SwHawk
Posts: 133
Joined: 23 August 2015, 16:45

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by SwHawk »

muntzer wrote: 02 November 2022, 01:39 All research in the real world works by running trials. You don't study every smoker in the world, you study 100 smokers or 1000 smokers, depending on your budget.
Since you're bringing the subject of trials up, actually one things that clinical trials have to account for is that their sample must be representative. I won't go into details about what criteria make a sample representative or not. But ask yourself, with the criteria you've set, is the sample representative? In my humble opinion, it is at least too small to be representative. I would also say that the level of the players involved make it unrepresentative, but that's another subject entirely. Also, they need to compare to a reference group. The quality of the reference group also influences the outcome of the trial. The reference group you're using is pure mathematics. As other have elaborated, the independence of the events can be questioned. So, in my humble opinion again, there's a problem here as well, as to me, to prove that the implementation is flawed you would have to check against a sample of physical game. Should those events happen with a similar probability on a real game, then there the fault lies with the game itself, rather than the implementation. Should there be a significant difference, in with direction does it lie? Is it fairer than the real game, or is it, as you're trying to show, that the implementation is skewed? If it is fairer then what are the consequences? Should we actually skew the implementation to better reflect the reality? Or should that improved fairness be accepted?
muntzer wrote: 02 November 2022, 01:39 If you find 45 out of 100 smokers subsequently develop cancer, compared to 10 non-smokers out of 100, do you rubbish the study for being too small? If the methodology is rigorous, it's still a statistically significant result. Very much so.
No it actually isn't if the sample isn't representative of what you're trying to observe. It only prompts for a more extensive study, with a larger sample.
muntzer wrote: 02 November 2022, 01:39 20 trials isn't comprehensive, but my budget - my time and energy - is limited. I can't claim proof that the dice implementation is flawed, but seeing so many anomalous results, particularly the 28 nines, seems like good evidence in that direction. The discussions are ongoing. But your objections would apply to any trial results, no matter how outlandish.
Also, you've done this trial once. You've observed 1 substantial anomaly. We can all agree that this anomaly has a non-zero chance of happening. Such is the way with probabilities, if the outcome can exist, then there is a probability it will occur, however small that may be. As euklid314 stated, you had roughly a chance in 100 to observe that event.

But drawing the conclusion that the implementation MUST be flawed because you happened to see it ONCE in your trial is a mistake. Instead of drawing this conclusion from this small a dataset, it should prompt you to study 20 more games. Would a similar event happen in this new sample of 20 more games? If it does, is it still enough? Show us that it happens in ten different samples of 20 games, and then we can start a discussion.
muntzer wrote: 02 November 2022, 01:39 The actual question is: what are the chances you will see such a strange outcome in your next 20 games? If the odds are very, very low, that's not PROOF but it is evidence. That's how testing a hypothesis works.
So for now, this is what this is, it's still a HYPOTHESIS that the system is flawed. You haven't found any definitive proof. You're just drawing conclusions from one small dataset which happens to show an irregularity. This is not a comprehensive study, by your own admission. So what it shows you is that you need to study a bigger sample.
muntzer wrote: 02 November 2022, 01:39 About PHP: that's a totally separate issue. As a developer, I know that any coding framework can be used well or used poorly. It's possible that the BGA use of PHP libraries is flawed in some way. It could be an environmental issue to do with the way the application is threaded or refreshed in some circumstances. Or maybe a well meaning developer has tried to make the dice rolls in Catan more even than a truly random set, to improve game play, but has inadvertently introduced a bug that has the opposite effect.
Then the problem wouldn't lie only with Catan, but with every game relying on some element of randomness. The only games being discussed are dice games, but card games would actually suffer from this as well, since the shuffling algorithm relies on randomness... Yet those threads only appear in dice based game... Also, for such a high profile game like Catan, I'm pretty sure the BGA admins would have reviewed the code surrounding the dice throws. So if the developer wasn't using the recommended functions, the game wouldn't have hit gold release.
muntzer wrote: 02 November 2022, 01:39 Whatever the case, I'm happy to be demonstrated to have made a mistake. What I can't accept is people making a priori judgements that "the dice can't be flawed, so I'm going to ignore any evidence to the contrary".

Sometimes five sevens come up in a row. Sure. But if it happens much more often than it should, there's clear evidence there COULD be a flaw in the BGA implementation. To say otherwise is to shut your eyes and close your mind.
I'm not ignoring the contrary, as you seem to be thinking. I'm just skeptical of results obtained once on a small/non representative sample. I wouldn't have had any problems with what you've written and studied should you have concluded that the results you've obtained prompt for a more extensive study to see if this anomaly was actually what it is, an anomaly with a small chance of appearing, or if the implementation is actually skewed. In fact, I'm asking you to definitely and without a shadow of a doubt prove me wrong. As I've said earlier, what you've observed is an anomaly, so I'm not discarding that fact. What I'm discarding is the conclusion you're trying to make using that fact.

EDIT: Furthermore, drawing the conclusion that the implementation must be flawed from that one dataset would mean that the observed frequency of the event in your dataset (1 in 20 games) is somewhat close from the theoritical frequency of 1 in 230,177, which is far from the truth. But it seems that you're drawing that conclusion anyway...
User avatar
Mathew5000
Posts: 661
Joined: 02 January 2021, 01:41

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by Mathew5000 »

Muntzer, I am impressed by this whole endeavour and your explanation of it. However, your results are not what you think they are. The problem is evident here:
muntzer wrote: 01 November 2022, 12:05 In my actual trials, the actual number of dice rolls per game was almost exactly 100 (1,999 rolls in 20 games = 99.95 per game). The number of rolls per game was clustered very tightly between 97 and 102 rolls: 98, 101, 101, 102, 101, 101, 98, 99, 99, 101, 97, 100, 101, 99, 100, 99, 100, 101, 100 and 101.
In three-player Catan games on BGA, it is uncommon to have the number of dice rolls as high as 100. In my own games, when I check the dice statistics, there are usually between 50 and 80 rolls. I looked at three of your games, and in two of them the number of dice rolls was not close to 100. In the statistics grid of the first game I looked at (https://boardgamearena.com/table?table=309875578), under the row for "number of turns", the entries are 16, 15, and 15 for a total of 46 dice rolls in the game. In the second game I looked at (https://boardgamearena.com/table?table=308936598) each player had 24 turns, for a total of 72 dice rolls. In the third of your games that I looked at (https://boardgamearena.com/table?table=309926046), there were a total of 90 rolls.

So what led you to misinterpret the dice rolls? BGA gives them as a percentage, but omits the % symbol. This is confusing, to say the least. When I first started looking at the dice statistics presented by BGA after each game, I made the same mistake you did. I made a suggestion for improvement a few weeks ago, but so far it has only three votes: https://boardgamearena.com/bug?id=72332

I think I first learned about the issue from the discussion in these bug reports:
https://boardgamearena.com/bug?id=67316
https://boardgamearena.com/bug?id=70809

I suspect that if you go back and re-examine the dice statistics for your 20 games, with the knowledge that there were not around 100 rolls in each of them but much fewer, then the results will not be so far outside the 95% confidence interval.
User avatar
euklid314
Posts: 680
Joined: 06 April 2020, 22:56

Re: Board Game Arena, Catan, virtual dice and randomness - a short trial

Post by euklid314 »

Nice find, Mathew. Problem solved, case closed. :-)

I hope, muntzer will recalculate his 20-game-series with the correct game lengths. His average game length is approx. 60 moves, I did not find any of his games with more than 77 moves, within a short search of approx. 20 of his games.

Probably muntzer will find that not 20 but only 10 outliers outside the 95%-interval did occur. If games have less rolls, a higher deviation is to be expected, i.e. the 95%-confidence intervals are broader. Thus, muntzers 20-game series was a very average series and not a 1%-rare case as my rough calculations would have suggested.

I want to go into more detail with the outliers, that bothered muntzer most:

The three single most striking outliers that muntzer did find were
*) 28% 9s in a game [muntzer gave a 1 in 230,177 probability, wrongly assuming a game of 102 rolls]
*) 1% 10s in a game [muntzer gave a 1 in 471 probability, wrongly assuming a game of 97 rolls]
*) 9% 2s in a game [muntzer game a 1 in 452 probability, wrongly assuming a game of 102 rolls]

Of course it would be nice to know which exact games muntzer did consider (I assume his 20-game series was before 16.10.2022), but lets assume for the moment that above games were 60-roll-games.

*) 28% or more 9s, i.e. 17 or more 9s in a 60-roll game occur in 1 of 4841 games
*) 1% or less 10s, i.e. 1 or less 10s in a 60-roll game occur in 1 of 29 games
*) 9% or more 2s, i.e. 5 or more 2s in a 60-roll game occur in 1 of 39 games

Note, that muntzer did consider 11 different numbers in 20 games, i.e., some (few) events with a probability of one out of 220 games are expected to occur.
Post Reply

Return to “Catan”