I think Frank and Mark are both right.
Elo was really designed for non-random games and works best there. By best, I mean the stability of the rating and the influence of winstreaks. In games with a high degree of randomness and also with a large k factor (arena), people often end up in places significantly above/below their average elo, which may confuse other people a little. When playing with a 700 elo checkers player, you know that his real level (true elo) is close to this value, however, even in such a serious game as conquering Mars, your 700 elo opponent can play much worse than those who are stable on this rating
FrankJones wrote: ↑04 February 2026, 06:35
Can anyone show me evidence of a game in which a player consistently maintains a 550 Elo but cannot consistently beat 150-Elo opponents?
If the true player ratings are 150 and 550, then by definition the stronger player should win ~90% of the time. However, as noted by Mark, in reality, the formula does not work so accurately for large differences in elo, and the weaker player wins much less, up to <1%. So for me, consistently winning with a difference of 400 elo is never or almost never losing. A game where this is not working is 7 wonders duel (base game), where I once lost to a person 355 (140 vs 495) ratings lower and knew that could happen again. This happends in this particular game because there is a strong randomness factor, and sometimes the strongest moves turn out to be the most intuitive, and even a weak player makes them possibly without even understanding why they are strong.
I haven't delved into this issue, but I know that Go uses a Glicko rating that works almost like elo, only without the strong influence of recent games. As far as I understand, it works like this: you take the last N (N is big number possibly containing all persons games) games of a person and assuming that his opponents played for their real rating, you calculate what the most probable ELo rating you should have in order to achieve those results for these N games