Multiple runs each — is that statistically meaningful?
It is not a poll and we do not present it as one. It is a record of what four systems said, on those questions, on those dates — and we publish how much they disagreed with themselves.
Every question is asked more than once of every engine — three times in the August 2026 edition. That is the difference between a measurement and a screenshot: ask an AI system once and you learn what it said once, with no way to tell a stable answer from a coin toss.
The number we publish because of it
In the August 2026 edition, when an AI system named a vendor at all, it named them again on every repeat of the same question 35% of the time.
That figure is on the edition page on purpose. It is not flattering and it is the most important thing on it: visibility in AI answers is a probability, not a position, and anyone presenting you a single-run ranking as a fixed order is showing you noise with the error bars cropped off.
It is also why month-over-month movement matters more than any one edition, and why movement is measured only over the questions both editions asked. The set does grow, and a question aimed at the wrong market gets retired — but a vendor's movement is computed across the questions the two editions have in common, so your score dropped can never quietly mean he changed the questions. Same questions on both sides, same engines, same number of runs. The noise stays constant from one edition to the next, so the signal is readable.
What it does not claim
Repeating a question is enough to detect instability and not enough to eliminate it. We report the stability figure rather than claiming precision we do not have, and we restate an edition in public when a figure turns out to be wrong. The method is published in full, including the parts that limit it.
If the number changes, the comparison gets published
The run count is a parameter, not a principle. It is fixed before an edition and stated on that edition’s page, so a reader can always see how many runs produced the figures in front of them — three, in August 2026.
If a later edition runs five instead of three, this page will say so, and the difference will be published under research rather than announced as an improvement: whether the extra runs moved any vendor’s score, or only cost more money. A method change that nobody checks is a claim, and we would rather publish the check.