How the Consensus score works
Consensus is not a headcount. Every alg site and sheet is scored first, on how close its own ranking of a case is to the order top solvers were reconstructed executing, and a resource that leads with what the field really runs counts for several times more than one that doesn't. Some sets rank by it because their own usage is too thin to order a table on.
The resulting distribution of Consensus
Simplified exampleThree invented resources, cut down to show the mechanism. A real case has far more of them, and the last figure on this page draws one. Every ribbon here is one mention, as thick as the credit it paid: thick where a resource ranks an algorithm first, thin where it ranks it third. Drag a resource's agreement and watch the case re-form around it - below 33% it leaves the arithmetic entirely.
The case and its algorithms are real (an Ortega PBL); the three resources and their scores are invented, and no real resource's standing is shown anywhere on this page. A longer list pays out more in total, because each mention pays independently.
How it gets there
Four steps. Open any one for its arithmetic.
1 Score every resource against real solves. Each site and sheet is measured on how close its ranking is to what solvers actually execute.
Three resources, auditioned on five real 3x3 cases carrying 2,921 reconstructed executions. Each column is one case, as wide as that case is used and as full as the resource's ranking of it deserved.
Everybody looks good on the T perm, where 92% of executions are one algorithm. The U perms decide it: solvers overwhelmingly run the R-move versions, so a resource that leads with the slice algorithms is giving sound advice the field does not take, across 1,433 executions. Resource B has the right algorithm in its list and still loses both, because a reader takes the one at the top.
What that score is worth, as credit to spend
Agreement compresses - almost every scorable resource lands between 50% and 90% - so the score has a floor subtracted and the remainder curved. At or below 33% a resource earns no credit at all: no citation, and no place in any case's total.
A 29-point spread in agreement becomes a 2.7x spread in credit - and these three are close together. Across every resource scored, best to worst-still-standing runs past thirty times.
2 Take a set with too few solves to rank. Every resource covering the case is read as an ordered list, matched by the state it solves.
The audition happened on 3x3, where reconstructions are plentiful. What those three earned is spent here on an Ortega PBL case, which almost nobody has been reconstructed solving - there is no usage to rank it by. That gap is the whole reason the metric exists.
- #1 (U2) R' U R' F2 R F' R
- #2 R U' R F2 R' U R'
- #3 (y2) L U' R U2 L' U R'
- #1 R' F R' F2 R U' R
- #2 (U2) R' U R' F2 R F' R
- #1 (y2) L U' R U2 L' U R'
- #2 (U2) R' U R' F2 R F' R
3 A top pick counts most. A mention pays the resource's credit over its rank, so naming one good algorithm beats padding.
Full, half, a third
The same ladder, at each resource's credit
Drawn to scale, the weakest resource's first choice is worth less than the strongest resource's third.
Length is not an advantage in itself. A resource that names one algorithm and nothing else is judged in step 1 on exactly that algorithm, and scores full marks if it is the one most solvers run, while a fourth pick almost nobody executes costs its author more than the sliver of influence it buys here.
4 Normalise per case. Each algorithm's credit over the case's total. That is the value in the Consensus column.
| Algorithm | Who paid, and at what rank | Score | Consensus |
|---|---|---|---|
| (U2) R' U R' F2 R F' R | #1#2#2 | 13.9 | 46% |
| (y2) L U' R U2 L' U R' | #1#3 | 7.0 | 23% |
| R U' R F2 R' U R' | #2 | 5.0 | 17% |
| R' F R' F2 R U' R | #1 | 4.1 | 14% |
(U2) R' U R' F2 R F' R leads not because more resources mention it, but because the resources that predict real usage best put it first. Shares need not reach 100%: a resource can recommend an algorithm we do not stock, and it still counts in the case's total. The live column blends every resource that covers the case, not three.
One real case, at full width
Everything above ran on three resources so the arithmetic would fit in a diagram. Here is a real one: 9 resources recommending 6 algorithms for a single case, at the percentages the algorithm database prints. Which case it is, and which algorithms, are left off on purpose - the reading here is the shape, and the shape needs no names. A lot of thin, scattered agreement; two algorithms taking most of it.
Ribbon thickness splits each algorithm's percentage across the resources that list it by rank: what any individual resource is worth stays internal and is never drawn. A further 2 algorithms sit below 2% and are left out of the picture. Deeper cases exist - the most-covered one in the database carries 24 resources - but only the sets ranked by Consensus publish enough structure to draw the fan.