How the Consensus score works

Consensus is not a headcount. Every alg site and sheet is scored first, on how close its own ranking of a case is to the order top solvers were reconstructed executing, and a resource that leads with what the field really runs counts for several times more than one that doesn't. Some sets rank by it because their own usage is too thin to order a table on.

The resulting distribution of Consensus

Simplified example

Three invented resources, cut down to show the mechanism. A real case has far more of them, and the last figure on this page draws one. Every ribbon here is one mention, as thick as the credit it paid: thick where a resource ranks an algorithm first, thin where it ranks it third. Drag a resource's agreement and watch the case re-form around it - below 33% it leaves the arithmetic entirely.

Resource A ranks it #1, paying 10.0 creditResource B ranks it #2, paying 2.0 creditResource C ranks it #2, paying 1.8 creditResource C ranks it #1, paying 3.7 creditResource A ranks it #3, paying 3.3 creditResource A ranks it #2, paying 5.0 creditResource B ranks it #1, paying 4.1 creditResource A10.0 credit, 93% agreementResource B4.1 credit, 66% agreementResource C3.7 credit, 64% agreement(U2) R' U R' F2 R F' R46% consensus(y2) L U' R U2 L' U R'23% consensusR U' R F2 R' U R'17% consensusR' F R' F2 R U' R14% consensus
93% agreement
66% agreement
64% agreement

The case and its algorithms are real (an Ortega PBL); the three resources and their scores are invented, and no real resource's standing is shown anywhere on this page. A longer list pays out more in total, because each mention pays independently.

How it gets there

Four steps. Open any one for its arithmetic.

1 Score every resource against real solves. Each site and sheet is measured on how close its ranking is to what solvers actually execute.

Three resources, auditioned on five real 3x3 cases carrying 2,921 reconstructed executions. Each column is one case, as wide as that case is used and as full as the resource's ranking of it deserved.

Resource A Leads with what fast solvers actually run judged on 2,921 executions
Ua
Ub
T
OLL 26
Ra
92.7% of the credit its ranking could have earned on the cases it covers
Resource B Strong on some cases, weak on others judged on 2,921 executions
Ua
Ub
T
OLL 26
Ra
65.9% of the credit its ranking could have earned on the cases it covers
Resource C Covers less, and leads with the wrong one judged on 1,824 executions across the 3 cases it covers
Ua
T
OLL 26
63.7% of the credit its ranking could have earned on the cases it covers

Everybody looks good on the T perm, where 92% of executions are one algorithm. The U perms decide it: solvers overwhelmingly run the R-move versions, so a resource that leads with the slice algorithms is giving sound advice the field does not take, across 1,433 executions. Resource B has the right algorithm in its list and still loses both, because a reader takes the one at the top.

What that score is worth, as credit to spend

Agreement compresses - almost every scorable resource lands between 50% and 90% - so the score has a floor subtracted and the remainder curved. At or below 33% a resource earns no credit at all: no citation, and no place in any case's total.

93% 10.0 credit
66% 4.1 credit
64% 3.7 credit

A 29-point spread in agreement becomes a 2.7x spread in credit - and these three are close together. Across every resource scored, best to worst-still-standing runs past thirty times.

2 Take a set with too few solves to rank. Every resource covering the case is read as an ordered list, matched by the state it solves.

The audition happened on 3x3, where reconstructions are plentiful. What those three earned is spent here on an Ortega PBL case, which almost nobody has been reconstructed solving - there is no usage to rank it by. That gap is the whole reason the metric exists.

Resource A 10.0 credit
  1. #1 (U2) R' U R' F2 R F' R
  2. #2 R U' R F2 R' U R'
  3. #3 (y2) L U' R U2 L' U R'
Resource B 4.1 credit
  1. #1 R' F R' F2 R U' R
  2. #2 (U2) R' U R' F2 R F' R
Resource C 3.7 credit
  1. #1 (y2) L U' R U2 L' U R'
  2. #2 (U2) R' U R' F2 R F' R
3 A top pick counts most. A mention pays the resource's credit over its rank, so naming one good algorithm beats padding.

Full, half, a third

#1 full
#2 half
#3 a third

The same ladder, at each resource's credit

Resource A
#1
#2
#3
Resource B
#1
#2
Resource C
#1
#2

Drawn to scale, the weakest resource's first choice is worth less than the strongest resource's third.

Length is not an advantage in itself. A resource that names one algorithm and nothing else is judged in step 1 on exactly that algorithm, and scores full marks if it is the one most solvers run, while a fourth pick almost nobody executes costs its author more than the sliver of influence it buys here.

4 Normalise per case. Each algorithm's credit over the case's total. That is the value in the Consensus column.
AlgorithmWho paid, and at what rankScoreConsensus
(U2) R' U R' F2 R F' R
#1#2#2
13.946%
(y2) L U' R U2 L' U R'
#1#3
7.023%
R U' R F2 R' U R'
#2
5.017%
R' F R' F2 R U' R
#1
4.114%

(U2) R' U R' F2 R F' R leads not because more resources mention it, but because the resources that predict real usage best put it first. Shares need not reach 100%: a resource can recommend an algorithm we do not stock, and it still counts in the case's total. The live column blends every resource that covers the case, not three.

One real case, at full width

Everything above ran on three resources so the arithmetic would fit in a diagram. Here is a real one: 9 resources recommending 6 algorithms for a single case, at the percentages the algorithm database prints. Which case it is, and which algorithms, are left off on purpose - the reading here is the shape, and the shape needs no names. A lot of thin, scattered agreement; two algorithms taking most of it.

One resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithmOne resource's recommendation of one algorithm9 resourcesunnamedAlgorithm 135% consensus · 7 of 9 list itAlgorithm 228% consensus · 7 of 9 list itAlgorithm 317% consensus · 6 of 9 list itAlgorithm 412% consensus · 4 of 9 list itAlgorithm 53% consensus · 2 of 9 list itAlgorithm 62% consensus · 1 of 9 list it

Ribbon thickness splits each algorithm's percentage across the resources that list it by rank: what any individual resource is worth stays internal and is never drawn. A further 2 algorithms sit below 2% and are left out of the picture. Deeper cases exist - the most-covered one in the database carries 24 resources - but only the sets ranked by Consensus publish enough structure to draw the fan.