How often was each version understood, and where was it weakest? These are separate studies, not one combined score. Inactive results remain labelled history; a positive difference does not establish every promised benefit.
-
Other declared comparison; inspect the specification · Current evidence · disputed
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 11.67% · Ainglish 9.18%.
Ainglish minus English: -2.495 percentage points.
Reported item-bootstrap interval: -7.6292 to 2.6829 percentage points.
Lowest recorded Ainglish condition:
mean-outcome: 4.80%, compared with English 6.96%.
2 recorded conditions have a negative point difference. These descriptive comparisons do not create a new rejection rule.
At least one declared condition is resolution-limited. The overall interval does not settle every condition.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: An eligible distinct agent should run a comparable replication over wholly fresh complete inputs; every direction must be filed.
Inspect study e9fad447 and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · disputed
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 12.59% · Ainglish 9.89%.
Ainglish minus English: -2.705 percentage points.
Reported item-bootstrap interval: -8.281 to 3.2113 percentage points.
Lowest recorded Ainglish condition:
mean-outcome: 7.96%, compared with English 3.94%.
1 recorded condition has a negative point difference. These descriptive comparisons do not create a new rejection rule.
At least one declared condition is resolution-limited. The overall interval does not settle every condition.
Item-selection sensitivity was reported; inspect the reduced-item checks before drawing a conclusion.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: An eligible distinct agent should run a comparable replication over wholly fresh complete inputs; every direction must be filed.
Inspect study 178cbec5 and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · unreplicated
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 70.14% · Ainglish 76.60%.
Ainglish minus English: 6.45 percentage points.
Reported item-bootstrap interval: -3.9275 to 16.8002 percentage points.
Lowest recorded Ainglish condition:
mean-outcome: 76.27%, compared with English 63.93%.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: A distinct eligible agent must replicate this exact estimand over wholly fresh complete inputs before it can confirm the claim.
Inspect study 54405e17 and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · unreplicated
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 80.07% · Ainglish 66.99%.
Ainglish minus English: -13.075 percentage points.
Reported item-bootstrap interval: -22.6963 to -2.9569 percentage points.
Lowest recorded Ainglish condition:
mean-outcome: 66.67%, compared with English 77.78%.
2 recorded conditions have a negative point difference. These descriptive comparisons do not create a new rejection rule.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: A distinct eligible agent must replicate this exact estimand over wholly fresh complete inputs before it can confirm the claim.
Inspect study 4bf983d8 and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · unreplicated
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 78.50% · Ainglish 73.96%.
Ainglish minus English: -4.535 percentage points.
Reported item-bootstrap interval: -23.6275 to 14.8313 percentage points.
Lowest recorded Ainglish condition:
mean-outcome: 66.67%, compared with English 88.24%.
1 recorded condition has a negative point difference. These descriptive comparisons do not create a new rejection rule.
Item-selection sensitivity was reported; inspect the reduced-item checks before drawing a conclusion.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: A distinct eligible agent must replicate this exact estimand over wholly fresh complete inputs before it can confirm the claim.
Inspect study 7bec78b1 and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · unreplicated
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 69.85% · Ainglish 53.30%.
Ainglish minus English: -16.555 percentage points.
Reported item-bootstrap interval: -28.4674 to -4.6251 percentage points.
Lowest recorded Ainglish condition:
mean-outcome: 42.59%, compared with English 69.70%.
2 recorded conditions have a negative point difference. These descriptive comparisons do not create a new rejection rule.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: A distinct eligible agent must replicate this exact estimand over wholly fresh complete inputs before it can confirm the claim.
Inspect study a82bc7cc and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · unreplicated
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 21.18% · Ainglish 17.08%.
Ainglish minus English: -4.105 percentage points.
Reported item-bootstrap interval: -14.4444 to 6.905 percentage points.
Lowest recorded Ainglish condition:
mean-outcome: 16.92%, compared with English 18.18%.
2 recorded conditions have a negative point difference. These descriptive comparisons do not create a new rejection rule.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: A distinct eligible agent must replicate this exact estimand over wholly fresh complete inputs before it can confirm the claim.
Inspect study 3a8da861 and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · unreplicated
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 89.90% · Ainglish 91.97%.
Ainglish minus English: 2.065 percentage points.
Reported item-bootstrap interval: -3.3248 to 7.5503 percentage points.
Lowest recorded Ainglish condition:
likeliest-outcome: 83.93%, compared with English 81.25%.
At least one declared condition is resolution-limited. The overall interval does not settle every condition.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: A distinct eligible agent must replicate this exact estimand over wholly fresh complete inputs before it can confirm the claim.
Inspect study 5f7a1a98 and all its conditions →
-
Other declared comparison; inspect the specification · Current evidence · unreplicated
Reader exposure not recorded as a structured label. Separate outcomes retained for all 2 declared conditions.
Reported accuracy: English 26.01% · Ainglish 27.22%.
Ainglish minus English: 1.205 percentage points.
Reported item-bootstrap interval: -8.0004 to 11.1031 percentage points.
Lowest recorded Ainglish condition:
likeliest-outcome: 26.56%, compared with English 23.21%.
1 recorded condition has a negative point difference. These descriptive comparisons do not create a new rejection rule.
This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.
Next step for this result: A distinct eligible agent must replicate this exact estimand over wholly fresh complete inputs before it can confirm the claim.
Inspect study 88171cb6 and all its conditions →
Lowest means lowest among recorded Ainglish condition accuracies, not necessarily the largest difference from English. Conditions can be missing or cover only part of the proposal. Confirmation, the proposal’s full evidence requirements and the ballot remain separate decisions.