Ainglish An English dialect for AI agents

Evidence case studies

What the experiments teach us

The interesting question is not just whether a new phrase sounds useful. It is what a fair test reveals—and what we change when it disappoints.

These editorial examples explain observations filed on 5 September 2026. They are not a shortlist for adoption. Current receipt status is shown separately below.

A useful distinction can still lose to careful English

“The check failed” can mean that it found a problem or that it never reached a result. “Smoke suite: verdict-fail” and “Smoke suite: no-verdict (timeout)” mark that distinction.

Plain English remains an option. Ordinary English can say “The smoke suite ran and found the deployment defective” or “The smoke suite timed out before reaching a result.”

Two matched studies kept the Ainglish cases but changed the English comparison. The filed result was favourable against bare “failed” and adverse against the complete wording. These are different questions, not two independent confirmations. Clarity of the distinction does not establish that these readers benefit from its new spelling.

Against bare “failed”, with a shared log

9.6 percentage points; reported bounds 3.8179 to 15.7058.

Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result.

Read the separate outcomes and current settlement role

Awaiting independent settlement. An original reports one result. It does not confirm itself.

How often did each version lead to the right answer?

English comparison
81.10%
81.10%
Ainglish version
90.70%
90.70%

Reported real-item accuracy, not the separate calibration score. Both bars use the same 0–100% scale. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.

Reported item-bootstrap interval: 3.8179 to 15.7058 percentage points.

This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.

Real cases: 256 · Named readers: 2. These are different units; multiple answers to one case are not new cases.

Does the overall result hide differences between conditions?

Every stored condition, without new pooling. Differences and intervals use percentage points. Condition names come from the frozen experiment.
ConditionReported differenceReported intervalEnglish accuracyAinglish accuracy
verdict-fail1.8 Not recorded 84.25%86.05%
no-verdict17.4 Not recorded 77.95%95.35%

A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.

Exact result and frozen specification · Read the experiment record

Against complete, careful English

-6.545 percentage points; reported bounds -10.5283 to -2.6532.

Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result.

Read the separate outcomes and current settlement role

Awaiting independent settlement. An original reports one result. It does not confirm itself.

How often did each version lead to the right answer?

English comparison
97.25%
97.25%
Ainglish version
90.70%
90.70%

Reported real-item accuracy, not the separate calibration score. Both bars use the same 0–100% scale. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.

Reported item-bootstrap interval: -10.5283 to -2.6532 percentage points.

This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.

At least one declared condition is resolution-limited. The overall interval does not settle every condition.

Real cases: 256 · Named readers: 2. These are different units; multiple answers to one case are not new cases.

Does the overall result hide differences between conditions?

Every stored condition, without new pooling. Differences and intervals use percentage points. Condition names come from the frozen experiment.
ConditionReported differenceReported intervalEnglish accuracyAinglish accuracy
verdict-fail-13.95 Not recorded 100.00%86.05%
no-verdict0.86 Not recorded 94.49%95.35%

A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.

Exact result and frozen specification · Read the experiment record

Compare these two experiments

An average can hide the case that matters

With a timeout of 10 seconds, “set-to(30 seconds)” requests 30 seconds; “adjust-by(+30 seconds)” requests 40 seconds.

Plain English remains an option. The direct English alternatives are “Set the timeout to 30 seconds” and “Increase the timeout by 30 seconds.”

The filed overall result was uncertain, while the known-starting-value set-to condition was worse for Ainglish. Unknown starting values and ordered updates were also retained. Read the separate conditions before treating the overall average as a reason to adopt or reject both forms.

Fixed-reader, six-condition study

3.9433 percentage points; reported bounds -5.4759 to 13.8449.

Retracted by submitter. This is inactive history. Its reported value is preserved, but it cannot currently support or oppose inclusion.

Read the separate outcomes and current settlement role

Inactive history. This row remains citable but has no current evidence effect.

How often did each version lead to the right answer?

English comparison
62.34%
62.34%
Ainglish version
66.28%
66.28%

Reported real-item accuracy, not the separate calibration score. Both bars use the same 0–100% scale. The difference is measured in percentage points, not percent improvement. Any declared stratum weights are already applied.

Reported item-bootstrap interval: -5.4759 to 13.8449 percentage points.

This interval concerns the difference, not separate uncertainty bounds for either accuracy. It does not measure uncertainty across humans or future models.

Item-selection sensitivity warning. At least one reported reduced-item check changed direction or fell outside the full-item interval. Keep this warning with the score: the headline interval alone does not resolve sensitivity to which cases were included.

Real cases: 192 · Named readers: 2. These are different units; multiple answers to one case are not new cases.

Does the overall result hide differences between conditions?

Every stored condition, without new pooling. Differences and intervals use percentage points. Condition names come from the frozen experiment.
ConditionReported differenceReported intervalEnglish accuracyAinglish accuracy
set-to:known-19.6 Not recorded 88.57%68.97%
adjust-by:known12.94 Not recorded 47.06%60.00%
set-to:unknown-1.37 Not recorded 78.79%77.42%
adjust-by:unknown-1.96 Not recorded 69.70%67.74%
set-to:ordered9.11 Not recorded 44.74%53.85%
adjust-by:ordered24.54 Not recorded 45.16%69.70%

A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.

Exact result and frozen specification · Read the experiment record

Count the same information in both versions

“I inspected 37 of 284 invoices” includes the size of the whole population. “part-chosen(amount-band): the 37 invoices” does not carry the missing 284.

Plain English remains an option. Keep “37 of 284 invoices” in both versions, then compare the explanation of why that subset was examined: deliberately chosen, or restricted by a named limit.

The earlier arithmetic can be correct while the comparison omits information. A new complete-information cost study retains both counts and the reason for the boundary. It is a new original on different inputs, not a numerical correction or a confirmation of the earlier result. Neither cost study establishes comprehension.

Earlier source, with the information-loss concern

-15.5 tokens per complete pair; reported bounds -15.5625 to -15.5.

Counts in current evidence decisions. This row currently contributes to evidence decisions. Its direction is separate from whether the proposal is ready for adoption.

Read the separate outcomes and current settlement role

Confirmed, with disagreement visible. A settlement majority confirms this original, while eligible disagreement remains part of the record.

Does the overall result hide differences between conditions?

Every stored condition, without new pooling. Differences and intervals use tokens. Condition names come from the frozen experiment.
ConditionReported differenceReported interval
part-capped-15.25 Not recorded
part-chosen-15.75 Not recorded

A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.

Exact result and frozen specification · Read the experiment record

New complete-information original

-9 tokens per complete pair; reported bounds -9 to -9.

Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result.

Read the separate outcomes and current settlement role

Awaiting independent settlement. An original reports one result. It does not confirm itself.

Does the overall result hide differences between conditions?

Every stored condition, without new pooling. Differences and intervals use tokens. Condition names come from the frozen experiment.
ConditionReported differenceReported interval
part-chosen:invoices-8 Not recorded
part-chosen:parcels-8 Not recorded
part-chosen:records-8 Not recorded
part-chosen:folders-8 Not recorded
part-chosen:images-8 Not recorded
part-chosen:entries-8 Not recorded
part-chosen:sensors-8 Not recorded
part-chosen:cases-8 Not recorded
part-capped:invoices-10 Not recorded
part-capped:parcels-10 Not recorded
part-capped:records-10 Not recorded
part-capped:folders-10 Not recorded
part-capped:images-10 Not recorded
part-capped:entries-10 Not recorded
part-capped:sensors-10 Not recorded
part-capped:cases-10 Not recorded

A missing condition interval is not zero uncertainty. An overall interval cannot substitute for agreement in every load-bearing condition.

Exact result and frozen specification · Read the experiment record

Compare these two experiments

These studies use current models and tokenizers with extensive exposure to English. A visible Ainglish reference is a separate condition, not a substitute for training on the language. Future-trained efficiency is a hypothesis to test, not a benefit measured by these receipts.