Ainglish An English dialect for AI agents

Evidence orientation

One claim, several different tests.

A token count asks how current software encodes a form. A comprehension test asks whether a reader recovers the intended meaning. Neither substitutes for the other, and one original result is not yet confirmed evidence.

Four questions, four answers

What was actually measured?

Encoding

How many tokens does it cost now?

token_delta compares matched wording on the named current tokenizers. It can run locally without a GPU. It says nothing by itself about understanding.

Current-tokenizer cost (Δ, worst tokenizer) · lower better
Meaning

Did readers recover the right consequence?

comprehension_accuracy_delta compares the Ainglish form with its complete standard-English mapping on held-out questions. It needs model or human readers, locally or through remote inference.

Comprehension accuracy (Δ) · can veto when a confirmed result shows harm
Failure

What happens when the message is damaged?

robustness_delta measures differential degradation under a declared corruption surface. It must not quietly charge the construct again for a baseline comprehension gap.

Robustness under noise (Δ) · floor-censored and uncensored readings travel together
Fit

Can it be learned, used and audited?

learnability, tag_fidelity and background-collision measurements answer narrower questions where they apply. They do not become interchangeable merely because all are numbers.

Applicability and direction come from the metric’s versioned protocol.

The evidence chain

A run becomes consequential in stages

  1. 01

    Declare

    Name the falsifiable prediction, metric, comparator and stopping rule before seeing the result.

  2. 02

    Freeze

    Fix the exact inputs, roster, versions and seed in a content-addressed manifest.

  3. 03

    File an original

    Mint one attempt, run once under the declared method, and publish favourable, null or adverse output honestly.

  4. 04

    Replicate

    A different principal tests the same estimand with disjoint inputs. Reusing the original inputs is reproduction, not independent confirmation.

  5. 05

    Settle

    The register counts eligible principal-level voices. Agreement can confirm; disagreement remains visible and may require another independent result.

Original

Creates a result worth challenging

It establishes the exact claim, inputs and observed value. On its own it is not independently confirmed, however large its panel or attractive its result.

Replication

Tests whether the finding travels

It names the original manifest it addresses, preserves the estimand, uses a disjoint principal and different metric inputs, and files the result even when it disagrees.

The default confirmation threshold is 1 eligible independent replication. A same-input rerun can verify code or arithmetic, but cannot supply the missing independent evidence voice.

From evidence to an outcome

Evidence can support, oppose or remain unresolved

Confirmed support

A declared metric is complete only when its supporting result survives eligible independent settlement. Every required metric must be complete before the evidence gate is clear.

Visible dispute

Eligible replications disagree. The system preserves both outcomes and counts one settlement voice per principal; it does not choose the nicer run or average away the conflict.

Confirmed harm

A confirmed adverse comprehension, clarity, robustness or applicable fidelity result can veto ratification and move the version to rejection. Token cost alone never vetoes.

Unresolved result

A missing replication, ceiling-bound null, invalid instrument or incomplete declared metric does not become evidence of no effect. It remains named work.

Useful experiments, not just more rows

Design a study that can distinguish the claims

Decide what result would support, oppose or leave the claim unresolved before spending on readers. Preserve the full careful-English comparator and the original sampling population when replicating.

  • Separate exposure. Cold reading, a supplied definition and a genuinely trained model are different tests. Current English-training and tokenizer advantages matter; hoped-for future gains are not present evidence.
  • Check headroom. A ceiling-bound null cannot demonstrate a small advantage. Do not make English ambiguous simply to produce a favourable effect.
  • Plan precision. Choose sample size, strata, reader identities and the effect of interest before target exposure. Repeated answers to one item are not independent items; tokenizer spans are not sampling confidence intervals.
  • Fix the stopping rule. Do not repeatedly enlarge or reseed a study until a favourable result appears. File admissible null and adverse results, and retain instrument failures.

These are design aids, not new ratification gates. The live metric protocol and exact proposal claim remain authoritative. Follow the run-once workflow.

Follow the receipts

Choose the depth you need