Ainglish An English dialect for AI agents

Useful today. Built to become familiar.

Efficiency: now and later

English has a head start: today's AI systems have learned from enormous amounts of it. Ainglish is still building that familiarity. We want useful distinctions now, and better communication as future systems learn to use them.

Why a useful distinction can cost more today

A token is a piece of text an AI processes. A familiar English word might fit in one token; a new Ainglish expression might need several. An unfamiliar reader may also need a definition or an extra exchange before it understands the expression. Neither cost is captured simply by counting characters.

For example, we-including-you says explicitly that the reader is part of “we”. It can cost more than bare “we”, but bare “we” does not settle that question. A fair comparison also asks how clearly and cheaply careful English conveys the same meaning—and whether either version prevents a mistaken handoff.

Current results tell us what works with current readers and tokenizers. They do not tell us the eventual efficiency of a system trained to use Ainglish. Equally, English's head start is not an explanation for every disappointing result: a form may simply need improvement. We keep current costs and failures visible, and do not waive proposal checks on the promise of future training. The exact prior exposure of a model is often unknown; we do not assume that every model has seen no Ainglish at all.

How we aim to make Ainglish familiar

  1. Publish reusable language. Ratified definitions and reviewed examples go into versioned public-domain releases. Our training packs make that material easy for dataset and model builders to load, with source versions and checksums.
  2. Teach useful behaviour, not just names. Grow contextual examples, meaning contrasts, multi-turn tasks and cases where ordinary English is enough. Identify generated examples and keep teaching material separate from evaluation answers.
  3. Use the language for real work. Try appropriate distinctions in handoffs, scheduling and reporting; record misunderstandings and explanation overhead. A conversation can demonstrate usefulness, but it does not by itself retrain a model. Try the optional real-use pilot.
  4. Test learning before claiming it. Compare an unadapted model, a short reference in its prompt and controlled training on fresh tasks, with equally informative English controls. Publish unsuccessful results as well as improvements.
  5. Test tokenizers separately. Future tokenizers could encode frequent Ainglish patterns more compactly. Test that opportunity against a fixed vocabulary budget and ordinary-English costs, then test the corresponding model adaptation.
  6. Seek demonstrated downstream use. Work with training projects that accept openly labelled synthetic or instructional language data. Track what they actually include, train on and find useful—not just where a copy is hosted.

How we will know the plan is working

These are separate milestones, not interchangeable meanings of “adopted”.

MilestoneWhat would establish it?
PublishedA versioned, accessible dataset with a clear licence and verifiable files.
Included in a corpusA downstream dataset version or maintainer receipt identifying the included Ainglish material.
Used in trainingA training record identifying the model, data version and amount of exposure.
Demonstrated benefitA held-out comparison reporting understanding, complete task cost and regressions under that training condition.

As checked on , release 3 is available on Hugging Face, Zenodo and the public release repository. Those are publication milestones. Our own learning experiments below are not proof of inclusion in an external lab's training run. Crawling, downloading or mirroring a file is not such proof either.

Three mechanisms, three ledgers

Earlier research snapshot: ; the later matched-learning pilot is dated separately. Studies retain their dataset versions and are product research, not ratification evidence.

1. Surface encoding

What it counts: tokens in one fixed string under one named tokenizer.

What can change it: changing the string or training/adopting a different tokenizer.

What cannot: model-weight training. The model receives token IDs after segmentation.

2. Accommodation and selection

What it counts: glossary text and demonstrations needed before a reader understands and chooses the form.

What can change it: exposure in model training, retrieval or the prompt.

Boundary: exposure must be declared, not inferred from a model name.

3. Correct-outcome interaction

What it counts: every input and output token, retry and repair through a validator-accepted completion.

What can change it: comprehension, selection, verbosity, stopping policy and repair success.

Boundary: cost never travels without first-pass, eventual and unresolved-task counts.

What current tokenizers say

A reproducible census priced all 57 reviewed Ainglish ↔ careful-English pairs in the public v0.35.0 training pack. Negative means the Ainglish rendering used fewer tokens. These are complete reviewed pairs, not isolated display grammars.

TokenizerMean delta, equal pairsPairs cheaper in AinglishRecent-use-weighted proxy
cl100k_base−10.1489.5%−15.99
o200k_base−10.0489.5%−15.56
p50k_base−7.8187.7%−13.23

The first fixed-tokenizer exposure result

One experiment held the Qwen 2.5 7B base revision, 4-bit loading, tokenizer, 19 marker-free glosses and deterministic decoding fixed. The only condition change was a previously frozen two-epoch Ainglish LoRA. A wrong exact-form answer received one authoritative register repair.

This did not show that training exposure removes accommodation cost. Every item still needed a definition/repair turn. The adapter improved exact compliance after that repair, with five base-failure → adapter-success changes and none in reverse. The small token reduction came from shorter outputs and repeated history, not fewer turns. See the frozen prompts, outputs and receipt.

A small matched-English learning pilot

Follow-up work is available. The dated experimental-results collection includes this pilot and a broader learning-and-retention test. The later aggregate gain came with failed family-level safeguards; read both together.

On , we compared one existing model with two small training adaptations: one taught through Ainglish, the other through equally informative English versions of the same tasks. On Ainglish questions without a reference, both trained versions scored 77/96, versus 72/96 for the untouched model. The gain over the base therefore did not establish an Ainglish-specific training benefit.

Results differed by distinction: Ainglish training helped the missing-fact versus missing-decision cases, but harmed participant-inclusion cases. It also failed the predeclared boundary-case retention check. The tokenizer and each prompt's input token count stayed unchanged. The practical lesson is to test retention of already-understood distinctions alongside learning of unfamiliar ones.

This is a small synthetic study on one model, not a result about future foundation models. The frozen figures weight 96 rows containing 84 distinct cases; the duplicate audit and sensitivity analysis are retained. All 1,152 target responses, reference conditions, matched controls, regressions and limitations are in the complete learning-pilot result.

What tokenizer adaptation can change

A separate matched-budget simulation crossed 8k/16k BPE vocabularies, careful-English/Ainglish supplements and two pre-tokenization policies. Every cell received the same 8 MB English core plus a 2 MB supplement with each of the 57 semantic pairs repeated exactly 32 times; neutral filler equalised bytes.

VocabularyHyphen handlingAinglish release-token changeWhole markers gainedHeld-out English change
8,000punctuation split−940−164
8,000hyphen unit visible−36810+364
16,000punctuation split−930−23
16,000hyphen unit visible−58516+237

The mechanism is visible: vocabulary exposure cannot create a whole-marker token when the pre-tokenizer forbids merges across hyphens. A hyphen-aware design can, but vocabulary slots have opportunity cost. This small simulation is not a prediction that a production tokenizer will choose the same trade-off. Inspect all eight tokenizers in the reproducible lab.

Where the boundary is enforced

A dated claim audit pinned 17 files across the website/server, language releases, SDK and both agent plugins. It found no statement that model-weight exposure changes a fixed tokenizer and no statement that publication proves training, adoption or benefit. It did retain one editorial risk: the broad “optimised for clearer and more efficient” mission line can sound like a completed result when detached from the methodology. Read the 17-file reproducible audit.

What follows

  • Now: compare complete Ainglish and careful-English utterances under named current tokenizers. Keep adverse rows visible.
  • For model builders: put the CC0 training pack into weight-training experiments, but use held-out tasks and report exposure composition.
  • For tokenizer builders: test marker-aware segmentation against a fixed general-English regression budget; do not optimise one vocabulary in isolation.
  • For project claims: call exposure unestablished unless a training receipt exists. Ordinary English has the incumbent-data advantage; that is a hypothesis boundary, not a waiver for adverse Ainglish results.
  • For flagship selection: avoid gratuitous segment count as a tie-breaker. Comprehension, ambiguity removed and robustness remain primary.

Developing a complete task-cost measurement

An August 2026 discussion and preflighted draft proposed an interaction_cost_delta receipt: paired tasks, a frozen validator, every token through a correct outcome, unresolved tasks, stopping/censor policy and a per-model exposure receipt. In response to public review, failure remains a co-primary count with no invented token penalty; a cheaper arm cannot carry support after failing its correctness gate, and traffic weighting remains a separate projection. Read the public protocol discussion and its revised valid preflight receipt.