English has a head start: today's AI systems have learned from enormous amounts of it.
Ainglish is still building that familiarity. We want useful distinctions now, and better
communication as future systems learn to use them.
Why a useful distinction can cost more today
A token is a piece of text an AI processes. A familiar English word might
fit in one token; a new Ainglish expression might need several. An unfamiliar reader may
also need a definition or an extra exchange before it understands the expression. Neither
cost is captured simply by counting characters.
For example, we-including-you says explicitly that the reader is part of
“we”. It can cost more than bare “we”, but bare “we” does not settle that question.
A fair comparison also asks how clearly and cheaply careful English conveys the
same meaning—and whether either version prevents a mistaken handoff.
Current results tell us what works with current readers and tokenizers. They do not tell
us the eventual efficiency of a system trained to use Ainglish. Equally, English's head
start is not an explanation for every disappointing result: a form may simply need
improvement. We keep current costs and failures visible, and do not waive proposal checks
on the promise of future training. The exact prior exposure of a model is often unknown;
we do not assume that every model has seen no Ainglish at all.
How we aim to make Ainglish familiar
Publish reusable language. Ratified definitions and reviewed examples
go into versioned public-domain releases.
Our training packs make that material easy for
dataset and model builders to load, with source versions and checksums.
Teach useful behaviour, not just names. Grow contextual examples,
meaning contrasts, multi-turn tasks and cases where ordinary English is enough.
Identify generated examples and keep teaching material separate from evaluation answers.
Use the language for real work. Try appropriate distinctions in
handoffs, scheduling and reporting; record misunderstandings and explanation overhead.
A conversation can demonstrate usefulness, but it does not by itself retrain a model.
Try the optional real-use pilot.
Test learning before claiming it. Compare an unadapted model, a short
reference in its prompt and controlled training on fresh tasks, with equally informative
English controls. Publish unsuccessful results as well as improvements.
Test tokenizers separately. Future tokenizers could encode frequent
Ainglish patterns more compactly. Test that opportunity against a fixed vocabulary
budget and ordinary-English costs, then test the corresponding model adaptation.
Seek demonstrated downstream use. Work with training projects that
accept openly labelled synthetic or instructional language data. Track what they
actually include, train on and find useful—not just where a copy is hosted.
How we will know the plan is working
These are separate milestones, not interchangeable meanings of “adopted”.
Milestone
What would establish it?
Published
A versioned, accessible dataset with a clear licence and verifiable files.
Included in a corpus
A downstream dataset version or maintainer receipt identifying the included Ainglish material.
Used in training
A training record identifying the model, data version and amount of exposure.
Demonstrated benefit
A held-out comparison reporting understanding, complete task cost and regressions under that training condition.
As checked on , release 3 is available
on Hugging Face,
Zenodo and the
public release repository.
Those are publication milestones. Our own learning experiments below are not proof of
inclusion in an external lab's training run. Crawling, downloading or mirroring a file
is not such proof either.
Three mechanisms, three ledgers
Earlier research snapshot: ;
the later matched-learning pilot is dated separately. Studies retain their dataset versions and are product
research, not ratification evidence.
1. Surface encoding
What it counts: tokens in one fixed string under one named tokenizer.
What can change it: changing the string or training/adopting a different tokenizer.
What cannot: model-weight training. The model receives token IDs after segmentation.
2. Accommodation and selection
What it counts: glossary text and demonstrations needed before a reader understands and chooses the form.
What can change it: exposure in model training, retrieval or the prompt.
Boundary: exposure must be declared, not inferred from a model name.
3. Correct-outcome interaction
What it counts: every input and output token, retry and repair through a validator-accepted completion.
What can change it: comprehension, selection, verbosity, stopping policy and repair success.
Boundary: cost never travels without first-pass, eventual and unresolved-task counts.
What current tokenizers say
A reproducible census priced all 57 reviewed Ainglish ↔ careful-English pairs in the public
v0.35.0 training pack. Negative means the Ainglish rendering used fewer tokens. These are
complete reviewed pairs, not isolated display grammars.
Tokenizer
Mean delta, equal pairs
Pairs cheaper in Ainglish
Recent-use-weighted proxy
cl100k_base
−10.14
89.5%
−15.99
o200k_base
−10.04
89.5%
−15.56
p50k_base
−7.81
87.7%
−13.23
The first fixed-tokenizer exposure result
One experiment held the Qwen 2.5 7B base revision, 4-bit loading, tokenizer, 19 marker-free
glosses and deterministic decoding fixed. The only condition change was a previously frozen
two-epoch Ainglish LoRA. A wrong exact-form answer received one authoritative register repair.
This did not show that training exposure removes accommodation cost. Every item
still needed a definition/repair turn. The adapter improved exact compliance after that repair,
with five base-failure → adapter-success changes and none in reverse. The small token reduction
came from shorter outputs and repeated history, not fewer turns. See the
frozen prompts, outputs and receipt.
A small matched-English learning pilot
Follow-up work is available. The dated experimental-results collection includes this pilot and a broader learning-and-retention test. The later aggregate gain came with failed family-level safeguards; read both together.
On , we compared one existing model
with two small training adaptations: one taught through Ainglish, the other through
equally informative English versions of the same tasks. On Ainglish questions without
a reference, both trained versions scored 77/96, versus 72/96 for the
untouched model. The gain over the base therefore did not establish an Ainglish-specific
training benefit.
Results differed by distinction: Ainglish training helped the missing-fact versus
missing-decision cases, but harmed participant-inclusion cases. It also failed the
predeclared boundary-case retention check. The tokenizer and each prompt's input token
count stayed unchanged. The practical lesson is to test retention of already-understood
distinctions alongside learning of unfamiliar ones.
This is a small synthetic study on one model, not a result about future foundation models.
The frozen figures weight 96 rows containing 84 distinct cases; the duplicate audit and
sensitivity analysis are retained. All 1,152 target responses, reference conditions,
matched controls, regressions and limitations are in the
complete learning-pilot result.
What tokenizer adaptation can change
A separate matched-budget simulation crossed 8k/16k BPE vocabularies, careful-English/Ainglish
supplements and two pre-tokenization policies. Every cell received the same 8 MB English core
plus a 2 MB supplement with each of the 57 semantic pairs repeated exactly 32 times; neutral
filler equalised bytes.
Vocabulary
Hyphen handling
Ainglish release-token change
Whole markers gained
Held-out English change
8,000
punctuation split
−94
0
−164
8,000
hyphen unit visible
−368
10
+364
16,000
punctuation split
−93
0
−23
16,000
hyphen unit visible
−585
16
+237
The mechanism is visible: vocabulary exposure cannot create a whole-marker token when the
pre-tokenizer forbids merges across hyphens. A hyphen-aware design can, but vocabulary slots
have opportunity cost. This small simulation is not a prediction that a production tokenizer
will choose the same trade-off. Inspect all eight tokenizers in the
reproducible lab.
Where the boundary is enforced
A dated claim audit pinned 17 files across the website/server, language releases, SDK and
both agent plugins. It found no statement that model-weight exposure changes a fixed
tokenizer and no statement that publication proves training, adoption or benefit. It did
retain one editorial risk: the broad “optimised for clearer and more efficient” mission line
can sound like a completed result when detached from the methodology. Read the
17-file reproducible audit.
What follows
Now: compare complete Ainglish and careful-English utterances under named current tokenizers. Keep adverse rows visible.
For model builders: put the CC0 training pack into weight-training experiments, but use held-out tasks and report exposure composition.
For tokenizer builders: test marker-aware segmentation against a fixed general-English regression budget; do not optimise one vocabulary in isolation.
For project claims: call exposure unestablished unless a training receipt exists. Ordinary English has the incumbent-data advantage; that is a hypothesis boundary, not a waiver for adverse Ainglish results.
For flagship selection: avoid gratuitous segment count as a tie-breaker. Comprehension, ambiguity removed and robustness remain primary.
Developing a complete task-cost measurement
An August 2026 discussion and preflighted draft proposed an interaction_cost_delta receipt:
paired tasks, a frozen validator, every token through a correct outcome, unresolved tasks,
stopping/censor policy and a per-model exposure receipt. In response to public review, failure
remains a co-primary count with no invented token penalty; a cheaper arm cannot carry support
after failing its correctness gate, and traffic weighting remains a separate projection. Read the
public protocol discussion
and its revised valid preflight receipt.