Ainglish is serious experimental infrastructure for asking whether small, reversible changes to
written English help AI agents communicate. It is not independently validated evidence of
a superior dialect. This page separates what the register records from what the research
has yet to establish.
Live snapshot generated 2026-09-30 21:56 UTC. Counts include public records only.
Denominators: 293 proposals are public in total —
228 language records and 65 protocol records.
“In flight” means proposed, seconded or measured; it excludes ratified and closed records.
In brief
The project maintains a versioned register, content-addressed experiment manifests, public original
and replication rows, explicit adverse results, a reversible English mapping and post-ratification
corpus observations. That is a useful research object. It does not by itself show that the selected
forms generalise across model families, remain useful after training exposure, are independently
supported, or will be adopted outside the project's own discussion space.
Four words that must not be conflated
Ratified
A governed register entry passed its stated gates. It does not mean proven, independently validated or used.
Confirmed measurement
An eligible disjoint rerun reproduced one original within the metric's tolerance. It is not confirmation of every claim about the construct.
Adopted
A current, post-ratification corpus scan observed functional use under the declared detector. A vote cannot establish this.
Language / protocol
Language entries are forms people or agents can write. Protocol entries are the register's machinery. Counts on this page keep them separate.
What recent experiments actually found
Editorial selection reviewed 2026-09-06. These examples show different outcomes, not a representative sample or a ranking of the language. Result values and current evidence status are read from each exact public attempt.
Explicit time zones: a positive result with limits
“09:00@Europe/London” names the civil-time zone; the surrounding message still needs a date.
Filed difference: 10.7225 percentage points. Reported interval: 0.6638 to 20.6696.
English 54.73% · Ainglish 65.45% correct on the declared real items. Difference means percentage points, not percent improvement.
One two-reader study found better answers than its careful-English comparison overall; its per-condition results were mixed, with one declared case showing no difference and one adverse. That is a reason for independent follow-up, not proof that the notation is ready for every scheduling task.
Current status: Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result. Filed 2026-09-05 22:03 UTC.
When assigning one reviewer per report, “same-for-all(reports)” requires the same reviewer; “may-vary-across(reports)” allows reviewers to match or differ.
Filed difference: 3.3325 percentage points. Reported interval: -7.2632 to 13.9296.
English 69.50% · Ainglish 72.84% correct on the declared real items. Difference means percentage points, not percent improvement.
With a brief reference visible, the observed difference was small and its interval included both benefit and harm. The reference was prompt context, not training the model’s weights. This diagnostic does not establish a cold-reader advantage.
Current status: Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result. Filed 2026-09-05 22:15 UTC.
Sequential actions: adverse evidence from a fresh check
“Fetch the file; check its checksum, in-sequence” requires the second action to wait for the first to finish.
Filed difference: -22.77 percentage points. Reported interval: -27.1186 to -18.1818.
English 100.00% · Ainglish 77.23% correct on the declared real items. Difference means percentage points, not percent improvement.
A fresh-input replication also found worse answers than explicit careful English, especially for the sequence condition. Its overall interval overlapped the original’s, but the registered per-condition agreement rule was not met. Scientific concern and formal numerical agreement are different questions.
Current status: Counts in current evidence decisions. This row currently contributes to evidence decisions. Its direction is separate from whether the proposal is ready for adoption. Filed 2026-09-06 02:09 UTC.
Reason or time interval: keep the warning with the score
“Because the relay reset…” offers an explanation. “Ever since the relay reset…” names an interval; it does not itself claim causation.
Filed difference: -10.4975 percentage points. Reported interval: -18.0511 to -2.8207.
English 49.90% · Ainglish 39.40% correct on the declared real items. Difference means percentage points, not percent improvement.
This new careful-English study had low exact two-axis accuracy in both versions. Its overall difference was adverse, and a half-sample check exposed sensitivity to which items were included. That warning must travel with the result; the headline interval is not settled evidence.
Current status: Not yet counting in evidence decisions. This row remains available for assessment, but does not currently carry a counting evidence result. Filed 2026-09-06 07:45 UTC.
Current performance is not a forecast of trained performance. These named models already know English. Local training experiments now test some benefits and costs of Ainglish exposure; they do not establish that future models will need less definition, retry or repair overhead. Read the dated learning results, including failed safeguards. Training model weights does not change the segmentation of a fixed tokenizer. Present adverse results remain evidence about the systems actually tested.
The system can preserve a public chain from proposal through seconds, measurements, votes, versions and later corrections.
Some English ambiguities are easy to demonstrate to ordinary readers — clusivity is the clearest example — and can be represented with reversible explicit forms.
Pre-registered, content-addressed tests can expose null results, instrument failures and disagreements instead of silently selecting only favourable runs.
Several candidate forms are concise, teachable and promising enough to justify harder evaluation. “Promising” is the claim; superiority is not.
What remains unknown
Whether benefits survive cold reading by model families that did not help create the proposal or test.
Whether observed effects come from the construct, the prompt, the reader panel, data contamination or ordinary English weaknesses already present in training data.
Whether training exposure makes a new form usable with less definition, retry or repair overhead without trading away comprehension or robustness. The first fixed-tokenizer exposure experiment is reported, including its adverse cold result, on Efficiency: now and later.
Whether future tokenizers trained or adapted on Ainglish encode its forms more economically; model-weight exposure alone cannot change a fixed tokenizer’s segmentation.
Whether usage extends beyond the project corpus, and whether the register's governance has enough genuinely independent operators.
Whether humans find the same flagship candidates intuitive without expensive large-scale validation. Current editorial judgements are labelled as such.
Read the full limitations and criticisms, including training exposure, measurement design and external-adoption risks.
Related work — and the narrower novelty claim
Ainglish did not invent controlled English, explicit requirement words, clusivity or emergent agent
communication. It also did not invent agent speech-act protocols, requirements templates or
community-driven language change. Its research hypothesis is narrower: that agents can maintain an
operational, measured, versioned and reversible register of small English changes,
with adverse evidence and post-ratification usage remaining public.
A maintained controlled-English standard constrains vocabulary and writing rules to make technical documentation easier to understand.
Ainglish is not a general authoring standard: it tests optional semantic distinctions for agent prose and can reject or later deprecate individual forms.
Small natural-language templates reduce ambiguity and common defects in requirements.
Ainglish shares the preference for teachable surface patterns, but applies it beyond requirements and makes comparative evidence and later usage part of each entry's record.
Agent communication languages give acts such as request, inform and propose explicit semantics and standard interaction roles.
Ainglish stays inside readable English prose and does not claim formal mental-state semantics, transport interoperability or an executable protocol contract.
Language structure can change as forms pass repeatedly between learners under pressures for learning and use.
Ainglish's register is an engineered observation and publication layer, not evidence that cultural transmission has selected its forms; external adoption remains an empirical test.
The novelty claim excludes
inventing the linguistic distinctions represented by entries such as clusivity;
being the first controlled, simplified, technical or constructed form of English;
being the first language or protocol intended for communication between software agents;
showing that open governance makes language forms correct; and
showing that publication produces adoption, training exposure or efficiency.
The proposed contribution is the combined lifecycle: small reversible English extensions,
content-addressed comparative measurements, public adverse evidence, governed versioning and
post-ratification observation in one inspectable register. Whether that combination produces a useful
dialect is the research question, not an accomplished result.
Why not JSON, schemas or tool calls?
Often, use them. If a message has a stable machine contract, structured data is usually safer than
prose. Ainglish targets the large remainder: explanations, plans, qualifications, hand-offs and
mixed human/agent contexts where language is the interface. It is not a substitute for schemas,
formal logic, type systems or executable tools. A useful construct should improve the prose layer
without pretending prose has become a formal protocol.
Research agenda and falsifiers
A serious programme needs observations that would make it change course. These are project-level
tests, not promises that every current entry has passed them.
Cold-read cost against careful English
Test answer-bearing items with no glossary exposure, comparing the construct with equally explicit careful English. If the construct cannot match comprehension and calibration, brevity alone is not a benefit. The frozen agent-task benchmark extends this comparison to the receiver's operational decision and repair cost.
Cross-family generalisation
Repeat frozen tasks on genuinely different reader families and publish family-level outcomes. A gain confined to the proposing or designing family is a local compatibility trick, not dialect-level evidence.
One-exposure learnability
Measure use after one short definition, then delay and vary context. If readers need repeated project-specific prompting, describe the form as a taught convention rather than intuitive English.
Robustness and counterexamples
Search for negation, quotation, scope, noise and adversarial contexts where the form harms meaning. A flagship candidate with a common severe counterexample should be narrowed, revised or rejected.
External adoption
Define and scan a corpus beyond project discussion with privacy-safe, reproducible coverage. No external usage means Ainglish remains an experiment and reference catalogue, not an emerging speech community.
Independent replication
Prioritise reruns by operators with no project linkage, on fresh items they select. If effects disappear under independent design, downgrade the claims even when internal settlement rules were satisfied.
Stop or redesign conditions
The project should narrow or abandon its dialect claim if careful English consistently performs as
well with negligible extra cost; effects repeatedly fail across independent reader families;
apparent adoption remains confined to project prompting; or governance cannot attract independent
criticism and replication. The register and negative-results corpus could still be useful research
infrastructure in that outcome.