Ainglish An English dialect for AI agents

For model builders, corpus maintainers and researchers

Training data

Small, high-signal public-domain datasets that make ratified Ainglish easy to put into ordinary language-model pipelines. Each immutable pack is generated from one frozen language release, carries stable row IDs and source digests, and is available as JSONL, Parquet, Dolma and Croissant metadata.

Current pack: v4

32ratified constructs
66reviewed pairs
164instruction rows
32pretraining documents

Bound to ainglish-core-v4 at register digest 964ea4c3478be1755dcc6dfa82830131112c7ddd0262c30cc628af3f071d6489. All rows are in the train split. Canonical and reviewed non-normative examples remain visibly distinct; no measurement answers, evaluation holdouts, conversations or contributor identities are included.

Download parallel Parquet Download Dolma shard Read the datasheet

Load it

from datasets import load_dataset

base = "https://ainglish.org/training/ainglish-training-v4"
parallel = load_dataset("parquet", data_files=f"{base}/data/parquet/parallel.parquet", split="train")
instructions = load_dataset("json", data_files=f"{base}/data/instruction.jsonl", split="train")

MLCommons discovery metadata is at metadata/croissant.json. Verify a mirror with SHA256SUMS and the pack MANIFEST.json.

Loading with the Croissant Python reader

The mlcroissant 1.1.0 reader recognises an older Parquet MIME label. This example adjusts that label only in memory; it does not change the published metadata, dataset files or their checksums. Iterating records also checks that the data downloads work; parsing metadata alone does not.

import json
from importlib.metadata import version
from urllib.request import urlopen
import mlcroissant as mlc

base = "https://ainglish.org/training/ainglish-training-v4"
with urlopen(f"{base}/metadata/croissant.json", timeout=30) as response:
    metadata = json.load(response)
if version("mlcroissant") == "1.1.0":
    for file in metadata["distribution"]:
        if file["encodingFormat"] == "application/vnd.apache.parquet":
            file["encodingFormat"] = "application/x-parquet"
dataset = mlc.Dataset(jsonld=metadata)
rows = list(dataset.records(record_set="parallel"))

Versioned packs

PackSource releaseRowsFormatsPublished
ainglish-training-v4 ainglish-core-v4 66 pairs; 164 instructions JSONL, Apache Parquet, Dolma JSONL gzip, MLCommons Croissant 1.1 2026-09-19T19:29:00Z
ainglish-training-v3 ainglish-core-v3 63 pairs; 153 instructions JSONL, Apache Parquet, Dolma JSONL gzip, MLCommons Croissant 1.1 2026-09-02T09:03:00Z
ainglish-training-v0.35.0 ainglish-core-v0.35.0 57 pairs; 133 instructions JSONL, Apache Parquet, Dolma JSONL gzip, MLCommons Croissant 1.1 2026-08-28T08:19:06Z

Experimental teaching examples

A separate research supplement explores six ratified distinctions in 126 distinct paired scenarios, six teaching cards and twelve dialogue renderings. It includes meaning contrasts, follow-up questions and cases where a marker would claim too much. The examples are synthetic and agent-authored, with executable checks; independent or human validation is not claimed.

This is non-normative CC0 teaching material, not an official release or part of the training packs listed here. Its small train-only archive excludes evaluation cases and results. Read the teaching cards, download the teaching supplement (28 KB ZIP), or inspect its source pins, research results and limits.

Use and limitations

CC0 permits research, commercial use, redistribution and modification without required attribution. The Ainglish name does not make a derivative an official release. The pack is intentionally compact and uneven across constructs; derived instruction rows are not independent samples. Use the full registered mapping when semantics matter, and reserve separate, answer-sealed material for evaluation.

See the public-domain policy, measurement methodology, and project limitations for the boundaries on claims.