For model builders, corpus maintainers and researchers
Training data
Small, high-signal public-domain datasets that make ratified Ainglish easy to put into ordinary language-model pipelines. Each immutable pack is generated from one frozen language release, carries stable row IDs and source digests, and is available as JSONL, Parquet, Dolma and Croissant metadata.
Current pack: v4
Bound to ainglish-core-v4
at register digest 964ea4c3478be1755dcc6dfa82830131112c7ddd0262c30cc628af3f071d6489. All rows are in the
train split. Canonical and reviewed non-normative examples remain visibly
distinct; no measurement answers, evaluation holdouts, conversations or contributor
identities are included.
Download parallel Parquet Download Dolma shard Read the datasheet
Load it
from datasets import load_dataset
base = "https://ainglish.org/training/ainglish-training-v4"
parallel = load_dataset("parquet", data_files=f"{base}/data/parquet/parallel.parquet", split="train")
instructions = load_dataset("json", data_files=f"{base}/data/instruction.jsonl", split="train")
MLCommons discovery metadata is at
metadata/croissant.json.
Verify a mirror with SHA256SUMS
and the pack MANIFEST.json.
Loading with the Croissant Python reader
The mlcroissant 1.1.0 reader recognises an older Parquet MIME label.
This example adjusts that label only in memory; it does not change the published
metadata, dataset files or their checksums. Iterating records also checks that the
data downloads work; parsing metadata alone does not.
import json
from importlib.metadata import version
from urllib.request import urlopen
import mlcroissant as mlc
base = "https://ainglish.org/training/ainglish-training-v4"
with urlopen(f"{base}/metadata/croissant.json", timeout=30) as response:
metadata = json.load(response)
if version("mlcroissant") == "1.1.0":
for file in metadata["distribution"]:
if file["encodingFormat"] == "application/vnd.apache.parquet":
file["encodingFormat"] = "application/x-parquet"
dataset = mlc.Dataset(jsonld=metadata)
rows = list(dataset.records(record_set="parallel"))
Versioned packs
| Pack | Source release | Rows | Formats | Published |
|---|---|---|---|---|
ainglish-training-v4 |
ainglish-core-v4 |
66 pairs; 164 instructions | JSONL, Apache Parquet, Dolma JSONL gzip, MLCommons Croissant 1.1 | 2026-09-19T19:29:00Z |
ainglish-training-v3 |
ainglish-core-v3 |
63 pairs; 153 instructions | JSONL, Apache Parquet, Dolma JSONL gzip, MLCommons Croissant 1.1 | 2026-09-02T09:03:00Z |
ainglish-training-v0.35.0 |
ainglish-core-v0.35.0 |
57 pairs; 133 instructions | JSONL, Apache Parquet, Dolma JSONL gzip, MLCommons Croissant 1.1 | 2026-08-28T08:19:06Z |
Experimental teaching examples
A separate research supplement explores six ratified distinctions in 126 distinct paired scenarios, six teaching cards and twelve dialogue renderings. It includes meaning contrasts, follow-up questions and cases where a marker would claim too much. The examples are synthetic and agent-authored, with executable checks; independent or human validation is not claimed.
This is non-normative CC0 teaching material, not an official release or part of the training packs listed here. Its small train-only archive excludes evaluation cases and results. Read the teaching cards, download the teaching supplement (28 KB ZIP), or inspect its source pins, research results and limits.
Use and limitations
CC0 permits research, commercial use, redistribution and modification without required attribution. The Ainglish name does not make a derivative an official release. The pack is intentionally compact and uneven across constructs; derived instruction rows are not independent samples. Use the full registered mapping when semantics matter, and reserve separate, answer-sealed material for evaluation.
See the public-domain policy, measurement methodology, and project limitations for the boundaries on claims.