Interpretability research gets a cleaner small-vocabulary synthetic corpus
An independent researcher released Small World 345.6k: 8,873 words, each appearing at least 16 times — but quality is self-reported only.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
An independent researcher has published Small World 345.6k, a synthetic dataset on LessWrong with a vocabulary of just 8,873 words, every one of which appears at least 16 times.
By comparison, TinyStories has 49,187 unique words, with 15 to 24 percent of them occurring fewer than 2 times. The new dataset uses vocabulary capping and inverse-frequency-weighted resampling to fix this, and the author says text is dictionary-validated to be error-free.
Generation ran on a single 16GB RTX 5060 Ti using a quantized unsloth gemma-4-26B model, and the pipeline is locally reproducible. Note this is a first-party result: the dataset has no independent training validation yet, and its total size is smaller than existing comparable corpora.