work/research · sep 2026

the paper

i wanted to see how much the amount of data actually matters if you keep everything else the same. so i trained the same model four times on more and more of tinystories, and wrote it up.

model
lilstory, 8.2m
runs
4, 5000 steps each
data
22k to 330k stories
published on
zenodo

train and val loss at the end of each run, lower is better

val loss, on stories it never sawtrain loss
see the numbers
datastoriestrain lossval loss
1%22k1.3292.402
5%110k1.7011.932
10%220k1.8261.816
15%330k1.7371.837

the setup

one model, lilstory, trained four times on 1%, 5%, 10% and 15% of tinystories. same size, same settings, 5000 steps each. the only thing that changes is how much it gets to read.

lower loss is better, and val loss is the one that matters, since its measured on stories the model never saw.

what happened

the 1% one had the best training loss and the worst val loss. it basically just memorised its 22k stories lol. the more data i gave it the smaller that gap got, and at 10% theyre pretty much the same.

the gap, val loss minus train loss

1%1.073
5%0.231
10%-0.010, basically none
15%0.100
1% run: train loss keeps dropping while validation loss flattens out and creeps back up
1% of the data. train loss (blue) keeps going down, val loss (orange) gives up and creeps back up
10% run: train and validation loss stay together the whole way
10% of the data. the two lines stay together the whole way

the weird one

15% was slightly worse than 10%, which i didnt expect. i think its because every model only got 5000 steps, so the 15% one probably never even got through all its data.

the difference is tiny too, 0.021, so it could just be noise.

read it

its on zenodo if you wanna read the whole thing: the effect of corpus size on the performance of a llm.

next case study playground

the page where all my models run in your browser.