writing/

i wrote a paper on how much data matters

i wanted to see how much the amount of data actually matters if you keep everything else the same. so i trained lilstory 4 times on 1%, 5%, 10% and 15% of tinystories. same model, same settings, 5000 steps each.

datastoriestrain lossval loss
1%22k1.3292.402
5%110k1.7011.932
10%220k1.8261.816
15%330k1.7371.837

lower is better, and val loss is the one that matters since its on stories the model never saw.

the 1% one had the best training loss but the worst val loss, it basically just memorised its 22k stories lol. the more data i gave it the smaller that gap got, and at 10% theyre pretty much the same.

1% run: train loss keeps dropping while validation loss flattens out and creeps back up
1% of the data. train loss (blue) keeps going down, val loss (orange) gives up and creeps back up
10% run: train and validation loss stay together the whole way
10% of the data. the two lines stay together the whole way

15% was slightly worse than 10% which i didnt expect. i think its because every model only got 5000 steps, so the 15% one probably never even got through all its data. the difference is tiny too so it could just be noise.

its on zenodo if you wanna read the whole thing: the effect of corpus size on the performance of a llm

theres a case study for thisthe longer version, with interactive charts read it