navthings/work
the models so far.
in the order i made them, starting with a transformer that only knew one sentence. the dots grow with the number of parameters. the whole thing took 29 days.
-
day 1
my first transformer
first ml thing i made. it trains on one sentence and guesses the next word in it, so really its memorising, but it has all the actual transformer parts: embeddings, attention, the loss and adam.
8 numbers per word1 layer, 2 headspytorch
-
day 2
bigtransformer
the same model a bit bigger, reading a whole text file instead of one sentence. it works on whole words, so cat and cats are unrelated to it and it cant say anything that isnt in the file. thats why i needed a tokenizer next.
16 numbers per wordno tokenizer
-
day 6
lilstory
my first real language model. llama style, trained on tinystories on my macbooks gpu, with a bpe tokenizer i trained myself. it writes little stories that mostly make sense and then kinda wander off.
8.2m params4 layers8000 token vocab
-
day 7
the paper
i trained lilstory four times on more and more of tinystories to see how much the data matters. at 1% it memorised its stories, by 10% train and val loss were basically the same.
4 runs5000 steps eachon zenodo
-
day 14
-
day 25
lilbase
a proper base model. i rewrote the training in jax and moved it to a kaggle tpu v5e-8. it beats gpt-2 small on hellaswag and arc-easy and loses on lambada.
297m params6.1b tokenstpu v5e-8
-
day 28
lilchat
lilbase finetuned into a chat model overnight on my macbook air, in mlx. it answers and stops now, it just doesnt know much.
val loss 2.230 to 1.33737m tokensone night
-
day 29
sprout
the next base model. 523m params, 12 billion tokens from four sources, back on the kaggle tpu. it finished on sep 27 with a loss of 2.39 on held-out fineweb text, then got the chat finetune on all of smol-smoltalk.
523m params12b tokens22,888 steps
they all run in the playground, and theres a case study on how that works.