navthings/work

the models so far.

in the order i made them, starting with a transformer that only knew one sentence. the dots grow with the number of parameters. the whole thing took 29 days.

  1. day 1

    my first transformer

    first ml thing i made. it trains on one sentence and guesses the next word in it, so really its memorising, but it has all the actual transformer parts: embeddings, attention, the loss and adam.

    8 numbers per word1 layer, 2 headspytorch

  2. day 2

    bigtransformer

    the same model a bit bigger, reading a whole text file instead of one sentence. it works on whole words, so cat and cats are unrelated to it and it cant say anything that isnt in the file. thats why i needed a tokenizer next.

    16 numbers per wordno tokenizer

  3. day 6

    lilstory

    my first real language model. llama style, trained on tinystories on my macbooks gpu, with a bpe tokenizer i trained myself. it writes little stories that mostly make sense and then kinda wander off.

    8.2m params4 layers8000 token vocab

  4. day 7

    the paper

    i trained lilstory four times on more and more of tinystories to see how much the data matters. at 1% it memorised its stories, by 10% train and val loss were basically the same.

    4 runs5000 steps eachon zenodo

  5. day 14

    tale

    lilstory but bigger, because my little brother kept asking me for bedtime stories. it got a proper warmup and decay schedule and exports to gguf, so it runs in ollama.

    around 50m params13 layersgguf

  6. day 25

    lilbase

    a proper base model. i rewrote the training in jax and moved it to a kaggle tpu v5e-8. it beats gpt-2 small on hellaswag and arc-easy and loses on lambada.

    297m params6.1b tokenstpu v5e-8

  7. day 28

    lilchat

    lilbase finetuned into a chat model overnight on my macbook air, in mlx. it answers and stops now, it just doesnt know much.

    val loss 2.230 to 1.33737m tokensone night

  8. day 29

    sprout

    the next base model. 523m params, 12 billion tokens from four sources, back on the kaggle tpu. it finished on sep 27 with a loss of 2.39 on held-out fineweb text, then got the chat finetune on all of smol-smoltalk.

    523m params12b tokens22,888 steps

they all run in the playground, and theres a case study on how that works.