everything before this was trained on my laptop, but i wanted to make a proper base model, like one thats actually read a decent chunk of the internet. my macbook air was never gonna do that. kaggle gives you free time on a tpu v5e-8 (8 of googles ai chips) so i rewrote the whole thing in jax and moved it there.
its 297m params, llama style, 24 layers, 1024 wide, 16 query heads and 4 kv heads. it read 6.1 billion tokens of fineweb-edu, which is about the right amount for its size (the chinchilla thing, around 20 tokens per param). each step is 524,288 tokens split over the 8 chips, and each kaggle session gets about 8.4 hours before it has to save and stop.

| benchmark | lilbase | gpt-2 124m |
|---|---|---|
| hellaswag | 41.0% | 31.1% |
| arc-easy | 50.5% | 39.5% |
| lambada | 28.7% | 32.6% |
it beats gpt-2 small on hellaswag and arc-easy and loses on lambada. to be fair gpt-2 is from 2019, trained on different data, and lilbase is more than twice its size, so its not exactly a fair fight lol
stuff that went wrong: kaggle sometimes gives you a broken tpu with only 1 chip working instead of 8, so now the script just refuses to run unless it sees all 8. also doubling the batch size ran out of memory by like 400mb.
ollama run navthings/lilbase "The water cycle begins when"
its a base model so it just continues whatever you give it instead of answering. it gets grammar right and makes up most facts.