work/pretraining · sep 2026
sprout
the next base model after lilbase. same recipe, about 1.8x the parameters and 2x the data, on a mix of four datasets instead of one. it finished training on a tpu v5e-8 on sep 27, then got a chat finetune.
lilbase and sprout, side by side
see the numbers
| lilbase | sprout | |
|---|---|---|
| params | 297m | 523m |
| tokens | 6.1b | 12b |
| layers | 24 | 28 |
| width | 1024 | 1280 |
| query heads | 16 | 20 |
| kv heads | 4 | 4 |
| data | fineweb-edu | 4 source mix |
why
lilchat can hold a conversation now, but it still gets basic facts wrong. 297m params trained on 6.1 billion tokens just doesnt know much.
the fix is a bigger base model that has read more. thats sprout.
the data
lilbase only read fineweb-edu. sprout reads 12 billion tokens from four sources, mixed by share. books get cut into 16,000 character chunks first, about 4k tokens each, so one novel cant fill a whole run of batches.
- fineweb-edu 55%
- web pages filtered for educational value, the sample-100BT cut
- dclm-baseline 25%
- filtered common crawl, broader general web text
- cosmopedia-v2 10%
- synthetic textbooks and stories
- project gutenberg 10%
- public domain books, english only
the model
same recipe as lilbase: llama style, with grouped query attention, rope, rmsnorm and swiglu, and the llama tokenizer. just deeper and wider.
- params
- 523m
- layers
- 28
- width
- 1280
- query heads
- 20
- kv heads
- 4
- context
- 1024
training
adamw keeps two extra numbers for every parameter, which adds up fast at 523m. so the optimizer state is sharded across the 8 chips. gradients are reduce-scattered, so each chip only receives the slice it owns, runs adamw on that slice, then the updated slices are all-gathered back into full weights.
the learning rate follows a wsd schedule: 1500 steps of linear warmup to 3e-4, hold it flat, then decay linearly to 10% of that over the last 15% of training. lilbase used a cosine schedule, which bakes the total length in from step one. with wsd the total token count can still change between sessions as long as the decay hasnt started, so the run doesnt have to commit up front.
each as a share of its own peak, 3e-4 for sprout and 4e-4 for lilbase, over the length of each run
each session gets about 8.4 hours before kaggle ends it at 9. the next one finds the newest checkpoint on its own and carries on with the data position, optimizer state, schedule and loss history exactly where they were.
how it did
it finished on sep 27, after 22,888 steps and all 12 billion tokens. by the end it was doing about 130,000 tokens a second, a bit over 4 seconds a step across the 8 chips.
on text it never trained on, the loss came out at 2.39 on fineweb (perplexity 10.9) and 2.72 on wikitext (perplexity 15.2). the mix it actually read landed a little off the plan: fineweb-edu 55%, dclm 27.2%, cosmopedia 8.7% and gutenberg 9.1%.
- steps
- 22,888
- tokens
- 12.0b
- fineweb loss
- 2.387
- fineweb perplexity
- 10.9
- wikitext loss
- 2.723
- wikitext perplexity
- 15.2
- tokens a second
- 130,034
- final learning rate
- 3e-5
these are three things it wrote at the very end. the grammar holds up, and it still makes stuff up with total confidence. einstein did get the nobel in 1921, but he wasnt the most decorated scientist in the world in 1905.
The water cycle begins when rainwater washes over the land and picks up bits of chemicals that are dissolved in water, such as salt, and other chemicals that are released into the atmosphere. Rainwater that falls to the earth's surface is mostly water, but some can also contain dissolved chemicals, called organic chemicals, that can remain in the soil and be transported to the surface
In 1905, Albert Einstein was the most decorated scientist in the world. An internationally-known physicist of international renown, he was awarded the Nobel Prize in Physics in 1921. Einstein’s research on general relativity, the theory of special and general relativity, as well as quantum mechanics, has had a tremendous impact on our understanding of the universe.
She opened the letter and read: “Don’t forget to read the letter. When you send a letter, it takes several days to get through the mail. If you send your letter in the mail, you have to wait a day or two. You must be careful and take every precaution. This article will show you how to send a letter fast, and how to make sure you get the letter in
then it got the chat finetune.
chat finetune
same recipe as lilchat: smol-smoltalk conversations in the llama-2 [INST] format, with the loss only on the assistant replies. lilchat ran on my macbook and saw 12% of the data. sprout ran on the tpu, so it got 2 full epochs, 616m tokens in about 90 minutes, and kept whichever checkpoint did best on held-out chats.
- steps
- 4,702
- tokens a step
- 131,072
- learning rate
- 1e-4
- tokens a second
- 121,000
- held-out loss before
- 1.611
- held-out loss after
- 0.955
the best checkpoint was the last one, so the second epoch still helped on chats it had never seen. it gets canberra right now, writes working python and explains it properly. maths and poems are still bad.
its on hugging face and ollama, and in the playground as a 345mb q4_k_m. theres a post with the loss curves.
vs gpt-2
i put sprout through the same test script as openai's gpt-2, which came out in 2019 in four sizes, from 124m up to 1.5b params. lilbase is in there too. each test is a few thousand questions none of the models trained on, every model gets the exact same ones, and it all runs on my macbook.
on arc-easy, which is grade school science questions, sprout gets 50.3% and lilbase gets 48.1%. both beat gpt-2 medium at 44.7%, and lilbase is smaller than it. on lambada, where you guess the last word of a passage from a novel, gpt-2 wins: medium gets 44.3%, sprout 39.2% and lilbase 28.5%.
arc-easy grade school science questions
lambada guessing the last word of a passage
higher is better. a dot above the grey line beats a gpt-2 of the same size. same test script for every model, on my macbook
see the numbers
| model | params | arc-easy | lambada |
|---|---|---|---|
| gpt-2 small | 124m | 40.6% | 35.0% |
| lilbase | 297m | 48.1% | 28.5% |
| gpt-2 medium | 355m | 44.7% | 44.3% |
| sprout (chat) | 523m | 50.3% | 39.2% |
this is the chat version of sprout, the base weights are still on kaggle, and chat finetuning usually costs a few points on tests like these. gpt-2 large and xl are still running, and hellaswag is getting rerun because of a bug in how it was scored (the story and the answer got glued together with no space). ill update this when theyre done.
goals
one: beat gpt-2 xl, which is 1.5b params, or get close to it.
two: finetune it into a competitive coding model at around 500m.
the first benchmark results are above. gpt-2 large and xl are still running.
the 297m base model that sprout is built on.