work/pretraining · sep 2026

sprout

the next base model after lilbase. same recipe, about 1.8x the parameters and 2x the data, on a mix of four datasets instead of one. it finished training on a tpu v5e-8 on sep 27, then got a chat finetune.

params
523m
read
12b tokens from 4 sources
trained on
kaggle tpu v5e-8
held-out loss
2.39 on fineweb

lilbase and sprout, side by side

lilbasesprout
params
297m
523m
tokens read
6.1b
12b
layers
24
28
width
1024
1280
query heads
16
20
see the numbers
lilbasesprout
params297m523m
tokens6.1b12b
layers2428
width10241280
query heads1620
kv heads44
datafineweb-edu4 source mix

why

lilchat can hold a conversation now, but it still gets basic facts wrong. 297m params trained on 6.1 billion tokens just doesnt know much.

the fix is a bigger base model that has read more. thats sprout.

the data

lilbase only read fineweb-edu. sprout reads 12 billion tokens from four sources, mixed by share. books get cut into 16,000 character chunks first, about 4k tokens each, so one novel cant fill a whole run of batches.

fineweb-edu 55%
web pages filtered for educational value, the sample-100BT cut
dclm-baseline 25%
filtered common crawl, broader general web text
cosmopedia-v2 10%
synthetic textbooks and stories
project gutenberg 10%
public domain books, english only

the model

same recipe as lilbase: llama style, with grouped query attention, rope, rmsnorm and swiglu, and the llama tokenizer. just deeper and wider.

params
523m
layers
28
width
1280
query heads
20
kv heads
4
context
1024

training

adamw keeps two extra numbers for every parameter, which adds up fast at 523m. so the optimizer state is sharded across the 8 chips. gradients are reduce-scattered, so each chip only receives the slice it owns, runs adamw on that slice, then the updated slices are all-gathered back into full weights.

8 chips, each with the full weightsand 1/8 of the optimizer state

the learning rate follows a wsd schedule: 1500 steps of linear warmup to 3e-4, hold it flat, then decay linearly to 10% of that over the last 15% of training. lilbase used a cosine schedule, which bakes the total length in from step one. with wsd the total token count can still change between sessions as long as the decay hasnt started, so the run doesnt have to commit up front.

peak 10% start of training decay, last 15%
sprout, wsdlilbase, cosine

each as a share of its own peak, 3e-4 for sprout and 4e-4 for lilbase, over the length of each run

each session gets about 8.4 hours before kaggle ends it at 9. the next one finds the newest checkpoint on its own and carries on with the data position, optimizer state, schedule and loss history exactly where they were.

how it did

it finished on sep 27, after 22,888 steps and all 12 billion tokens. by the end it was doing about 130,000 tokens a second, a bit over 4 seconds a step across the 8 chips.

on text it never trained on, the loss came out at 2.39 on fineweb (perplexity 10.9) and 2.72 on wikitext (perplexity 15.2). the mix it actually read landed a little off the plan: fineweb-edu 55%, dclm 27.2%, cosmopedia 8.7% and gutenberg 9.1%.

steps
22,888
tokens
12.0b
fineweb loss
2.387
fineweb perplexity
10.9
wikitext loss
2.723
wikitext perplexity
15.2
tokens a second
130,034
final learning rate
3e-5

these are three things it wrote at the very end. the grammar holds up, and it still makes stuff up with total confidence. einstein did get the nobel in 1921, but he wasnt the most decorated scientist in the world in 1905.

The water cycle begins when rainwater washes over the land and picks up bits of chemicals that are dissolved in water, such as salt, and other chemicals that are released into the atmosphere. Rainwater that falls to the earth's surface is mostly water, but some can also contain dissolved chemicals, called organic chemicals, that can remain in the soil and be transported to the surface

In 1905, Albert Einstein was the most decorated scientist in the world. An internationally-known physicist of international renown, he was awarded the Nobel Prize in Physics in 1921. Einstein’s research on general relativity, the theory of special and general relativity, as well as quantum mechanics, has had a tremendous impact on our understanding of the universe.

She opened the letter and read: “Don’t forget to read the letter. When you send a letter, it takes several days to get through the mail. If you send your letter in the mail, you have to wait a day or two. You must be careful and take every precaution. This article will show you how to send a letter fast, and how to make sure you get the letter in

then it got the chat finetune.

chat finetune

same recipe as lilchat: smol-smoltalk conversations in the llama-2 [INST] format, with the loss only on the assistant replies. lilchat ran on my macbook and saw 12% of the data. sprout ran on the tpu, so it got 2 full epochs, 616m tokens in about 90 minutes, and kept whichever checkpoint did best on held-out chats.

steps
4,702
tokens a step
131,072
learning rate
1e-4
tokens a second
121,000
held-out loss before
1.611
held-out loss after
0.955

the best checkpoint was the last one, so the second epoch still helped on chats it had never seen. it gets canberra right now, writes working python and explains it properly. maths and poems are still bad.

its on hugging face and ollama, and in the playground as a 345mb q4_k_m. theres a post with the loss curves.

vs gpt-2

i put sprout through the same test script as openai's gpt-2, which came out in 2019 in four sizes, from 124m up to 1.5b params. lilbase is in there too. each test is a few thousand questions none of the models trained on, every model gets the exact same ones, and it all runs on my macbook.

on arc-easy, which is grade school science questions, sprout gets 50.3% and lilbase gets 48.1%. both beat gpt-2 medium at 44.7%, and lilbase is smaller than it. on lambada, where you guess the last word of a passage from a novel, gpt-2 wins: medium gets 44.3%, sprout 39.2% and lilbase 28.5%.

arc-easy grade school science questions

35%40%45%50%55%124m355m774m1.5bmodel size (params)smallmediumsprout 50.3lilbase 48.1

lambada guessing the last word of a passage

25%30%35%40%45%50%124m355m774m1.5bmodel size (params)smallmediumsprout 39.2lilbase 28.5
sprout (chat version), 523mlilbase, 297mopenai gpt-2, 2 sizes so far

higher is better. a dot above the grey line beats a gpt-2 of the same size. same test script for every model, on my macbook

see the numbers
modelparamsarc-easylambada
gpt-2 small124m40.6%35.0%
lilbase297m48.1%28.5%
gpt-2 medium355m44.7%44.3%
sprout (chat)523m50.3%39.2%

this is the chat version of sprout, the base weights are still on kaggle, and chat finetuning usually costs a few points on tests like these. gpt-2 large and xl are still running, and hellaswag is getting rerun because of a bug in how it was scored (the story and the answer got glued together with no space). ill update this when theyre done.

goals

one: beat gpt-2 xl, which is 1.5b params, or get close to it.

two: finetune it into a competitive coding model at around 500m.

the first benchmark results are above. gpt-2 large and xl are still running.

next case study lilbase

the 297m base model that sprout is built on.