work/finetuning · sep 2026
lilchat
a base model just continues text. ask it a question and itll probably write 5 more questions instead of answering. lilchat is lilbase after supervised finetuning on conversations, done in mlx in one night.
loss on conversations it hadnt seen, lower is better
the idea
to make a chat model you finetune it on a bunch of conversations, so it learns the pattern: you say something, it answers, it stops.
every conversation gets flattened into one line of text with some markers in it. lilchat learned this exact format, so the playground builds prompts the same way.
training
i did it overnight on my macbook air with mlx. the data is smol-smoltalk, and it only gets trained on the assistant replies, so it learns to answer rather than to write your half of the chat.
1,136 steps of 32k tokens, about 37m tokens total. thats only like 12% of the dataset, but its what my laptop could do in a night.
what works
if you say "hi" it says "Hello! How can I help you today?" and stops. plain lilbase just repeats the chat formatting forever when you do that.
it does lists and code blocks, and it wrote a correct one line function to reverse a string. then explained it completely wrong.
what doesnt
it does not know things. it told me the capital of australia is melbourne, then perth on another try. 17 + 25 was 48, and then "so, the answer is 17".
i asked for a haiku about a cat that hates mondays and got 20 lines ending with "a reminder that we are all brothers" over and over.
thats not really a finetuning problem. 297m params just doesnt know much.
one knob
turning the temperature down helps a bit. out of 12 tries saying "hi", heres how often it gave a normal reply.
going all the way to zero doesnt work either. greedy decoding just loops, so it needs some temperature. it ships at 0.4, with top-k 40 and a repeat penalty of 1.1, and the playground uses the same.
shipping it
its on ollama in three sizes. the 8 bit one is the default because its the same quality as the full 16 bit weights at a bit over half the size.
next
the real fix is a bigger base model, which is sprout. 523m params, 12 billion tokens from fineweb-edu, dclm, cosmopedia and project gutenberg, back on the kaggle tpu. it finished on sep 27, and next it gets the same finetune.
ollama run navthings/lilchat
the same model trained on four amounts of data, written up as a paper.