a base model just continues text. if you ask it a question itll probably write 5 more questions instead of answering. to make it a chat model you finetune it on a bunch of conversations so it learns the pattern: you say something, it answers, it stops.
i did this overnight on my macbook air with mlx. used smol-smoltalk as the data, and it only gets trained on the assistant replies. it did 1,136 steps of 32k tokens, about 37m tokens total, which is only like 12% of the dataset but thats what my laptop could do in a night.
loss on conversations it hadnt seen went from 2.230 (plain lilbase) to 1.337.
what works
if you say "hi" it says "Hello! How can I help you today?" and stops. plain lilbase just repeats the chat formatting forever when you do that. it does lists and code blocks, and it wrote a correct one line function to reverse a string (then explained it completely wrong lol).
what doesnt
it does not know things. it told me the capital of australia is melbourne, then perth on another try. 17 + 25 was 48, and then "so, the answer is 17". i asked for a haiku about a cat that hates mondays and got 20 lines ending with "a reminder that we are all brothers" over and over.
thats not really a finetuning problem, 297m params just doesnt know much. turning the temperature down from 0.7 to 0.4 helps a bit tho, it gave a normal reply to "hi" 12 out of 12 times instead of 8 out of 12.
ollama run navthings/lilchat
or just try it in your browser, no install.
next
the real fix is a bigger base model, which is sprout. 523m params, 12 billion tokens from fineweb-edu, dclm, cosmopedia and project gutenberg, back on the kaggle tpu. it got the same finetune, heres how that went.