how this works

talk to my models.

they run right here in your browser. pick one, it downloads once, then runs on your own computer. nothing you type gets sent anywhere.

further down: lilstory's tokenizer and a digit net with no libraries

runs on wllama, which is llama.cpp in your browser. the models are on hugging face. theyre small, so they make stuff up a lot.

how a model reads.

before a model reads anything, the text gets cut into pieces and each piece becomes a number. this is the tokenizer i trained for lilstory, 8000 pieces it learned from about 2 million kids stories.

words it saw a lot are one piece, anything else gets chopped up. duck is one token, tpu is t, p, u. my all lowercase headline costs 3 more tokens than the capitalised one, because the stories it learned from use proper capitals.

a neural net with no libraries.

neural networks in plain python. no numpy, no pytorch, no mlx, just math, random, json and os, which all come with python. every weight is a number in a python list, and backprop is written out by hand as loops. 784 pixels in, 32 hidden neurons, 10 outputs, about 25k weights, and it gets 90.5% of the mnist test set right.

draw a digit, or show it a real one from the test set, which it never trained on. its running the exact weights it learned.

draw a digit

its guess?
what it sees, 28 x 28
  1. 0
  2. 1
  3. 2
  4. 3
  5. 4
  6. 5
  7. 6
  8. 7
  9. 8
  10. 9