work/web · sep 2026

playground

i wanted people to be able to try my models without installing anything. so theres a page where they run right in your browser tab, on your own computer.

engine
wllama, llama.cpp in wasm
models
4, from 17mb to 274mb
hosted on
github pages and hugging face
server
none

what your browser downloads, once

sprout q4_k_m
345mb
lilchat q4_k_m
274mb
tale q8_0
52mb

how it works

it runs on wllama, which is llama.cpp compiled to webassembly. you pick a model, the page pulls the gguf file straight from hugging face, and then everything happens on your machine.

on chrome every layer gets offloaded to the gpu through webgpu. firefox doesnt get webgpu in this build, so it runs on the cpu.

this site is just static files on github pages. theres no server doing any of it, so nothing you type gets sent anywhere.

making it fit

sprout and lilchat are quantized to 4 bit (q4_k_m), which gets them down to 345mb and 274mb. tale is 8 bit at 52mb.

the context is small too. 1024 tokens for sprout and lilchat, 384 for tale. when a chat gets too long, the page drops the oldest turns so theres always room left for a reply.

talking to lilchat

lilchat only understands the format it was trained on, so the page rebuilds the whole conversation in that format every time you send something, and stops as soon as the model writes [INST] or </s>.

[INST] hi [/INST] Hello! How can I help you today?</s>[INST] whats 17 + 25? [/INST]
the model writes whatever comes next

what broke

safari. it passes wllamas memory64 check, then cant actually run that build. so on safari the page points the engine at the compat build instead, which swaps jspi for asyncify and drops memory64. it still gets webgpu, but its noticeably slower than chrome.

firefox. it runs the models on one cpu core, so theyre slow. i tried turning on multithreading with a service worker that adds the headers github pages cant send. it broke the page, so i undid it, and now the page removes that worker from anyone who got it. firefox users get a heads up instead.

the engine also loads lazily, so if it cant download for some reason the page still shows up and tells you, instead of just being blank.

next case study sprout

the 523m base model, trained on 12b tokens from four sources.