MicroLLM Lab – Try 7 tiny LLM's in the browser
MicroLLM Lab is a browser experiment for running seven tiny LLMs locally, described as Q4-quantized and executed through a WebGPU engine. The page initializes a model catalog and shows an active model, indicating which model is loaded on the device and that it is saved in a browser IndexedDB cache. Controls include loading all models and unloading all models, plus a generate-max-new-tokens setting and a next-action selector. Benchmarking is central to the tool. The interface reports "Benchmarking in progress," prepares tests, and states that evaluation uses objective checks based on regex or exact tokens rather than writing quality. It offers running a suite on the active model, running a suite on every loaded model, and a sustained speed test of 256 tokens. Two speed readouts are defined: highest peak speed for the fastest single test, and highest sustained speed for continuous 256-token decode. The page also notes that a 135M model is allowed to fail, describing that failure as the measurement, and that time estimates use the last tokens-per-second figure when one is available.
Developers can benchmark tiny Q4 LLMs locally in the browser via WebGPU, with objective token checks instead of subjective writing-quality judgments.