Running and optimizing open-source LLMs on a single Apple Silicon laptop — and measuring what actually matters: speed, memory, and real task quality (coding & math). Everything here is reproducible; no cloud GPU required. New to the terms? Every i explains itself, and there's a full Glossary.
Compare famous & coding-specialized models on coding, math, speed, and memory. Which local model is the best assistant?
Quantization, speculative decoding, KV-cache compression, and batching — how each makes a model faster or smaller.
See the actual questions and what each model answered — good and bad — so the scores make sense.
Plain-English definitions of every term: tokens/sec, quantization, perplexity, pass@1, and more.