Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together

Kwon Crash

Published Aug 26, 2026, 1:51 AM UTC

Source: AISource
- Liquid AI just shipped Pipette — an open-source benchmarking suite that tests how foundation models actually perform on edge devices instead of parroting server-class fairy tales from model cards. The premise is almost offensive in its honesty: on-device behavior depends on the full deployment stack — model plus quantization plus runtime plus hardware — not some isolated benchmark number you bragged about in a press release. They launched with 1,000+ configurations across 30+ models, verified by Artificial Analysis. The findings are brutal: two 350M models at the same quantization on the same phone retain 78.4% versus 33.8% of decode throughput. Translation — your spec sheet is a hash manifest nobody verified. Apache 2.0, no waitlist, real devices. This is what happens when someone actually reads the cargo manifest instead of trusting the seal.