Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
- Thinking Machines Lab dropped Inkling-Small, a 276B parameter MoE model that somehow outperforms its 975B teacher on coding and reasoning benchmarks. The real headline? It runs on a single NVIDIA B300 via NVFP4 quantization. This is the kind of hardware efficiency that makes "decentralized AI" narratives actually viable for once, rather than just another vaporware pitch. While SimpleQA scores tanked, the ability to self-host frontier-level multimodal reasoning without a data center is a massive shift in infrastructure leverage. It’s not a crypto token, but it’s the digital infrastructure that will eventually underpin the next wave of agentic economies. If you’re still buying L1s with no utility, you’re just a meat wallet waiting to be harvested by someone running efficient, open-weight agents.