Thinking Machines Lab releases Inkling, a 975B-parameter open-weight model
Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, released Inkling on July 15, 2026 — its first open-weight model and a direct challenge to the "one-size-fits-all" approach dominating frontier AI development.
What's new
Inkling is a mixture-of-experts model with 975 billion total parameters and roughly 41 billion active per task, trained on 45 trillion tokens of text, image, audio, and video data. Per the company's own model page, it is "an efficient open model for text, images, and audio, available to fine-tune," with a 1 million token context window (64K or 256K when used through Tinker, the company's training platform).
The model is natively multimodal on input — "Natively multimodal on input, and at the open-source frontier for speech" — and lets users dial thinking time up or down to trade speed for performance. Thinking Machines also emphasizes calibration: the model "makes predictions with well-calibrated confidence" rather than guessing when uncertain.
On benchmarks, the company says Inkling uses roughly a third as many tokens as Nvidia's Nemotron 3 Ultra to hit equivalent coding performance. In an early enterprise deployment, Bridgewater Associates reportedly hit 84.7% on financial reasoning tests — beating top proprietary models — while running at roughly one-fourteenth the cost.
Inkling is open-weight and downloadable now, with Tinker available for teams that want to fine-tune it for specific domains. Together AI added same-day platform support, making it available via API without local deployment.
Context
Thinking Machines Lab has published research on model customization and efficient training (On-Policy Distillation, LoRA Without Regret) since its founding, but Inkling marks its first shipped foundation model. The launch lands amid a broader wave of open-weight releases from well-funded labs — DeepSeek, Google's DiffusionGemma, and others — competing on cost-efficiency rather than raw scale alone.
Why it matters
Thinking Machines' own positioning is notably candid: "Inkling is not the strongest overall model available today, open or closed." Rather than chasing the top of leaderboards, the company is betting that a well-rounded, customizable, open-weight model — paired with Tinker's fine-tuning tools — is more valuable to enterprises than a slightly higher score on a static benchmark. The Bridgewater efficiency numbers, if they hold up under broader scrutiny, would support that thesis: cheaper inference at near-parity performance matters more to most production deployments than marginal capability gains. It's also a signal that the fight for enterprise AI budgets is shifting toward total cost of ownership and customizability, not just raw model quality.
Corroborating sources
- Techcrunch
https://techcrunch.com/2026/07/15/thinking-machines-amps-up-its-bet-against-one-size-fits-all-ai-with-its-first-open-model-inkling/
“Inkling uses a third as many tokens as Nvidia's Nemotron 3 Ultra”
- Theregister
https://www.theregister.com/ai-and-ml/2026/07/16/former-openai-cto-does-what-altman-wont-releases-a-frontier-ai-model-thats-actually-open/5272177
- Thinkingmachines
https://thinkingmachines.ai/inkling