Together AI brings Thinking Machines Lab's Inkling model to its platform on day 0
Together AI is now serving Inkling, a new multimodal mixture-of-experts model from Thinking Machines Lab, on its inference platform starting the day of the model's release. The launch gives developers immediate API access to Inkling without waiting for a separate hosting rollout.
What's new
Inkling is a large multimodal MoE model with 975 billion total parameters and roughly 40 billion active parameters per token, paired with a 1 million token context window. According to Together AI's announcement, the model accepts text, image, and audio inputs and produces text outputs through a unified decoder architecture. It also supports controllable inference effort, letting developers dial reasoning depth up or down depending on the task.
Together AI frames the release as a "day 0" partnership: rather than the usual lag between a lab's own release and third-party hosting, Inkling is live on Together's infrastructure at launch, positioning the platform as a fast on-ramp for developers who want to try the model without going through Thinking Machines Lab directly.
Context
Thinking Machines Lab has been one of the most closely watched new entrants in frontier AI, and Inkling is a significant public signal of the technical direction the company is taking — a large-scale, multimodal MoE architecture rather than a narrower, single-modality model. Together AI, meanwhile, has built its business on being an early, low-friction host for new and open models, a role it has played for labs ranging from Meta to DeepSeek to smaller research outfits. Landing a same-day hosting deal with a lab as high-profile as Thinking Machines Lab is a notable win for that strategy.
The model's size and shape — a very large total parameter count with a much smaller active count — follows the industry-wide shift toward sparse MoE architectures that keep inference costs manageable while scaling total capacity. The 1 million token context window and multimodal input handling (text, image, and audio) also put Inkling in direct competition with the context and modality capabilities frontier labs like Google and OpenAI have been racing to expand.
Why it matters
For developers, day-0 availability on a major inference platform lowers the barrier to evaluating a brand-new model — no waiting on capacity, regional availability, or a separate vendor relationship with the lab itself. For Together AI, it reinforces a distribution advantage: being the fastest place to try a hot new model is a meaningful draw in a market where inference providers otherwise compete heavily on price and latency.
More broadly, Inkling's launch adds to a growing list of well-funded challengers building frontier-scale multimodal models outside the four major chat labs (OpenAI, Anthropic, Google, and xAI). Whether Inkling's actual performance holds up against those incumbents will depend on independent benchmarking, but the technical specs — large MoE scale, million-token context, tri-modal input, controllable reasoning effort — signal that Thinking Machines Lab is building for the same class of workloads the frontier labs are targeting, not a niche use case.
Corroborating sources
- Together
https://www.together.ai/blog/together-ai-brings-thinking-machines-labs-new-model-inkling-on-day-0
“Inkling accepts text, image, and audio inputs and produces text outputs through a unified decoder architecture.”