Inkling isn't interesting because it tops a benchmark — it doesn't.
It's interesting because of who released it, how (open weights), and what that says about where the frontier is heading.
Here's the framing a tweet won't give you. The announcement.
Inkling is a trillion-scale multimodal Mixture-of-Experts — 975B parameters, 41B firing per token, 1M context, trained on 45T tokens of text, images, audio and video.
There's also Inkling-Small (12B active).
But the launch post says the quiet part out loud: "Inkling is not the strongest overall model available today, open or closed."
That's deliberate.
Their bet is that most real-world value comes from customization, not raw capability — a model you can fine-tune to your workflow beats a smarter one you can only rent through an API.
"Different users need models that can adapt to very different workflows, not just excel on benchmarks." Open weights are the product, not a giveaway.
Mira Murati was OpenAI's CTO — she ran the org that shipped ChatGPT.
She left in 2024 and founded Thinking Machines Lab in Feb 2025, taking a cohort of the people who built the closed frontier with her: John Schulman (a father of RLHF), Barret Zoph, Lilian Weng, Andrew Tulloch, Luke Metz.
The lab raised one of the largest seeds ever and is now valued like an incumbent — before shipping a flagship model.
~$2B seed at a $12B valuation; reportedly seeking $50B by late 2025 — an incumbent-sized bet on a 30-person team, placed before a public model existed.
The lab world splits into two camps. Inkling matters because it comes from the closed camp's alumni — and lands in the open one.
Their earlier product, Tinker (Oct 2025), already leaned open — a platform to fine-tune Llama, gpt-oss and DeepSeek without managing training infra. Inkling is the house model for that world: built to be adapted, not just called.
When the ex-CTO of the closed leader ships open weights under Apache 2.0, that's a market signal: open isn't the budget option or the fast-follower anymore — it's a deliberate path chosen by frontier-caliber people.
And the runtime they picked to launch on is vLLM, the open inference stack.
Models change hands; the open runtime is the constant.
"The former OpenAI CTO's first move was to give the weights away. Open isn't the underdog play anymore — it's a strategy the frontier is choosing."
The whole model ran on vLLM — the open inference runtime — from launch day. When a frontier lab picks the open runtime as its launch partner, that's a real vote for the open stack. The specs, quickly:
Inkling lands in the middle of the OpenAI exodus and the open-vs-closed war — the same week a second open frontier model (Kimi K3) dropped.
The people who built the closed leader are now betting the value is in customization and the open stack, not owning the biggest model.
When the frontier's own alumni go open, "open" stops being the underdog play.