Mira Murati’s Thinking Machines Lab has officially entered the fray with Inkling, an open-weight model built to challenge the dominance of centralized, “one-size-fits-all” AI giants. By betting on customizability instead of raw, general-purpose power, the startup says the future of enterprise AI lies in specialized, locally optimized intelligence.

The Architecture of Inkling: Efficiency Through MoE

Inkling is a Mixture-of-Experts system with a massive 975 billion total parameters. Yet it activates only about 41 billion parameters for any single task, trimming latency and cost while preserving high-level reasoning. The model learned from 45 trillion tokens of text, image, audio and video, giving it native multimodal abilities. The company claims Inkling matches Nvidia’s Nemotron 3 Ultra on coding benchmarks while using just one-third of the tokens, a clear sign of training efficiency.

A Strategic Pivot Toward Enterprise Customization

Unlike the flagship products from OpenAI, Anthropic or Google, which function mainly as general-purpose chatbots, Inkling is marketed as a foundation for organizations. Thinking Machines pitches the model as a starting point, urging enterprises to fine-tune it with proprietary data via its “Tinker” platform.

The move tackles the industry’s “double payment” trap. Microsoft CEO Satya Nadella warned that companies pay twice—once for a subscription and again by feeding unique business logic into a model that ultimately enriches the provider’s training set. By offering an open-weight model, Thinking Machines lets firms keep that intelligence in-house.

High-Stakes Infrastructure and Rapid Development

Thinking Machines sprinted to market. OpenAI took roughly five years and Anthropic about three; Thinking Machines says it reached a comparable milestone in just nine months.

To power that speed, the firm leaned on hardware. After a partnership with Nvidia in March, Inkling trained entirely on Nvidia’s GB300 NVL72 systems. The company used distillation tricks—leveraging open-weight models like Moonshot AI’s Kimi K2.5 for early post-training—but vows to run future post-training fully in-house.

Key Takeaways

  • Open-Weight Advantage: Developers can download and modify the 975 B-parameter MoE model, keeping control away from closed ecosystems.
  • Efficiency-First Design: Activating only 41 B parameters per task and trimming token use lets the model deliver top-tier reasoning at a fraction of the operational cost.
  • The Customization Bet: Thinking Machines shifts from generic chatbots to “Tinker,” a platform that turns foundation models into specialized assets for each organization.