Inkling bets on customization

Thinking Machines introduced Inkling, a Mixture-of-Experts model with 975 billion total parameters, 41 billion active parameters, up to one million tokens of context, and text, image, and audio input.

The company acknowledges it is not the strongest model in absolute terms. Its bet is different: Apache 2.0 weights, configurable reasoning, and direct integration with Tinker to adapt behavior to specific domains.

Kimi K3 scales toward long-horizon work

Kimi announced K3 as a 2.8-trillion-parameter model with native vision and a million-token context window. It activates a fraction of its experts and is positioned for coding, knowledge, and long-horizon reasoning.

Published results are provider evaluations and require independent validation. At announcement time, Kimi said full weights would arrive on July 27: available in a product does not automatically mean available for self-hosting.

Large, open, and useful are not synonyms

Size describes potential capacity, not serving cost, latency, licensing fit, or adaptation effort. It also does not tell you how the model will behave with a company’s tools and data.

Useful evaluation starts with the workflow: two or three real tasks, quality criteria, cost per result, response time, and recovery behavior say more than a general table.

  • Compare cost per result, not isolated token pricing.
  • Verify the license, actual weight availability, and infrastructure needs.
  • Test adaptation and tool use inside your own process.