Thinking Machines Lab on July 15 released Inkling, a 975-billion-parameter multimodal artificial intelligence model with open weights under an Apache 2.0 license on Hugging Face—an unusually permissive launch that positions the system for immediate use in self-hosted crypto trading stacks, blockchain research workflows, and other data-sensitive applications. The company, founded by former OpenAI executive Mira Murati, says the model was trained from scratch and is built for agentic tool use, an area of growing interest across digital asset markets where automation and integrations matter.
Technology Use Case
Inkling is a mixture-of-experts model, an architecture in which only a subset of the network activates for any given input. That design aims to preserve depth while keeping inference efficient—attributes relevant for crypto-facing applications that need to parse heterogeneous inputs and trigger external tools. Multimodality means the system accepts text, images, and audio, while a 1 million-token context window—roughly 750,000 words—allows it to hold long-form discussions, persistent instructions, or compliance documentation in memory without frequent truncation. Thinking Machines says the model was pretrained on 45 trillion tokens spanning text, images, audio, and video.
Agentic performance is where the company highlights early gains. Inkling scores 74.1% on MCP Atlas, a benchmark that measures how reliably an AI agent completes real-world tasks using Model Context Protocol to connect assistants to external services. For crypto users, that orientation to tool use maps to common needs like orchestrating pipelines that analyze market data, triaging research across multiple sources, or coordinating back-end services without manual overhead. The open-weights approach allows these workflows to run in a controlled environment, which is often preferred when handling proprietary strategies or regulated information.
AI Integration
Thinking Machines describes Inkling as a generalist system intended to avoid over-optimizing for one task at the expense of others. On SWE-Bench Verified—a test of whether an AI agent can autonomously fix real GitHub software bugs—the model scores 77.6%. For developer teams in digital assets that maintain complex open-source codebases, agentic help on routine issues can reduce context switching and accelerate fixes, while the large context window supports longer traces, logs, and documentation during troubleshooting.
Fine-tuning is available through Tinker, Thinking Machines’ cloud platform. Because the full weights are openly hosted on Hugging Face, teams can also adapt the model privately under the Apache 2.0 license, a setup compatible with institutions that require control over both model behavior and deployment perimeter. That combination—open weights plus fine-tune tooling—aligns with how many crypto organizations run infrastructure, where self-hosting and configuration control are common requirements.
Benchmark Landscape
Inkling’s agentic showing comes with caveats that matter for teams selecting a base model. Thinking Machines notes that Chinese models retain the edge on several fronts: Z.ai’s GLM 5.2 scores 82.7% on Terminal Bench 2.1, a benchmark for autonomous coding agents operating in a real terminal environment, compared with Inkling’s 63.8%. Kimi K2.6 leads on Humanity’s Last Exam, a test of PhD-level scientific reasoning. The company acknowledges that Inkling is not the strongest model available today, open or closed, even as it seeks to balance breadth and reliability.
Within the open-weights category in the West, however, Inkling’s maker highlights relative strengths. The 74.1% MCP Atlas result sits nearly 30 points above Nvidia’s Nemotron 3 Ultra in the comparison presented. On safety, the model posts 78.0% on FORTRESS Adversarial, a benchmark that assesses whether a system consistently refuses genuinely harmful prompts without over-blocking legitimate ones. For crypto applications where automated agents may interact with user inputs, public channels, or developer tooling, a safety profile that aims to filter harmful instructions while preserving task completion is a practical consideration.
Market Impact
The licensing and distribution choices are likely to draw attention from teams that cannot route workloads through models developed in Beijing, or that simply prefer to avoid external dependencies. Thinking Machines frames Inkling as the most capable open-weights model built by a Western lab for agentic tool use, offering an alternative to self-hosting Chinese models. For crypto market participants—who often operate cross-border while observing strict internal controls—the ability to deploy an open, fine-tunable system can simplify procurement, auditing, and integration.
The open-weights decision also intersects with budget and infrastructure planning. Because Inkling uses a mixture-of-experts architecture with 975 billion total parameters and 41 billion active per task, it is far beyond the scale of a typical local deployment. That reality keeps the model in data centers, but the absence of usage restrictions and the availability of weights still support private deployments, including environments where market data, user information, or research artifacts cannot leave a firm’s boundary.
Corporate Backdrop
Inkling arrives roughly two years after Mira Murati left OpenAI in September 2024. She had served briefly as interim CEO when the OpenAI board fired Sam Altman in November 2023, before Altman returned five days later and she resumed the CTO role; Murati departed about ten months after that. Thinking Machines Lab was founded in February 2025 and raised $2 billion at a $12 billion valuation in July 2025, led by Andreessen Horowitz with Nvidia, Accel, ServiceNow, Cisco, AMD, and Jane Street participating—described at the time as one of Silicon Valley’s largest seed rounds. Reports in November 2025 placed the company in talks for a new round at a $50 billion valuation, which collapsed by January 2026.
Product Line
Alongside the flagship model, the company previewed Inkling-Small, which has 276 billion total parameters with 12 billion active. According to Thinking Machines, the smaller system already matches the larger model on most reasoning benchmarks. The team says the weights for Inkling-Small will be released once testing is complete, with no timeline provided.
Industry Response
Thinking Machines presents Inkling as a “well-rounded” generalist that does not sacrifice capability in one area to excel in another. For crypto developers, integrators, and research desks who have kept a close eye on open-weights progress, the release offers a Western option organized around agentic workflows and permissive licensing. While benchmark leaders from China retain a performance lead on several key tests, Inkling’s combination of tool-use reliability, safety scores, and open distribution may broaden the pool of models considered for blockchain analytics, exchange infrastructure, and other market-facing systems where autonomy and control are prized.
Ultimately, the launch adds a high-capacity, openly available model to the mix just as more teams in digital assets are experimenting with agent-based automation. Thinking Machines’ choice to open the weights, support fine-tuning on Tinker, and emphasize tool-use benchmarks positions Inkling as an option for organizations that need to keep data and decision logic in-house while connecting AI agents to external services. Whether and how teams adapt it for their production environments will depend on their tolerance for trade-offs noted in the benchmark landscape, their operational constraints, and their appetite for running large-scale models under a permissive license.

