The free Inkling endpoint is only available for use with agentic harnesses. Do not upload any confidential information or personal data (e.g., voices and images of people's faces). Your usage of this free endpoint, including prompts and outputs, is logged and used to improve Thinking Machines Lab's models, products, and services. The logged session data will be disassociated from your account and other persistent identifiers before being used for these purposes.
By using this free endpoint, you agree to the TML Free Research API Terms of ServiceOpens in new tab. For more information about Thinking Machines Lab's data processing practices, see this Privacy NoticeOpens in new tab.
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems, retrieval-augmented generation, instruction following, and multilingual conversational applications. Its native image and audio understanding supports multimodal analysis alongside text.
| $0.95 | $4.05 | $0.16 | 0.73s | 48 tps | ||
| $1.00 | $4.05 | $0.17 | 1.29s | 82 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
