MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for greater computational efficiency. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers strong performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.
| $0.14 | $0.28 | $0.0028 | 4.03s | 27 tps | ||
| $0.14 | $0.28 | $0.0028 | 2.33s | 31 tps | ||
| $0.14 | $0.28 | $0.0028 | 4.00s | 38 tps | ||
| $0.14 | $0.28 | $0.003 | 4.86s | 3 tps | ||
| $0.175 | $0.35 | $0.00375 | 3.81s | 22 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
