
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts architecture that accepts text and image input, and it always operates in a thinking mode, preserving full reasoning content across multi-turn conversations. With a 256K-token context window, it targets long-horizon coding, agentic task decomposition, and multi-turn dialogue. The model activates 32B parameters out of roughly 1T total.
| $0.6712 | $3.35 | $0.18 | 1.01s | 46 tps | ||
| $0.71 | $3.50 | $0.15 | 0.72s | 142 tps | ||
25% off | $0.95$0.7125 | $4.00$3.00 | $0.19$0.1425 | 1.99s | 54 tps | |
| $0.75 | $3.50 | $0.16 | 2.55s | 18 tps | ||
| $0.85916 | $3.80 | $0.17993 | 2.33s | 25 tps | ||
4% off | $0.95$0.912 | $4.00$3.84 | $0.19$0.1824 | 1.81s | 59 tps | |
| $0.95 | $4.00 | -- | 1.30s | 15 tps | ||
| $0.95 | $4.00 | $0.19 | 5.28s | 55 tps | ||
| $0.95 | $4.00 | $0.19 | 1.59s | 34 tps | ||
| $0.95 | $4.00 | $0.19 | 1.14s | 33 tps | ||
| $0.95 | $4.00 | $0.19 | 1.01s | 65 tps | ||
| $1.90 | $8.00 | $0.38 | 1.22s | 133 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.
