
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and agentic tool use. Post-trained on instruction data, it demonstrates competitive performance across reasoning (AIME, ZebraLogic), coding (MultiPL-E, LiveCodeBench), and alignment (IFEval, WritingBench) benchmarks. It outperforms its non-instruct variant on subjective and open-ended tasks while retaining strong factual and coding performance.
55% off | $0.107$0.04815 | $0.429$0.1931 | 0.80s | 52 tps | |
|---|---|---|---|---|---|
| $0.09 | $0.30 | 0.42s | 43 tps | ||
| $0.09 | $0.30 | 0.94s | 25 tps | ||
| $0.10 | $0.30 | 0.38s | 25 tps | ||
| $0.13 | $0.52 | 0.32s | 47 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.