
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated DeltaNet linear attention with standard gated attention layers, enabling efficient inference at a fraction of the compute cost. The model supports a 262K token native context window (extensible to 1M via YaRN) and accepts text, image, and video inputs. It includes integrated thinking mode with reasoning traces preserved across multi-turn conversations, function calling, and structured output. Released under the Apache 2.0 license.
| $0.05 | $0.70 | $0.025 | 1.01s | 25 tps | ||
| $0.10 | $0.90 | $0.05 | 0.41s | 44 tps | ||
| $0.10 | $0.95 | -- | 0.91s | 22 tps | ||
| $0.10 | $1.00 | $0.05 | 0.44s | 67 tps | ||
| $0.10 | $1.00 | -- | 0.72s | 76 tps | ||
| $0.15 | $1.00 | $0.05 | 0.66s | 24 tps | ||
| $0.186 | $1.11375 | $0.186 | 1.09s | 38 tps | ||
| $0.20 | $1.27 | $0.056 | 1.88s | 41 tps | ||
| $0.24 | $1.80 | $0.15 | 1.25s | 33 tps | ||
| $0.25 | $1.25 | $0.25 | 0.32s | 76 tps |
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.