GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency use cases such as classification, data extraction, ranking, and sub-agent execution.
The model prioritizes responsiveness and efficiency over deep reasoning, making it ideal for pipelines that require fast, reliable outputs at scale. GPT-5.4 nano is well suited for background tasks, real-time systems, and distributed agent architectures where minimizing cost and latency is essential.
| $0.20 | $1.25 | $0.02 | 2.06s | 52 tps | ||
| $0.20 | $1.25 | $0.02 | 1.02s | 53 tps | ||
Not used in Standard routing:Why these endpoints are not used | ||||||
Flex | $0.10 | $0.625 | $0.01 | 1.68s | 55 tps | |
| $0.22 | $1.375 | $0.022 | -- | -- | ||
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.