GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and GPT-4.5 across coding (54.6% SWE-bench Verified), instruction compliance (87.4% IFEval), and multimodal understanding benchmarks. It is tuned for precise code diffs, agent reliability, and high recall in large document contexts, making it ideal for agents, IDE tooling, and enterprise knowledge retrieval.
| $2.00 | $8.00 | $0.50 | 1.51s | 39 tps | ||
| $2.00 | $8.00 | $0.50 | 1.17s | 58 tps | ||
Not used in Standard routing:Why these endpoints are not used | ||||||
| $2.20 | $8.80 | $0.55 | 0.63s | 43 tps | ||
P50, best across providers
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.