Explore the pricing for our model API. With transparent rates and flexible options, find the right plan to meet your needs.
Anthropic's Claude model offers advanced AI safety capabilities, focusing on useful, harmless, and honest AI assistants with powerful reasoning and conversational abilities.
| Model Name | Input Token Range | Context | Input (/Mt) | Cache Write (/Mt) | Cache read (/Mt) | Output (/Mt) | Operation |
|---|---|---|---|---|---|---|---|
| claude-haiku-4-5-20251001 | 1-2K | 20K | $5 | $13 (5m) | $12 | $10 | |
| 2K-10K | 20K | $3 | $8 (5m) | $8 | $4 | ||
| claude-3-7-sonnet-20250219 | - | 200K | $3 | $3.75 (5 m) | $0.3 | $15 | |
| claude-sonnet-4-20250514 | - | 200K | $3 | $3.75 (5 min) · $6.60 (1 hr) | $0.3 | $15 | |
| claude-opus-4-20250514 | - | 200K | $15 | $18.75 (5 m) | $1.5 | $75 | |
| claude-opus-4-1-20250805 | - | 200K | $15 | $18.75 (5 m) | $1.5 | $75 | |
| claude-sonnet-4-5-20250929官方资源 | 1-2K | 200K | $2 | $4(5m) × $8(1h)$ | $4 | $2 | |
| 2K-20K | 200K | $4 | $6(5 min) × $10(1 hr)$ | $6 | $4 | ||
| claude-3-5-sonnet-20241022 | - | 200K | $3 | $3.75 (5 m) | $0.3 | $15 | |
| claude-3-haiku-20240307 | - | 200K | $0.25 | - | - | $1.25 | |
| claude-3-5-haiku-20241022 | - | 200K | $0.8 | - | - | $4 |
OpenAI's GPT series of models offer state-of-the-art language understanding and generation capabilities, delivering outstanding performance across a wide range of tasks, and are among the industry's leading AI models.
| Model Name | Context | Input (/Mt) | Cache Write (/Mt) | Cache read (/Mt) | Output (/Mt) | Operation |
|---|---|---|---|---|---|---|
| gpt-5-codex | 400K | $1.25 | - | $0.125 | $10 | |
| OpenAI GPT OSS 120B | 131.1K | $0.1 | - | - | $0.5 | |
| OpenAI: GPT OSS 20B | 131.1K | $0.05 | - | - | $0.2 | |
| gpt-5-mini | 400K | $0.25 | - | $0.025 | $2 | |
| gpt-5-nano | 400K | $0.05 | - | $0.005 | $0.4 | |
| gpt-5-pro | 400K | $15 | $1 (1 hour) | - | $120 | |
| gpt-5.6-sol | 400K | $1 | $5(30m) | $3 | $2 | |
| gpt-5-chat-latest | 400K | $1.25 | - | $0.125 | $10 | |
| gpt-5 | 500K | $1.25 | $0.1(5m) \times $0.2(1h) | $0.125 | $10 | |
| gpt-4.1-mini | 1M | $0.4 | - | $0.1 | $1.6 | |
| gpt-4.1-nano | 1M | $0.1 | - | $0.025 | $0.4 | |
| gpt-4.1 | 1M | $2 | - | $0.5 | $8 | |
| gpt-4o-mini | 131.1K | $0.15 | - | $0.075 | $0.6 | |
| gpt-4o | 131.1K | $2.5 | - | $1.25 | $10 |
Google's Gemini model offers high-quality natural language processing capabilities, performs exceptionally well across a wide range of NLP tasks, and boasts powerful multimodal capabilities.
| Model Name | Input Token Range | Context | Input (/Mt) | Cache Write (/Mt) | Cache read (/Mt) | Output (/Mt) | Operation |
|---|---|---|---|---|---|---|---|
| gemini-3.1-flash-lite-preview—naer官方资源 | 1-24.8M | 1M | Free | - | - | Free | |
| Gemma3 12B | - | 131.1K | $0.05 | - | - | $0.1 | |
| gemini-2.5-flash | - | 1M | $0.3 | $0.083 (5m) | $0.075 | $2.5 | |
| gemini-2.5-pro | - | 1M | $1.25 | $0.375 (5m) | $0.3125 | $10 | |
| Gemma 3 27B | - | 32.8K | $0.119 | - | - | $0.2 | |
| gemini-3.1-flash-lite-preview低价专区 | 1-24.8M | 1M | $10000 | $10002(5m)·$10003(1h) | $10001 | $10007 | |
| gemini-2.5-flash-lite-preview-09-2025 | - | 1M | $0.1 | $0.083 (5m) | $0.01 | $0.4 | |
| gemini-2.0-flash-lite | - | 1M | $0.075 | $0.083 (5m) | $0.0188 | $0.3 | |
| gemini-2.5-flash-lite | - | 1M | $0.1 | $0.083 (5m) | $0.025 | $0.4 | |
| gemini-2.5-flash-lite-preview-06-17 | - | 1M | $0.1 | - | - | $0.4 | |
| gemini-2.5-flash-preview-05-20 | - | 1M | $0.15 | - | - | $3.5 | |
| gemini-2.5-pro-preview-06-05 | - | 1M | $1.25 | - | - | $10 | |
| gemini-2.0-flash-20250609 | - | 1M | $0.15 | - | - | $0.6 |
Meta's Llama model offers state-of-the-art language understanding capabilities and features an open architecture, making it suitable for a wide range of applications.
The Qwen series of models offers powerful natural language processing capabilities and is available in a range of parameter sizes, from lightweight to enterprise-grade solutions.
| Model Name | Context | Input (/Mt) | Cache Write (/Mt) | Cache read (/Mt) | Output (/Mt) | Operation |
|---|---|---|---|---|---|---|
| Qwen/Qwen3-8B | - | Free | - | - | Free | |
| Qwen3 Next 80B A3B Thinking低价专区 | 65.5K | $6 | $6(5m)·$6(1h) | $6 | $6 | |
| Qwen3 Coder 480B A35B Instruct | 262.1K | $0.29 | - | - | $1.2 | |
| Qwen3 235B A22b Thinking 2507 | 131.1K | $0.3 | - | - | $3 | |
| Qwen3 235B A22B Instruct 2507 | 131.1K | $0.15 | - | - | $0.8 | |
| Qwen 2.5 72B Instruct | 32K | $0.38 | - | - | $0.4 | |
| Qwen3 235B A22B | 41K | $0.2 | - | - | $0.8 | |
| Qwen2.5 VL 72B Instruct | 32.8K | $0.8 | - | - | $0.8 | |
| Qwen3 32B | 41K | $0.1 | - | - | $0.45 | |
| Qwen3 30B A3B | 41K | $0.09 | - | - | $0.45 | |
| Qwen3 Next 80B A3B Instruct低价专区 | 65.5K | $0.15 | - | - | $1.5 | |
| Qwen MT Plus | 4.1K | $0.25 | - | - | $0.75 | |
| Qwen3 8B | 128K | $0.035 | - | - | $0.138 | |
| Qwen2.5 7B Instruct | 32K | $0.07 | - | - | $0.07 |
A fine-tuned model specifically optimized for creative and role-playing applications, featuring enhanced storytelling capabilities.
Advanced AI models from DeepSeek, offering cutting-edge inference capabilities and competitive pricing for enterprise and research applications.
| Model Name | Input Token Range | Context | Input (/Mt) | Cache Write (/Mt) | Cache read (/Mt) | Output (/Mt) | Operation |
|---|---|---|---|---|---|---|---|
| deepseek/deepseek-v3.1-test | - | 20K | Free | - | - | Free | |
| DeepSeek V3.1 | - | 163.8K | $20 | $1 (5m) | $1 | $100 | |
| DeepSeek R1 0528 | 1-32.8K | 163.8K | $0.015 | $0.006(5m) | $0.009001 | $0.06 | |
| 131.1K-204.8K | 163.8K | $0.03 | $0.007(5m) | $0.005001 | $0.06 | ||
| 32.8K-131.1K | 163.8K | $0.08 | $0.005(5m) | $0.003001 | $0.04 | ||
| DeepSeek V3 0324 | - | 163.8K | $0.28 | $0.14 (5m) | $0.14 | $1.14 |
MiniMax AI's advanced language model delivers powerful conversational AI capabilities, excelling in customer service, content generation, and creative applications, with robust multilingual support and enterprise-grade scalability.
A sophisticated collection of state-of-the-art AI models, featuring advanced reasoning and mathematical proof capabilities, as well as cutting-edge language understanding across multiple domains.
| Model Name | Input Token Range | Context | Input (/Mt) | Cache Write (/Mt) | Cache read (/Mt) | Output (/Mt) | Operation |
|---|---|---|---|---|---|---|---|
| DeepSeek V3.1 | - | 163.8K | $20 | $1 (5m) | $1 | $100 | |
| OpenAI GPT OSS 120B | - | 131.1K | $0.1 | - | - | $0.5 | |
| GLM-4.5 | - | 131.1K | $0.6 | - | - | $2.2 | |
| Qwen3 235B A22b Thinking 2507 | - | 131.1K | $0.3 | - | - | $3 | |
| GLM 4.5V | - | 65.5K | $0.6 | - | - | $1.8 | |
| OpenAI: GPT OSS 20B | - | 131.1K | $0.05 | - | - | $0.2 | |
| MiniMax M1 | - | 1M | $0.55 | - | - | $2.2 | |
| DeepSeek R1 0528 | 1-32.8K | 163.8K | $0.015 | $0.006(5m) | $0.009001 | $0.06 | |
| 131.1K-204.8K | 163.8K | $0.03 | $0.007(5m) | $0.005001 | $0.06 | ||
| 32.8K-131.1K | 163.8K | $0.08 | $0.005(5m) | $0.003001 | $0.04 | ||
| Qwen3 235B A22B | - | 41K | $0.2 | - | - | $0.8 | |
| Llama 4 Maverick Instruct | - | 1M | $0.17 | - | - | $0.85 | |
| Llama 4 Scout Instruct | - | 131.1K | $0.1 | - | - | $0.5 | |
| 222 | 1–200 | 222 | $2 | $5 (5m) | $4 | $3 | |
| 200-50K | 222 | $3 | $6 (5m) | $5 | $4 | ||
| 50K-250K | 222 | $4 | $7 (5m) | $6 | $5 | ||
| ERNIE 4.5 VL 424B A47B | - | 123K | $0.42 | - | - | $1.25 | |
| ERNIE 4.5 300B A47B | - | 123K | $0.28 | - | - | $1.1 | |
| Qwen3 32B | - | 41K | $0.1 | - | - | $0.45 | |
| Qwen3 30B A3B | - | 41K | $0.09 | - | - | $0.45 | |
| Kimi K2 Instruct | - | 131.1K | $0.57 | - | - | $2.3 | |
| DeepSeek V3 0324 | - | 163.8K | $0.28 | $0.14 (5m) | $0.14 | $1.14 | |
| test-model-interface-2 | 1-32K | 65K | $5.1 | $8.1 (5m) | $7.1 | $6.1 | |
| 32K-128K | 65K | $5.3 | $8.3 (5m) | $7.3 | $6.3 | ||
| 128K-256K | 65K | $5.2 | $8.2 (5m) | $7.2 | $6.2 |