Large Model API Pricing

Explore the pricing for our model API. With transparent rates and flexible options, find the right plan to meet your needs.

Anthropic logo

Anthropic

Anthropic's Claude model offers advanced AI safety capabilities, focusing on useful, harmless, and honest AI assistants with powerful reasoning and conversational abilities.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
claude-haiku-4-5-202510011-2K20K$5$13 (5m)$12$10
2K-10K20K$3$8 (5m)$8$4
claude-3-7-sonnet-20250219-200K$3$3.75 (5 m)$0.3$15
claude-sonnet-4-20250514-200K$3$3.75 (5 min) · $6.60 (1 hr)$0.3$15
claude-opus-4-20250514-200K$15$18.75 (5 m)$1.5$75
claude-opus-4-1-20250805-200K$15$18.75 (5 m)$1.5$75
claude-sonnet-4-5-20250929官方资源1-2K200K$2$4(5m) × $8(1h)$$4$2
2K-20K200K$4$6(5 min) × $10(1 hr)$$6$4
claude-3-5-sonnet-20241022-200K$3$3.75 (5 m)$0.3$15
claude-3-haiku-20240307-200K$0.25--$1.25
claude-3-5-haiku-20241022-200K$0.8--$4
OpenAI

OpenAI

OpenAI's GPT series of models offer state-of-the-art language understanding and generation capabilities, delivering outstanding performance across a wide range of tasks, and are among the industry's leading AI models.

Model NameContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
gpt-5-codex400K$1.25-$0.125$10
OpenAI GPT OSS 120B131.1K$0.1--$0.5
OpenAI: GPT OSS 20B131.1K$0.05--$0.2
gpt-5-mini400K$0.25-$0.025$2
gpt-5-nano400K$0.05-$0.005$0.4
gpt-5-pro400K$15$1 (1 hour)-$120
gpt-5.6-sol400K$1$5(30m)$3$2
gpt-5-chat-latest400K$1.25-$0.125$10
gpt-5500K$1.25$0.1(5m) \times $0.2(1h)$0.125$10
gpt-4.1-mini1M$0.4-$0.1$1.6
gpt-4.1-nano1M$0.1-$0.025$0.4
gpt-4.11M$2-$0.5$8
gpt-4o-mini131.1K$0.15-$0.075$0.6
gpt-4o131.1K$2.5-$1.25$10
Gemini logo

Gemini

Google's Gemini model offers high-quality natural language processing capabilities, performs exceptionally well across a wide range of NLP tasks, and boasts powerful multimodal capabilities.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
gemini-3.1-flash-lite-preview—naer官方资源1-24.8M1M
Free
--
Free
Gemma3 12B-131.1K$0.05--$0.1
gemini-2.5-flash-1M$0.3$0.083 (5m)$0.075$2.5
gemini-2.5-pro-1M$1.25$0.375 (5m)$0.3125$10
Gemma 3 27B-32.8K$0.119--$0.2
gemini-3.1-flash-lite-preview低价专区1-24.8M1M$10000$10002(5m)·$10003(1h)$10001$10007
gemini-2.5-flash-lite-preview-09-2025-1M$0.1$0.083 (5m)$0.01$0.4
gemini-2.0-flash-lite-1M$0.075$0.083 (5m)$0.0188$0.3
gemini-2.5-flash-lite-1M$0.1$0.083 (5m)$0.025$0.4
gemini-2.5-flash-lite-preview-06-17-1M$0.1--$0.4
gemini-2.5-flash-preview-05-20-1M$0.15--$3.5
gemini-2.5-pro-preview-06-05-1M$1.25--$10
gemini-2.0-flash-20250609-1M$0.15--$0.6
Llama logo

Llama

Meta's Llama model offers state-of-the-art language understanding capabilities and features an open architecture, making it suitable for a wide range of applications.

Model NameContextInput (/Mt)Output (/Mt)Operation
Llama 3.1 8B Instruct16.4K$0.02$0.05
Llama 3.3 70B Instruct131.1K$0.13$0.39
Llama 4 Maverick Instruct1M$0.17$0.85
Llama 4 Scout Instruct131.1K$0.1$0.5
Llama 3.2 3B Instruct32.8K$0.03$0.05
Qwen logo

Qwen

The Qwen series of models offers powerful natural language processing capabilities and is available in a range of parameter sizes, from lightweight to enterprise-grade solutions.

Model NameContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
Qwen/Qwen3-8B-
Free
--
Free
Qwen3 Next 80B A3B Thinking低价专区65.5K$6$6(5m)·$6(1h)$6$6
Qwen3 Coder 480B A35B Instruct262.1K$0.29--$1.2
Qwen3 235B A22b Thinking 2507131.1K$0.3--$3
Qwen3 235B A22B Instruct 2507131.1K$0.15--$0.8
Qwen 2.5 72B Instruct32K$0.38--$0.4
Qwen3 235B A22B41K$0.2--$0.8
Qwen2.5 VL 72B Instruct32.8K$0.8--$0.8
Qwen3 32B41K$0.1--$0.45
Qwen3 30B A3B41K$0.09--$0.45
Qwen3 Next 80B A3B Instruct低价专区65.5K$0.15--$1.5
Qwen MT Plus4.1K$0.25--$0.75
Qwen3 8B128K$0.035--$0.138
Qwen2.5 7B Instruct32K$0.07--$0.07
Wenxin

Baidu

Baidu's ERNIE model offers advanced Chinese language understanding and multimodal capabilities, is optimized for Chinese applications, and is competitively priced.

Model NameContextInput (/Mt)Output (/Mt)Operation
ERNIE 4.5 VL 424B A47B123K$0.42$1.25
ERNIE 4.5 300B A47B123K$0.28$1.1
ChatGLM

THUDM

The GLM series of models from Tsinghua University feature advanced Chinese language understanding and generation capabilities.

Model NameContextInput (/Mt)Output (/Mt)Operation
GLM-4.5131.1K$0.6$2.2
GLM 4.5V65.5K$0.6$1.8
GLM 4.1V 9B Thinking65.5K$0.035$0.138
Sao10K logo

Sao10K

A fine-tuned model specifically optimized for creative and role-playing applications, featuring enhanced storytelling capabilities.

Model NameContextInput (/Mt)Output (/Mt)Operation
L3 70B Euryale V2.1 8.2K$1.48$1.48
Sao10k L3 8B Lunaris 8.2K$0.05$0.05
L3 8B Stheno V3.28.2K$0.05$0.05
L31 70B Euryale V2.28.2K$1.48$1.48
Mistralai logo

Mistralai

A powerful and efficient language model from Mistral AI, designed for both commercial and open-source applications.

Model NameContextInput (/Mt)Output (/Mt)Operation
Mistral Nemo60.3K$0.04$0.17
Mistral 7B Instruct32.8K$0.029$0.059
Deepseek logo

Deepseek

Advanced AI models from DeepSeek, offering cutting-edge inference capabilities and competitive pricing for enterprise and research applications.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
deepseek/deepseek-v3.1-test-20K
Free
--
Free
DeepSeek V3.1-163.8K$20$1 (5m)$1$100
DeepSeek R1 05281-32.8K163.8K$0.015$0.006(5m)$0.009001$0.06
131.1K-204.8K163.8K$0.03$0.007(5m)$0.005001$0.06
32.8K-131.1K163.8K$0.08$0.005(5m)$0.003001$0.04
DeepSeek V3 0324-163.8K$0.28$0.14 (5m)$0.14$1.14
MiniMax logo

MiniMax

MiniMax AI's advanced language model delivers powerful conversational AI capabilities, excelling in customer service, content generation, and creative applications, with robust multilingual support and enterprise-grade scalability.

Model NameContextInput (/Mt)Output (/Mt)Operation
MiniMax M11M$0.55$2.2
Gryphe logo

Gryphe

An innovative AI model from Gryphe that offers professional-grade language understanding capabilities, with a focus on efficiency and adaptability, making it ideal for niche applications.

Model NameContextInput (/Mt)Output (/Mt)Operation
Mythomax L2 13B4.1K$0.09$0.09

Mixture of Experts

A sophisticated collection of state-of-the-art AI models, featuring advanced reasoning and mathematical proof capabilities, as well as cutting-edge language understanding across multiple domains.

Model NameInput Token RangeContextInput (/Mt)Cache Write (/Mt)Cache read (/Mt)Output (/Mt)Operation
DeepSeek V3.1-163.8K$20$1 (5m)$1$100
OpenAI GPT OSS 120B-131.1K$0.1--$0.5
GLM-4.5-131.1K$0.6--$2.2
Qwen3 235B A22b Thinking 2507-131.1K$0.3--$3
GLM 4.5V-65.5K$0.6--$1.8
OpenAI: GPT OSS 20B-131.1K$0.05--$0.2
MiniMax M1-1M$0.55--$2.2
DeepSeek R1 05281-32.8K163.8K$0.015$0.006(5m)$0.009001$0.06
131.1K-204.8K163.8K$0.03$0.007(5m)$0.005001$0.06
32.8K-131.1K163.8K$0.08$0.005(5m)$0.003001$0.04
Qwen3 235B A22B-41K$0.2--$0.8
Llama 4 Maverick Instruct-1M$0.17--$0.85
Llama 4 Scout Instruct-131.1K$0.1--$0.5
2221–200222$2$5 (5m)$4$3
200-50K222$3$6 (5m)$5$4
50K-250K222$4$7 (5m)$6$5
ERNIE 4.5 VL 424B A47B-123K$0.42--$1.25
ERNIE 4.5 300B A47B-123K$0.28--$1.1
Qwen3 32B-41K$0.1--$0.45
Qwen3 30B A3B-41K$0.09--$0.45
Kimi K2 Instruct-131.1K$0.57--$2.3
DeepSeek V3 0324-163.8K$0.28$0.14 (5m)$0.14$1.14
test-model-interface-21-32K65K$5.1$8.1 (5m)$7.1$6.1
32K-128K65K$5.3$8.3 (5m)$7.3$6.3
128K-256K65K$5.2$8.2 (5m)$7.2$6.2
Contact Us