DeepSeek V4.1 Flash
DeepSeek V4.1 Flash: multimodal reasoning, coding and tool use with a 1M-token context window. Separate from DeepSeek V4 Flash.
Input
Output
Input
≈ $0.375
Output
≈ $1.5
Cache
≈ $0.0075
Showing 21 groups
DeepSeek V4.1 Flash: multimodal reasoning, coding and tool use with a 1M-token context window. Separate from DeepSeek V4 Flash.
Input
Output
Input
≈ $0.375
Output
≈ $1.5
Cache
≈ $0.0075
DeepSeek V4 Flash: a distinct text-only V4 model for reasoning, coding and tool use. This model ID does not select V4.1.
Input
Output
Input
≈ $0.525
Output
≈ $1.57
Cache
≈ $0.0175
GLM 5.3 Flash: multimodal reasoning, coding and tool use with a 1M-token context window. Thinking is always enabled with low, high or max effort.
Input
Output
Input
≈ $0.14
Output
≈ $0.49
Cache
≈ $0.0405
Google Gemini 3.5 Flash — a fast, cost-efficient multimodal chat model with image+text input and a million-token context window, for high-concurrency workloads
Input
Output
Input
≈ $0.938
Output
≈ $6.25
Cache
≈ $0.0938
MiniMax M3 — a large language model with a million-token context window, excelling at long-document understanding, complex reasoning, and tool use
Input
Output
Input
≈ $1.2
Output
≈ $5
Cache
≈ $1.5
DeepSeek V4 — a reasoning-focused LLM with a 128K-token context window, excelling at code generation and complex logic at a highly competitive price
Input
Output
Input
≈ $0.212
Output
≈ $0.475
Cache
≈ $0.025
OpenAI GPT-5.5 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $4.25
Output
≈ $25
Cache
≈ $0.425
OpenAI GPT-5.6 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $6
Output
≈ $40
Cache
≈ $0.85
OpenAI GPT-6 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $3.5
Output
≈ $30
Cache
≈ $0.531
Grok 4.5 is an advanced AI model designed to deliver fast, writing, coding, research, data analysis, and creative problem-solving through natural conversation
Input
Output
Input
≈ $0.6
Output
≈ $1.2
Cache
≈ $0.6
Grok 4.6 is an advanced AI model designed to deliver fast, writing, coding, research, data analysis, and creative problem-solving through natural conversation
Input
Output
Input
≈ $0.6
Output
≈ $1.2
Cache
≈ $0.6
A premium reasoning route for visual front-end prototyping, repository-scale coding, large evidence sets, long-running agents, and complex knowledge work that benefits from a 1.05M-token working context.
Input
Output
Input
≈ $2
Output
≈ $10
Cache
≈ $0.2
Grok 4.7 for text conversations and coding through an OpenAI-compatible API. Priced at the same credit rates as Grok 4.6.
Input
Output
Input
≈ $0.6
Output
≈ $1.2
Cache
≈ $0.6
Google Gemini 2.5 Flash Lite — an ultra-low-cost, low-latency chat model with a million-token context window, ideal for high-frequency everyday tasks
Input
Output
Input
≈ $0.1
Output
≈ $0.15
Cache
≈ $0.01
Google Gemini 3.1 Flash Lite — improved reasoning over the 2.5 generation while staying economical, balancing speed and quality for lightweight tasks
Input
Output
Input
≈ $0.15
Output
≈ $1
Cache
≈ $0.0125
Google Gemini 3.1 Pro — the flagship multimodal Gemini model, offering strong complex reasoning and long-context capability for demanding production apps
Input
Output
Input
≈ $1.25
Output
≈ $7.5
Cache
≈ $0.625
OpenAI GPT-4o mini — a fast, affordable multimodal chat model with quick responses and low cost, ideal for everyday Q&A and lightweight coding help
Input
Output
Input
≈ $0.275
Output
≈ $0.412
Cache
≈ $0.0138
OpenAI GPT-5.4 — a high-capability model for advanced reasoning, code generation, and agentic workflows, with a 400K-token context window for production use
Input
Output
Input
≈ $2.5
Output
≈ $15
Cache
≈ $0.25
Anthropic Claude Opus 4.8 — the flagship Claude model for the most demanding reasoning, coding, and long-form writing tasks, with excellent long context
Input
Output
Input
≈ $2.5
Output
≈ $12
Cache
≈ $0.25
Anthropic Claude Opus 5 from APIAny — Anthropic’s newest Opus-tier flagship for the hardest coding, long-running agents, and judgment-heavy review.
Input
Output
Input
≈ $3.75
Output
≈ $18.75
Cache
≈ $0.375
Anthropic Claude Sonnet 4.6 — a balanced flagship model delivering strong reasoning at production speed and cost, for large-scale, stable deployment
Input
Output
Input
≈ $1.5
Output
≈ $7.5
Cache
≈ $1
APIAny lists public chat and text models on one OpenAI-compatible endpoint, with live credit prices on each model page.
By APIAny Editorial · Updated
Filter by type or provider, compare credit prices, then open a model page for playground tests and request examples.
Sources
The quotations below are from official API documentation that APIAny implements against.
The Chat Completions API endpoint will generate a model response from a list of messages comprising a conversation.
The Gemini API provides access to Google's most capable generative AI models.
The Images API provides several endpoints that let you generate images from text prompts or create edits of existing images.
The APIAny models catalog is the live list of public chat, image, video, audio, and safety models you can call with one API key.