September 2026 · Updated pricing

OpenAI API pricing: complete cost comparison

Every model, every price tier, every saving mechanism — broken down so you can pick the right model and pay the least for it.

OpenAI API pricing comparison and cost monitoring dashboard

Understanding per-token pricing is the first step to controlling AI costs.

How OpenAI API pricing works

OpenAI charges per token. One token is roughly four characters of English text, or about three-quarters of a word. Every API call costs two things: input tokens (the text you send to the model) and output tokens (the text the model generates back). The two are priced separately because generation is computationally more expensive than reading.

As of September 2026, OpenAI offers eight production models. They range from the frontier GPT-6 Astra down to the lightweight GPT-5.4 Mini, with a 66× price spread between the most and least expensive. Choosing the right model for your workload is the single biggest lever for controlling API spend.

Beyond the base per-token rates, OpenAI provides three cost-reduction mechanisms: prompt caching (automatic 50–90% input-token discount on repeated prefixes), Batch mode (50% off for non-real-time requests), and Flex mode (50% off for requests that tolerate best-effort delivery). These stack and can dramatically lower effective costs for workloads that fit their constraints.

Full pricing table: every OpenAI model (September 2026)

The table below lists all current production models with standard rates, cached-input rates, and batch/Flex mode rates. All prices are per one million tokens.

Model Input Output Cached input Batch / Flex input Batch / Flex output
GPT-6 Astra $10.00 $30.00 $1.00 $5.00 $15.00
GPT-5.6 Sol $2.50 $10.00 $0.25 $1.25 $5.00
GPT-5.6 Terra $2.50 $10.00 $0.25 $1.25 $5.00
GPT-5.6 Luna $2.50 $10.00 $0.25 $1.25 $5.00
GPT-5.5 $1.25 $5.00 $0.125 $0.625 $2.50
GPT-5.4 $0.75 $3.00 $0.075 $0.375 $1.50
GPT-5.4 Mini $0.15 $0.60 $0.015 $0.075 $0.30
GPT-5.3 Codex Spark $1.00 $4.00 $0.10 $0.50 $2.00

Key takeaway: output tokens are always 3–4× more expensive than input tokens. When you see a cost spike, the first thing to check is whether your prompts are generating longer outputs than necessary.

Batch and Flex mode: 50% off

Batch mode lets you send a collection of requests as a single asynchronous job. Results are delivered within 24 hours. Because OpenAI can schedule these jobs during off-peak capacity, they offer a straight 50% discount on both input and output token rates.

Flex mode works similarly for real-time requests: you agree to occasional latency spikes (the request may be temporarily queued) in exchange for 50% off. For applications where a two-second delay is acceptable — background processing, scheduled analysis, non-interactive summarization — Flex mode is essentially free money.

Neither Batch nor Flex mode restricts which models you can use. Every model in the table above qualifies for both discounts. Combined with prompt caching, Batch mode on GPT-5.4 can bring effective input-token costs below $0.04 per million.

Prompt caching: how it works and what it saves

Prompt caching is automatic. When you send a request whose initial text (the "prefix") matches a recently used prefix of at least 1,024 tokens, OpenAI serves the cached portion from memory instead of recomputing it. You pay only 10% of the standard input-token price for the cached portion.

This is most impactful for workloads with large, repeated system prompts. If your agent includes 4,000 tokens of system instructions and tool definitions, and you make 500 API calls per day, prompt caching saves you roughly 90% of the input cost on that 4,000-token prefix across every call after the first one.

Prompt caching applies at the prefix level, not the entire prompt. If your system prompt is 3,000 tokens and your user message is 500 tokens, the 3,000-token prefix is cached and the 500-token suffix is always priced at the full rate. Structuring your prompts to put static content first and variable content last maximizes cache hits.

Cost calculator: three scenarios

The following estimates assume a 3:1 output-to-input ratio (common for chat and agent workloads) and no caching. Costs drop significantly when caching and Batch mode are applied.

Light use: personal tools and experiments

You build a personal summarizer and a few small scripts. Usage: roughly 2 million input tokens and 500,000 output tokens per month.

Model Monthly cost
GPT-6 Astra$35.00
GPT-5.5$5.00
GPT-5.4 Mini$0.60

For light use, even GPT-6 Astra stays under $40/month. If you're experimenting, GPT-5.4 Mini at $0.60/month is virtually free.

Moderate use: startup production workload

A small SaaS product uses the API for customer-facing chat, content generation, and internal data analysis. Usage: roughly 10 million input tokens and 3 million output tokens per month.

Model Monthly cost
GPT-6 Astra$190.00
GPT-5.5$27.50
GPT-5.4$16.50
GPT-5.4 Mini$3.30

At moderate volume, model choice has a dramatic impact. GPT-6 Astra costs 57× more than GPT-5.4 Mini. For most production workloads, GPT-5.5 or GPT-5.4 provides the best quality-to-cost ratio. Reserve Astra for tasks that genuinely require frontier-level reasoning.

Heavy use: coding agent at scale

A development team runs an AI coding agent that reviews every pull request, suggests refactors, and generates tests. Usage: roughly 50 million input tokens and 20 million output tokens per month.

Model Monthly cost
GPT-6 Astra$1,100.00
GPT-5.6 Sol$325.00
GPT-5.5$162.50
GPT-5.4$97.50
GPT-5.3 Codex Spark$130.00
GPT-5.4 Mini (Batch)$2.06

At heavy volume, Batch mode and caching become essential. Running GPT-5.4 Mini in Batch mode with caching for non-interactive parts of the pipeline cuts costs from over a thousand dollars to under three. Even GPT-5.5 with Batch mode drops to roughly $81/month — an 80% savings over standard pricing.

How model choice affects cost: same task, different models

The following table shows what a single typical task costs across models. The task: a 2,000-token system prompt, a 1,000-token user message, and a 500-token response — roughly equivalent to a structured data extraction call.

Model Per-call cost 1,000 calls vs Astra
GPT-6 Astra$0.045$45.001.0×
GPT-5.6 Sol$0.0125$12.500.28×
GPT-5.5$0.00625$6.250.14×
GPT-5.4$0.00375$3.750.08×
GPT-5.3 Codex Spark$0.005$5.000.11×
GPT-5.4 Mini$0.00075$0.750.017×

A thousand calls on GPT-6 Astra cost $45. The same thousand calls on GPT-5.4 Mini cost $0.75 — a 60× difference. The right strategy is to use the smallest model that delivers acceptable quality for each task, and reserve frontier models for the minority of requests that truly need them.

Postpaid vs prepaid billing

OpenAI's standard billing is postpaid: you use the API through the month, and at the end of the billing period you pay for everything you consumed. This model has two practical problems for cost-sensitive teams.

No hard spending ceiling. If an automated agent enters a loop or a batch job runs unexpectedly, your bill can spike with no cap. There is no mechanism to say "stop at $500." You can set usage alerts, but alerts are warnings, not circuit breakers.

Delayed cost visibility. Postpaid billing surfaces costs after the fact. By the time you see the monthly bill, the tokens are already consumed. Teams that rely on postpaid billing often discover overspending two to four weeks after it happens.

Prepaid token packs invert this model. You buy a fixed number of tokens upfront, load them into your account, and spend against that balance. When the balance runs out, requests fail until you top up. There is no mechanism for a surprise bill because spending cannot exceed the prepaid balance. You also get real-time dashboard visibility into per-model consumption, which makes it possible to spot cost trends before they become problems.

Cheaper alternative: APItokendeal token packs

APItokendeal provides OpenAI-compatible API access through prepaid token packs. You get one API key that works with GPT, Claude, DeepSeek, Gemini, and other models — all through a single OpenAI-compatible endpoint. The per-token costs can undercut standard OpenAI postpaid rates because APItokendeal negotiates volume pricing and passes the savings through.

Plan Price Best for
Explorer$9.99First paid experiments
Starter$24.99Experiments, personal tools
Growth$49.99Small production apps
Scale$149.99High-volume workloads
EnterpriseCustomDedicated volume pricing

The comparison is straightforward. OpenAI postpaid billing has no spending ceiling, delayed cost visibility, and requires a separate account per provider. APItokendeal provides a hard spending cap (your prepaid balance), real-time per-model cost tracking in a dashboard, multi-provider access through one key, and immediate provisioning with no sales call required. For teams that want predictable AI costs without vendor lock-in, prepaid token packs are the simpler option.

Need startup-friendly AI API cost control?

Use token plans and one managed API key to keep usage visible from the start.