Base pricing Current base rates for supported model families.
Copy pageThese base rates mirror the HALO production pricing catalog checked on September 3, 2026. Text prices are in USD per 1 million tokens unless a provider-specific unit is shown.
# How to read pricesUSD per 1M tokens Token rates are metered separately by the usage categories shown in each table. Paired values show standard and long-context rates in that order.
# Google GeminiModel Input Cached input Output gemini-3.8-flash$0.75 $0.075 $3.75 gemini-3.7-flash$0.75 $0.075 $3.75 gemini-3.6-flash$0.75 $0.075 $3.75 gemini-3.5-flash$1.50 $0.15 $9.00 gemini-3.5-flash-lite$0.30 $0.03 $2.50 gemini-3.1-pro-preview$2.00 / $4.00 $0.20 / $0.40 $12.00 / $18.00 gemini-3.1-flash-lite$0.25 $0.025 $1.50 gemini-3-flash-preview$0.50 $0.05 $3.00 gemini-2.5-pro$1.25 / $2.50 $0.125 / $0.25 $10.00 / $15.00 gemini-2.5-flash$0.30 $0.03 $2.50 gemini-2.5-flash-lite$0.10 $0.01 $0.40
Paired Gemini rates apply at up to 200,000 input tokens and above 200,000 input tokens. Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.6 Flash use introductory pricing through December 31, 2026; Google lists $1.50 input, $0.15 cached input, and $7.50 output starting January 1, 2027.
# OpenAIModel Uncached input Cached input Cache write Output gpt-5.6$4.00 / $8.00 $0.40 / $0.80 $5.00 / $10.00 $20.00 / $30.00 gpt-5.6-sol$4.00 / $8.00 $0.40 / $0.80 $5.00 / $10.00 $20.00 / $30.00 gpt-5.6-terra$2.00 / $4.00 $0.20 / $0.40 $2.50 / $5.00 $12.00 / $18.00 gpt-5.6-luna$0.20 / $0.40 $0.02 / $0.04 $0.25 / $0.50 $1.20 / $1.80 gpt-5.5$5.00 / $10.00 $0.50 / $1.00 — $30.00 / $45.00 gpt-5.4$2.50 / $5.00 $0.25 / $0.50 — $15.00 / $22.50 gpt-5.4-mini$0.75 $0.075 — $4.50 gpt-5.4-nano$0.20 $0.02 — $1.25 gpt-5.2$1.75 $0.175 — $14.00 gpt-5.1$1.25 $0.125 — $10.00 gpt-5$1.25 $0.125 — $10.00 gpt-5-mini$0.25 $0.025 — $2.00 gpt-5-nano$0.05 $0.005 — $0.40 gpt-4.1$2.00 $0.50 — $8.00 gpt-4.1-mini$0.40 $0.10 — $1.60 gpt-4.1-nano$0.10 $0.025 — $0.40 gpt-4o$2.50 $1.25 — $10.00 gpt-4o-mini$0.15 $0.075 — $0.60 o3$2.00 $0.50 — $8.00 o4-mini$1.10 $0.275 — $4.40
OpenAI Chat Completions reports cached input separately. All remaining prompt input is billed at the uncached-input rate. Paired rates apply at up to 272,000 input tokens and above 272,000 input tokens. GPT-5.6 Sol uses OpenAI's current promotional rates.
# OpenAI GPT ImageGPT Image rates are in USD per 1 million tokens and are metered separately by text and image modality.
Model Text input Cached text input Text output Image input Cached image input Image output gpt-image-2$5.00 $1.25 — $8.00 $2.00 $30.00 gpt-image-1.5$5.00 $1.25 $10.00 $8.00 $2.00 $32.00 gpt-image-1-mini$2.00 $0.20 — $2.50 $0.25 $8.00 gpt-image-1$5.00 $1.25 — $10.00 $2.50 $40.00 chatgpt-image-latest$5.00 $1.25 $10.00 $8.00 $2.00 $32.00
# AnthropicModel Input Cached input Cache write Output claude-fable-5-1$10.00 $0.25 $12.50 $50.00 claude-opus-5$5.00 $0.50 $6.25 $25.00 claude-sonnet-5$2.00 $0.20 $2.50 $10.00 claude-fable-5$10.00 $1.00 $12.50 $50.00 claude-opus-4-8$5.00 $0.50 $6.25 $25.00 claude-opus-4-7$5.00 $0.50 $6.25 $25.00 claude-opus-4-6$5.00 $0.50 $6.25 $25.00 claude-opus-4-5-20251101$5.00 $0.50 $6.25 $25.00 claude-sonnet-4-6$3.00 $0.30 $3.75 $15.00 claude-sonnet-4-5-20250929$3.00 $0.30 $3.75 $15.00 claude-sonnet-4-20250514$3.00 $0.30 $3.75 $15.00 claude-haiku-4-5-20251001$1.00 $0.10 $1.25 $5.00 claude-3-5-haiku-20241022$0.80 $0.08 $1.00 $4.00
# DeepSeekModel Input Cached input Output deepseek-v4-flash$0.44 / $0.22 $0.014 / $0.007 $1.32 / $0.66 deepseek-v4-pro$1.32 / $0.66 $0.044 / $0.022 $3.96 / $1.98
Paired DeepSeek rates show peak and off-peak pricing. Peak pricing applies on weekdays from 01:00–04:00 and 06:00–10:00 UTC; all other times use the off-peak rate. HALO snapshots the applicable rate when dispatch starts.
# Fish AudioFish Audio TTS is metered by the UTF-8 byte length of the input text, not by language-model tokens. The table shows HALO's 0 bps base price; the selected platform key's buyer price adjustment is applied to the exact request price.
Model USD per 1M UTF-8 bytes s2.1-pro$15.00 s2-pro$15.00 s1$15.00
# Open-sourceOpen-source rates are HALO catalog rates for the listed runtime pool. Project visibility and healthy provider capacity still apply.
# AllenAI · OLMoModel Input Cached input Output allenai/olmo-3-32b-think$0.15 $0.15 $0.50
# DeepSeekModel Input Cached input Output deepseek/deepseek-v4-flash$0.135883 $0.045979 $0.266656 deepseek/deepseek-v4-pro$1.310879 $0.133455 $2.632670
# Google · GemmaModel Input Cached input Output google/gemma-4-31b-it$0.29 $0.272222 $0.668889 google/gemma-4-26b-a4b-it$0.11 $0.075556 $0.368889
# Meta · LlamaModel Input Cached input Output meta-llama/llama-4-maverick$0.284 $0.248 $0.934 meta-llama/llama-4-scout$0.16 $0.146250 $0.482500
# MiniMaxModel Input Cached input Output minimax/minimax-m3$0.36 $0.168 $1.44
# Mistral AIModel Input Cached input Output mistralai/mistral-small-2603$0.168750 $0.101250 $0.675 mistralai/mistral-large-2512$0.50 $0.05 $1.50
# Moonshot AI · KimiModel Input Cached input Output moonshotai/kimi-k3$3.00 $0.30 $15.00 moonshotai/kimi-k2.7-code$0.928744 $0.186155 $4.059333
# NVIDIA · NemotronModel Input Cached input Output nvidia/nemotron-3-nano-30b-a3b$0.055 $0.055 $0.22 nvidia/nemotron-3-ultra-550b-a55b$0.55 $0.15 $2.90
# OpenAI · GPT-OSSModel Input Cached input Output openai/gpt-oss-120b$0.112800 $0.104467 $0.478333 openai/gpt-oss-20b$0.051167 $0.043458 $0.190833
# QwenModel Input Cached input Output qwen/qwen3.6-35b-a3b$0.176 $0.159333 $1.160625 qwen/qwen3-coder-next$0.182 $0.135200 $1.10
# Xiaomi · MiMoModel Input Cached input Output xiaomi/mimo-v2.5-pro$0.549707 $0.062359 $1.566080 xiaomi/mimo-v2.5$0.138250 $0.021050 $0.294
# Z.ai · GLMModel Input Cached input Output z-ai/glm-5.2$1.354196 $0.260500 $4.348842
# Imagen · per generated imageModel Price imagen-3.0-generate-002$0.03 imagen-4.0-fast-generate-001$0.02 imagen-4.0-generate-001$0.04 imagen-4.0-ultra-generate-001$0.06
# Veo · USD per generated secondModel 720p 1080p 4K veo-3.1-lite-generate-preview$0.05 $0.08 — veo-3.1-fast-generate-preview$0.10 $0.12 $0.30 veo-3.1-generate-preview$0.40 $0.40 $0.60 veo-3.0-fast-generate-001$0.10 $0.12 $0.30 veo-3.0-generate-001$0.40 $0.40 $0.40 veo-2.0-generate-001$0.35 — —
A dash means that resolution is unavailable and is rejected before provider dispatch.
# Billing notesBatch, Flex, Priority, and regional processing tiers are not listed. Search, cache storage, and modality-specific audio, image, or video token rates can add separate usage. The completed upstream usage and the billing snapshot attached to the request determine the final debit. Use the dashboard Models page for the current project-visible catalog.