Dashboard
Model Gateway

Base pricing

Current base rates for supported model families.

These base rates mirror the HALO production pricing catalog checked on September 3, 2026. Text prices are in USD per 1 million tokens unless a provider-specific unit is shown.

#How to read prices

#Google Gemini

ModelInputCached inputOutput
gemini-3.8-flash$0.75$0.075$3.75
gemini-3.7-flash$0.75$0.075$3.75
gemini-3.6-flash$0.75$0.075$3.75
gemini-3.5-flash$1.50$0.15$9.00
gemini-3.5-flash-lite$0.30$0.03$2.50
gemini-3.1-pro-preview$2.00 / $4.00$0.20 / $0.40$12.00 / $18.00
gemini-3.1-flash-lite$0.25$0.025$1.50
gemini-3-flash-preview$0.50$0.05$3.00
gemini-2.5-pro$1.25 / $2.50$0.125 / $0.25$10.00 / $15.00
gemini-2.5-flash$0.30$0.03$2.50
gemini-2.5-flash-lite$0.10$0.01$0.40

Paired Gemini rates apply at up to 200,000 input tokens and above 200,000 input tokens. Gemini 3.8 Flash, Gemini 3.7 Flash and Gemini 3.6 Flash use introductory pricing through December 31, 2026; Google lists $1.50 input, $0.15 cached input, and $7.50 output starting January 1, 2027.

#OpenAI

ModelUncached inputCached inputCache writeOutput
gpt-5.6$4.00 / $8.00$0.40 / $0.80$5.00 / $10.00$20.00 / $30.00
gpt-5.6-sol$4.00 / $8.00$0.40 / $0.80$5.00 / $10.00$20.00 / $30.00
gpt-5.6-terra$2.00 / $4.00$0.20 / $0.40$2.50 / $5.00$12.00 / $18.00
gpt-5.6-luna$0.20 / $0.40$0.02 / $0.04$0.25 / $0.50$1.20 / $1.80
gpt-5.5$5.00 / $10.00$0.50 / $1.00$30.00 / $45.00
gpt-5.4$2.50 / $5.00$0.25 / $0.50$15.00 / $22.50
gpt-5.4-mini$0.75$0.075$4.50
gpt-5.4-nano$0.20$0.02$1.25
gpt-5.2$1.75$0.175$14.00
gpt-5.1$1.25$0.125$10.00
gpt-5$1.25$0.125$10.00
gpt-5-mini$0.25$0.025$2.00
gpt-5-nano$0.05$0.005$0.40
gpt-4.1$2.00$0.50$8.00
gpt-4.1-mini$0.40$0.10$1.60
gpt-4.1-nano$0.10$0.025$0.40
gpt-4o$2.50$1.25$10.00
gpt-4o-mini$0.15$0.075$0.60
o3$2.00$0.50$8.00
o4-mini$1.10$0.275$4.40

OpenAI Chat Completions reports cached input separately. All remaining prompt input is billed at the uncached-input rate. Paired rates apply at up to 272,000 input tokens and above 272,000 input tokens. GPT-5.6 Sol uses OpenAI's current promotional rates.

#OpenAI GPT Image

GPT Image rates are in USD per 1 million tokens and are metered separately by text and image modality.

ModelText inputCached text inputText outputImage inputCached image inputImage output
gpt-image-2$5.00$1.25$8.00$2.00$30.00
gpt-image-1.5$5.00$1.25$10.00$8.00$2.00$32.00
gpt-image-1-mini$2.00$0.20$2.50$0.25$8.00
gpt-image-1$5.00$1.25$10.00$2.50$40.00
chatgpt-image-latest$5.00$1.25$10.00$8.00$2.00$32.00

#Anthropic

ModelInputCached inputCache writeOutput
claude-fable-5-1$10.00$0.25$12.50$50.00
claude-opus-5$5.00$0.50$6.25$25.00
claude-sonnet-5$2.00$0.20$2.50$10.00
claude-fable-5$10.00$1.00$12.50$50.00
claude-opus-4-8$5.00$0.50$6.25$25.00
claude-opus-4-7$5.00$0.50$6.25$25.00
claude-opus-4-6$5.00$0.50$6.25$25.00
claude-opus-4-5-20251101$5.00$0.50$6.25$25.00
claude-sonnet-4-6$3.00$0.30$3.75$15.00
claude-sonnet-4-5-20250929$3.00$0.30$3.75$15.00
claude-sonnet-4-20250514$3.00$0.30$3.75$15.00
claude-haiku-4-5-20251001$1.00$0.10$1.25$5.00
claude-3-5-haiku-20241022$0.80$0.08$1.00$4.00

#DeepSeek

ModelInputCached inputOutput
deepseek-v4-flash$0.44 / $0.22$0.014 / $0.007$1.32 / $0.66
deepseek-v4-pro$1.32 / $0.66$0.044 / $0.022$3.96 / $1.98

Paired DeepSeek rates show peak and off-peak pricing. Peak pricing applies on weekdays from 01:00–04:00 and 06:00–10:00 UTC; all other times use the off-peak rate. HALO snapshots the applicable rate when dispatch starts.

#Fish Audio

Fish Audio TTS is metered by the UTF-8 byte length of the input text, not by language-model tokens. The table shows HALO's 0 bps base price; the selected platform key's buyer price adjustment is applied to the exact request price.

ModelUSD per 1M UTF-8 bytes
s2.1-pro$15.00
s2-pro$15.00
s1$15.00

#Open-source

Open-source rates are HALO catalog rates for the listed runtime pool. Project visibility and healthy provider capacity still apply.

#AllenAI · OLMo

ModelInputCached inputOutput
allenai/olmo-3-32b-think$0.15$0.15$0.50

#DeepSeek

ModelInputCached inputOutput
deepseek/deepseek-v4-flash$0.135883$0.045979$0.266656
deepseek/deepseek-v4-pro$1.310879$0.133455$2.632670

#Google · Gemma

ModelInputCached inputOutput
google/gemma-4-31b-it$0.29$0.272222$0.668889
google/gemma-4-26b-a4b-it$0.11$0.075556$0.368889

#Meta · Llama

ModelInputCached inputOutput
meta-llama/llama-4-maverick$0.284$0.248$0.934
meta-llama/llama-4-scout$0.16$0.146250$0.482500

#MiniMax

ModelInputCached inputOutput
minimax/minimax-m3$0.36$0.168$1.44

#Mistral AI

ModelInputCached inputOutput
mistralai/mistral-small-2603$0.168750$0.101250$0.675
mistralai/mistral-large-2512$0.50$0.05$1.50

#Moonshot AI · Kimi

ModelInputCached inputOutput
moonshotai/kimi-k3$3.00$0.30$15.00
moonshotai/kimi-k2.7-code$0.928744$0.186155$4.059333

#NVIDIA · Nemotron

ModelInputCached inputOutput
nvidia/nemotron-3-nano-30b-a3b$0.055$0.055$0.22
nvidia/nemotron-3-ultra-550b-a55b$0.55$0.15$2.90

#OpenAI · GPT-OSS

ModelInputCached inputOutput
openai/gpt-oss-120b$0.112800$0.104467$0.478333
openai/gpt-oss-20b$0.051167$0.043458$0.190833

#Qwen

ModelInputCached inputOutput
qwen/qwen3.6-35b-a3b$0.176$0.159333$1.160625
qwen/qwen3-coder-next$0.182$0.135200$1.10

#Xiaomi · MiMo

ModelInputCached inputOutput
xiaomi/mimo-v2.5-pro$0.549707$0.062359$1.566080
xiaomi/mimo-v2.5$0.138250$0.021050$0.294

#Z.ai · GLM

ModelInputCached inputOutput
z-ai/glm-5.2$1.354196$0.260500$4.348842

#Media generation

#Imagen · per generated image

ModelPrice
imagen-3.0-generate-002$0.03
imagen-4.0-fast-generate-001$0.02
imagen-4.0-generate-001$0.04
imagen-4.0-ultra-generate-001$0.06

#Veo · USD per generated second

Model720p1080p4K
veo-3.1-lite-generate-preview$0.05$0.08
veo-3.1-fast-generate-preview$0.10$0.12$0.30
veo-3.1-generate-preview$0.40$0.40$0.60
veo-3.0-fast-generate-001$0.10$0.12$0.30
veo-3.0-generate-001$0.40$0.40$0.40
veo-2.0-generate-001$0.35

A dash means that resolution is unavailable and is rejected before provider dispatch.

#Billing notes

  • Batch, Flex, Priority, and regional processing tiers are not listed.
  • Search, cache storage, and modality-specific audio, image, or video token rates can add separate usage.
  • The completed upstream usage and the billing snapshot attached to the request determine the final debit.
  • Use the dashboard Models page for the current project-visible catalog.