---
title: Model costs
description: Public, validated pricing and capability data for AI model providers (LLMs, STT, TTS).
license: CC-BY-4.0
version: 2
verified: 2026-09-02
generated_at: 2026-09-02T17:21:39.271Z
source: https://github.com/hail-hq/hail/tree/main/costs
---

# Model costs

Public, validated pricing and capability data for AI model providers — large language models, speech-to-text, and text-to-speech. Schema-validated, dual-licensed CC-BY-4.0 for free reuse.

## At a glance

- **Models:** 166 (88 LLM · 43 STT · 35 TTS)
- **Providers:** 25
- **LLM output range:** $0.10 – $180.00 / Mtok
- **STT range:** $0.000667 – $0.0170 / min
- **TTS range:** $4.00 – $160.00 / 1M chars
- **Verified:** 2026-09-02

## For agents

Fetch the raw JSON for programmatic use — this is the source of truth, schema-validated on every PR:

- LLMs: <https://raw.githubusercontent.com/hail-hq/hail/main/costs/llm.json>
- STT: <https://raw.githubusercontent.com/hail-hq/hail/main/costs/stt.json>
- TTS: <https://raw.githubusercontent.com/hail-hq/hail/main/costs/tts.json>

JSON Schemas: <https://github.com/hail-hq/hail/tree/main/costs/schema>

## Large language models

| Provider | Model | Output $/MTok | Input $/MTok | Cached $/MTok | Context | Out cap | Tools | Modalities | Verified | Source |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Anthropic | Claude Opus 4.7 (`claude-opus-4-7`) | $25.00 | $5.00 | $0.5000 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://www.anthropic.com/pricing) |
| Anthropic | Claude Sonnet 4.6 (`claude-sonnet-4-6`) | $15.00 | $3.00 | $0.3000 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://www.anthropic.com/pricing) |
| Anthropic | Claude Fable 5 (`claude-fable-5`) | $50.00 | $10.00 | $1.0000 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Fable 5.1 (`claude-fable-5-1`) | $50.00 | $10.00 | $0.2500 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Opus 4.8 (`claude-opus-4-8`) | $25.00 | $5.00 | $0.5000 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Sonnet 5 (`claude-sonnet-5`) | $10.00 | $2.00 | $0.2000 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Opus 5 (`claude-opus-5`) | $25.00 | $5.00 | $0.5000 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/models) |
| OpenAI | GPT-5 (`gpt-5`) | $10.00 | $1.25 | $0.1250 | 400,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5) |
| Google | Gemini 2.5 Pro (`gemini-2.5-pro`) | $10.00 | $1.25 | $0.1250 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |
| DeepSeek | DeepSeek V4 Flash (`deepseek-v4-flash`) | $1.32 | $0.44 | $0.0140 | 1,048,576 | 384,000 | yes | in: text / out: text | 2026-09-02 | [link](https://api-docs.deepseek.com/quick_start/pricing) |
| DeepSeek | DeepSeek V4 Pro (`deepseek-v4-pro`) | $3.96 | $1.32 | $0.0440 | 1,048,576 | 384,000 | yes | in: text / out: text | 2026-09-02 | [link](https://api-docs.deepseek.com/quick_start/pricing) |
| DeepSeek | DeepSeek V3 (`deepseek-chat`) | $0.28 | $0.14 | $0.0028 | 1,048,576 | 384,000 | yes | in: text / out: text | 2026-09-02 | [link](https://api-docs.deepseek.com/quick_start/pricing) |
| DeepSeek | DeepSeek R1 (`deepseek-reasoner`) | $0.28 | $0.14 | $0.0028 | 1,048,576 | 384,000 | yes | in: text / out: text | 2026-09-02 | [link](https://api-docs.deepseek.com/quick_start/pricing) |
| DeepSeek | DeepSeek V4 Flash Vision (Experimental) (`deepseek-v4-flash-vision-exp`) | $1.32 | $0.44 | $0.0140 | 1,048,576 | 384,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://api-docs.deepseek.com/quick_start/pricing) |
| Anthropic | Claude Haiku 4.5 (`claude-haiku-4-5-20251001`) | $5.00 | $1.00 | $0.1000 | 200,000 | 64,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Sonnet 4.5 (`claude-sonnet-4-5-20250929`) | $15.00 | $3.00 | $0.3000 | 200,000 | 64,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Opus 4.5 (`claude-opus-4-5-20251101`) | $25.00 | $5.00 | $0.5000 | 200,000 | 64,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Sonnet 3.7 (`claude-3-7-sonnet-20250219`) | $15.00 | $3.00 | $0.3000 | 200,000 | 64,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/model-deprecations) |
| Anthropic | Claude Haiku 3.5 (`claude-3-5-haiku-20241022`) | $4.00 | $0.80 | $0.0800 | 200,000 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| OpenAI | GPT-5 mini (`gpt-5-mini`) | $2.00 | $0.25 | $0.0250 | 400,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5-mini) |
| OpenAI | GPT-5 nano (`gpt-5-nano`) | $0.40 | $0.05 | $0.0050 | 400,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5-nano) |
| OpenAI | GPT-4.1 (`gpt-4.1`) | $8.00 | $2.00 | $0.5000 | 1,047,576 | 32,768 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-4.1) |
| OpenAI | GPT-4o (`gpt-4o`) | $10.00 | $2.50 | $1.2500 | 128,000 | 16,384 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-4o) |
| OpenAI | GPT-4o mini (`gpt-4o-mini`) | $0.60 | $0.15 | $0.0750 | 128,000 | 16,384 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-4o-mini) |
| OpenAI | OpenAI o3 (`o3`) | $8.00 | $2.00 | $0.5000 | 200,000 | 100,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/o3) |
| OpenAI | OpenAI o4-mini (`o4-mini`) | $4.40 | $1.10 | $0.2750 | 200,000 | 100,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/o4-mini) |
| OpenAI | OpenAI o1 (`o1`) | $60.00 | $15.00 | $7.5000 | 200,000 | 100,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/o1) |
| OpenAI | GPT-5.5 (`gpt-5.5`) | $30.00 | $5.00 | $0.5000 | 1,050,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.5) |
| OpenAI | GPT-5.4 (`gpt-5.4`) | $15.00 | $2.50 | $0.2500 | 1,050,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.4) |
| OpenAI | GPT-5.4 mini (`gpt-5.4-mini`) | $4.50 | $0.75 | $0.0750 | 400,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.4-mini) |
| OpenAI | GPT-5.4 nano (`gpt-5.4-nano`) | $1.25 | $0.20 | $0.0200 | 400,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.4-nano) |
| OpenAI | GPT-5.6 Sol (`gpt-5.6-sol`) | $20.00 | $4.00 | $0.4000 | 1,050,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.6-sol) |
| OpenAI | GPT-5.6 Terra (`gpt-5.6-terra`) | $12.00 | $2.00 | $0.2000 | 1,050,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.6-terra) |
| OpenAI | GPT-5.6 Luna (`gpt-5.6-luna`) | $1.20 | $0.20 | $0.0200 | 1,050,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.6-luna) |
| OpenAI | GPT-5.5 Pro (`gpt-5.5-pro`) | $180.00 | $30.00 | — | 1,050,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.5-pro) |
| OpenAI | GPT-5.4 Pro (`gpt-5.4-pro`) | $180.00 | $30.00 | — | 1,050,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.4-pro) |
| OpenAI | GPT-5.3-Codex (`gpt-5.3-codex`) | $14.00 | $1.75 | $0.1750 | 400,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-5.3-codex) |
| Google | Gemini 2.5 Flash (`gemini-2.5-flash`) | $2.50 | $0.30 | $0.0300 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/gemini-api/docs/pricing) |
| Google | Gemini 2.5 Flash-Lite (`gemini-2.5-flash-lite`) | $0.40 | $0.10 | $0.0100 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/gemini-api/docs/pricing) |
| Google | Gemini 2.0 Flash (`gemini-2.0-flash`) | $0.40 | $0.10 | $0.0250 | 1,048,576 | 8,192 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/gemini-api/docs/pricing) |
| Google | Gemini 3.5 Flash (`gemini-3.5-flash`) | $9.00 | $1.50 | $0.1500 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |
| Google | Gemini 3.1 Pro Preview (`gemini-3.1-pro-preview`) | $12.00 | $2.00 | $0.2000 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |
| Google | Gemini 3 Flash Preview (`gemini-3-flash-preview`) | $3.00 | $0.50 | $0.0500 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |
| Google | Gemini 3.1 Flash-Lite (`gemini-3.1-flash-lite`) | $1.50 | $0.25 | $0.0250 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |
| Meta | Llama 4 Maverick (`llama-4-maverick`) | $0.85 | $0.27 | — | 1,048,576 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct) |
| Meta | Llama 4 Scout (`llama-4-scout`) | $0.59 | $0.18 | — | 10,485,760 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct) |
| Meta | Llama 3.3 70B Instruct (`llama-3.3-70b`) | $1.04 | $1.04 | — | 131,072 | 8,192 | yes | in: text / out: text | 2026-09-02 | [link](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct) |
| Mistral | Mistral Large 2 (24.11) (`mistral-large-2411`) | $6.00 | $2.00 | — | 131,072 | 8,192 | yes | in: text / out: text | 2026-09-02 | [link](https://mistral.ai/news/mistral-large-2407) |
| Mistral | Mistral Medium 3 (`mistral-medium-2505`) | $2.00 | $0.40 | — | 131,072 | 8,192 | yes | in: text / out: text | 2026-09-02 | [link](https://mistral.ai/news/mistral-medium-3) |
| Mistral | Mistral Small 3 (`mistral-small-2501`) | $0.30 | $0.10 | — | 32,768 | 8,192 | yes | in: text / out: text | 2026-09-02 | [link](https://mistral.ai/news/mistral-small-3) |
| Mistral | Codestral 25.08 (`codestral-2508`) | $0.90 | $0.30 | — | 131,072 | 8,192 | yes | in: text / out: text | 2026-09-02 | [link](https://mistral.ai/news/codestral-25-08) |
| Mistral | Pixtral Large (`pixtral-large-2411`) | $6.00 | $2.00 | — | 131,072 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://mistral.ai/news/pixtral-large) |
| Mistral | Mistral Large 3 (`mistral-large-2512`) | $1.50 | $0.50 | — | 262,144 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://mistral.ai/news/mistral-3) |
| Mistral | Mistral Medium 3.5 (`mistral-medium-3-5`) | $7.50 | $1.50 | — | 262,144 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5) |
| Mistral | Mistral Small 4 (`mistral-small-2603`) | $0.60 | $0.15 | — | 262,144 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03) |
| Mistral | Ministral 3 3B (`ministral-3b-2512`) | $0.10 | $0.10 | — | 262,144 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://mistral.ai/news/mistral-3) |
| Mistral | Ministral 3 8B (`ministral-8b-2512`) | $0.15 | $0.15 | — | 262,144 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://mistral.ai/news/mistral-3) |
| Mistral | Ministral 3 14B (`ministral-14b-2512`) | $0.20 | $0.20 | — | 262,144 | 8,192 | yes | in: text,image / out: text | 2026-09-02 | [link](https://mistral.ai/news/mistral-3) |
| Cohere | Command A (`command-a-03-2025`) | $10.00 | $2.50 | — | 256,000 | 8,000 | yes | in: text / out: text | 2026-09-02 | [link](https://docs.cohere.com/docs/command-a) |
| Cohere | Command R+ (`command-r-plus-08-2024`) | $10.00 | $2.50 | — | 128,000 | 4,000 | yes | in: text / out: text | 2026-09-02 | [link](https://cohere.com/pricing) |
| Cohere | Aya Expanse 32B (`c4ai-aya-expanse-32b`) | $1.50 | $0.50 | — | 128,000 | 4,000 | no | in: text / out: text | 2026-09-02 | [link](https://cohere.com/pricing) |
| xAI | Grok 4 (`grok-4-0709`) | $15.00 | $3.00 | $0.7500 | 256,000 | 256,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/developers/migration/may-15-retirement) |
| xAI | Grok 3 (`grok-3`) | $15.00 | $3.00 | $0.7500 | 131,072 | 131,072 | yes | in: text / out: text | 2026-09-02 | [link](https://docs.x.ai/developers/migration/may-15-retirement) |
| xAI | Grok Code Fast 1 (`grok-code-fast-1`) | $1.50 | $0.20 | $0.0200 | 256,000 | 256,000 | yes | in: text / out: text | 2026-09-02 | [link](https://docs.x.ai/developers/migration/may-15-retirement) |
| xAI | Grok 4.3 (`grok-4.3`) | $2.50 | $1.25 | $0.2000 | 1,000,000 | 1,000,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/docs/models) |
| xAI | Grok 4.5 (`grok-4.5`) | $6.00 | $2.00 | $0.3000 | 500,000 | 500,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/docs/models) |
| xAI | Grok 4.6 (`grok-4.6`) | $6.00 | $2.00 | $0.5000 | 500,000 | 500,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/docs/models) |
| xAI | Grok Build 0.1 (`grok-build-0.1`) | $2.00 | $1.00 | $0.2000 | 256,000 | 256,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/docs/models) |
| xAI | Grok 4.20 Reasoning (`grok-4.20-0309-reasoning`) | $2.50 | $1.25 | $0.2000 | 1,000,000 | 1,000,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/docs/models) |
| xAI | Grok 4.20 Non-Reasoning (`grok-4.20-0309-non-reasoning`) | $2.50 | $1.25 | $0.2000 | 1,000,000 | 1,000,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/docs/models) |
| xAI | Grok 4.20 Multi-Agent (`grok-4.20-multi-agent-0309`) | $2.50 | $1.25 | $0.2000 | 1,000,000 | 1,000,000 | no | in: text,image / out: text | 2026-09-02 | [link](https://docs.x.ai/docs/models) |
| Alibaba | Qwen3-Max (`qwen3-max`) | $6.00 | $1.20 | — | 262,144 | 32,768 | yes | in: text / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |
| Alibaba | Qwen3-Coder-Plus (`qwen3-coder-plus`) | $5.00 | $1.00 | — | 1,048,576 | 65,536 | yes | in: text / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |
| Alibaba | Qwen3.7-Max (`qwen3.7-max`) | $7.50 | $2.50 | — | 1,048,576 | 65,536 | yes | in: text / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/model-pricing) |
| Alibaba | Qwen3.7-Plus (`qwen3.7-plus`) | $1.60 | $0.40 | — | 1,048,576 | 65,536 | yes | in: text,image / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/billing) |
| Alibaba | Qwen3.6-Flash (`qwen3.6-flash`) | $1.50 | $0.25 | — | 1,048,576 | 65,536 | yes | in: text,image,video / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/billing) |
| Alibaba | Qwen3.8-Max (`qwen3.8-max`) | $6.00 | $2.00 | — | 1,048,576 | 131,072 | yes | in: text,image,video / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/billing) |
| Alibaba | Qwen3.8-Flash (`qwen3.8-flash`) | $0.47 | $0.15 | — | 1,048,576 | 131,072 | yes | in: text,image,video / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/billing) |
| Alibaba | Qwen3.6-27B (`qwen3.6-27b`) | $3.60 | $0.60 | — | 262,144 | 65,536 | yes | in: text,image,video / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/billing) |
| Alibaba | Qwen3.8-27B (`qwen3.8-27b`) | $3.00 | $0.50 | — | 1,048,576 | 131,072 | yes | in: text,image,video / out: text | 2026-09-02 | [link](https://www.alibabacloud.com/help/en/model-studio/billing) |
| Perplexity | Sonar (`sonar`) | $1.00 | $1.00 | — | 128,000 | 8,000 | no | in: text / out: text | 2026-09-02 | [link](https://docs.perplexity.ai/getting-started/pricing) |
| Perplexity | Sonar Pro (`sonar-pro`) | $15.00 | $3.00 | — | 200,000 | 8,000 | no | in: text / out: text | 2026-09-02 | [link](https://docs.perplexity.ai/getting-started/pricing) |
| Perplexity | Sonar Reasoning Pro (`sonar-reasoning-pro`) | $8.00 | $2.00 | — | 128,000 | 8,000 | no | in: text / out: text | 2026-09-02 | [link](https://docs.perplexity.ai/getting-started/pricing) |
| Anthropic | Claude Opus 4.6 (`claude-opus-4-6`) | $25.00 | $5.00 | $0.5000 | 1,000,000 | 128,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Anthropic | Claude Opus 4.1 (`claude-opus-4-1-20250805`) | $75.00 | $15.00 | $1.5000 | 200,000 | 32,000 | yes | in: text,image / out: text | 2026-09-02 | [link](https://platform.claude.com/docs/en/about-claude/pricing) |
| Google | Gemini 3.6 Flash (`gemini-3.6-flash`) | $3.75 | $0.75 | $0.0750 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |
| Google | Gemini 3.5 Flash-Lite (`gemini-3.5-flash-lite`) | $2.50 | $0.30 | — | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |
| Google | Gemini 3.7 Flash (`gemini-3.7-flash`) | $3.75 | $0.75 | $0.0750 | 1,048,576 | 65,536 | yes | in: text,image,audio,video / out: text | 2026-09-02 | [link](https://ai.google.dev/pricing) |

**Notes:**

- **Anthropic Claude Opus 4.7** — Cache hit pricing is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x base ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. Knowledge cutoff Jan 2026. Uses a new tokenizer vs prior Claude models (may use up to 35% more tokens for identical text). Supports adaptive thinking (no extended-thinking toggle); thinking output tokens are billed at the output rate. Legacy listing (superseded by Claude Opus 4.8 as of 2026-05-28) but still fully active with unchanged pricing.
- **Anthropic Claude Sonnet 4.6** — Cache hit pricing is 0.1x base input ($0.30/MTok); 5-minute cache write is 1.25x base ($3.75/MTok); 1-hour cache write is 2x ($6/MTok). Batch API discounts both input and output by 50%. Knowledge cutoff Aug 2025. Supports extended thinking and adaptive thinking; thinking output tokens are billed at the output rate. Corrected max_output_tokens to 128000 (was incorrectly 64000) per platform.claude.com/docs/en/docs/about-claude/models; not a price field, so last_changed_at is not bumped. Still active; legacy listing alongside Claude Sonnet 5 (launched 2026-06-30).
- **Anthropic Claude Fable 5** — Anthropic's most capable widely released model as of its GA on 2026-06-09; now a legacy listing (superseded by Claude Fable 5.1 as of 2026-09-01) but still fully active with unchanged pricing. GA on the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud, and Microsoft Foundry (Foundry has no deployment_options enum value, so omitted). Cache hit is 0.1x base input ($1.00/MTok); 5-minute cache write is 1.25x ($12.50/MTok); 1-hour cache write is 2x ($20/MTok). Batch API discounts both input and output by 50%. Adaptive thinking is always on (thinking:{type:'disabled'} returns 400); raw chain of thought is never returned (thinking.display defaults to 'omitted'). Safety classifiers may decline requests (stop_reason:'refusal'). Requires 30-day data retention; not available under zero data retention. Uses the same tokenizer as Opus 4.7/4.8 (~30% more tokens than pre-4.7 models for identical text). Access to Fable 5 was briefly suspended and restored on 2026-07-01 (no price change). Claude Mythos 5 (claude-mythos-5) shares identical specs and pricing but is invitation-only via Project Glasswing (not self-serve) — deferred, not added as a separate row.
- **Anthropic Claude Fable 5.1** — New row this refresh. Released 2026-09-01 as the latest, recommended model for demanding reasoning and long-horizon agentic work (Anthropic recommends starting with Opus 5 for most workloads; use Fable 5.1 when Opus 5 at higher effort falls short). Extends Claude Fable 5 at the same base input/output, 5-minute cache write, 1-hour cache write, and batch prices, but cache reads (hits/refreshes) are priced at 0.025x base input ($0.25/MTok) instead of the standard 0.1x multiplier used by every other current Claude model — this is the one structured price field that differs from Fable 5, hence the new row rather than an in-place update. 1M context window, 128k max output. Adaptive thinking always on; forced tool use (tool_choice any/tool) returns an error (breaking change vs Fable 5); thinking blocks are tied to the model that produced them and are invalidated by editing earlier turns. Available on Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud, and Microsoft Foundry. Uses the same tokenizer as Fable 5/Opus 4.7+ (~30% more tokens than pre-4.7 models for identical text). Reliable knowledge cutoff and training data cutoff both Jun 2026. Retirement commitment: not sooner than 2027-09-01. Claude Mythos 5.1 (claude-mythos-5-1) shares identical specs and pricing but is invitation-only via Project Glasswing (not self-serve) — deferred, not added as a separate row.
- **Anthropic Claude Opus 4.8** — Anthropic's most capable generally available model, launched 2026-05-28 as the successor to Opus 4.7; still fully active with unchanged pricing. Same base pricing as Opus 4.7. Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. 1M context window at standard pricing (no long-context premium), 128k max output. Adaptive thinking supported; effort parameter defaults to 'high' on all surfaces. Fast mode available as a research preview (Claude API only) at $10/$50 per MTok input/output. Knowledge cutoff Jan 2026.
- **Anthropic Claude Sonnet 5** — Launched 2026-06-30, next-generation Sonnet. The $2/$10 per MTok input/output pricing (cache hit $0.20, 5-min cache write $2.50, 1-hour cache write $4, batch $1.00/$5.00) was announced at launch as introductory pricing through 2026-08-31; Anthropic's pricing page now states this is the standard price and the previously scheduled increase to $3/$15 on 2026-09-01 will not occur. No structured price field changed this refresh, so last_changed_at is not bumped. 1M context window, 128k max output. Adaptive thinking on by default (omitting thinking runs adaptive); manual extended thinking (budget_tokens) removed and returns 400; non-default temperature/top_p/top_k also return 400. New tokenizer produces ~30% more tokens than Sonnet 4.6 for identical text. Not available with Priority Tier. Knowledge cutoff Jan 2026.
- **Anthropic Claude Opus 5** — Launched 2026-06-09, recommended for complex agentic coding and enterprise work; still active with unchanged pricing. 1M context window, 128k max output. Pricing: $5/$25 per MTok base input/output. Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50% ($2.50/$12.50 per MTok). Adaptive thinking on by default (omitting thinking runs adaptive); manual extended thinking not supported (returns 400); non-default temperature/top_p/top_k also return 400. Uses the new tokenizer as of Opus 4.7/4.8 (~30% more tokens than pre-4.7 models for identical text). Available on Claude API, Amazon Bedrock, Google Cloud Vertex AI, Claude Platform on AWS, and Microsoft Foundry. Knowledge cutoff May 2026 (training data cutoff May 2026).
- **OpenAI GPT-5** — Reasoning model with adjustable reasoning_effort; reasoning tokens are billed at the output rate. Context window 400k, max output 128k confirmed against developers.openai.com/api/docs/models/gpt-5. Cached input at 10% of base ($0.125/MTok). Batch API at flat 50% off input and output. PDF input via the Files API; image input native; audio is NOT supported on this model_id. Knowledge cutoff Sept 2024. Deprecation announced 2026-06-11: dated snapshot gpt-5-2025-08-07 scheduled for API removal 2026-12-11, recommended replacement gpt-5.5 (per developers.openai.com/api/docs/deprecations); bare 'gpt-5' still serving as of 2026-07-19, prices unchanged. 2026-07-13 re-verify: model card prose now points to the newer GPT-5.6 family ('We recommend using the latest GPT-5.6') as of GPT-5.6's release, but the structured deprecations table still redirects the dated gpt-5-2025-08-07 snapshot to gpt-5.5; kept replaced_by_model_id at gpt-5.5 pending an update to the formal deprecation entry. 2026-07-19 re-verify: model card live with prices/specs unchanged and deprecations table unchanged; note this model no longer appears on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices/specs unchanged (back on the main pricing table this pass, matching the model card); the deprecations table now redirects gpt-5-2025-08-07 to gpt-5.6-sol (GPT-5.6 family's flagship, launched since the last pass) instead of gpt-5.5 — replaced_by_model_id updated to match.
- **Google Gemini 2.5 Pro** — Input pricing shown is for prompts <=200k tokens; prompts >200k tokens are billed per pricing_tiers; cached input also tiers at 200k (the >200k rate is captured in pricing_tiers[0].cache_read_per_mtok_usd). Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-2.5-pro. Thinking is always on and cannot be disabled; thinking tokens are billed at the output rate. Knowledge cutoff January 2025. Batch Mode discount is a flat 50% off input/output. Audio input is billed at the standard input rate of $1.25/MTok (no separate audio premium, unlike 2.5 Flash/Flash-Lite/2.0 Flash); audio_input_per_mtok_usd omitted. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. PDF input via the Files API; image, audio, and video native. Re-verified 2026-07-02: prices, context window, and modalities unchanged; no new pricing tiers. Re-verified 2026-07-13: prices, tiers, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices, tiers, and modalities unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **DeepSeek DeepSeek V4 Flash** — MoE architecture: 284B total parameters, 13B activated. Replaces deepseek-chat (V3-era alias). Supports both non-thinking and thinking (default) modes; thinking output is billed at the same output rate, so reasoning_tokens_billed: true. deepseek-chat and deepseek-reasoner legacy aliases were discontinued 2026-07-24. Re-verified 2026-07-02 through 2026-08-11: prices unchanged; snapshot updated to DeepSeek-V4-Flash-0731 on 2026-08-11 check (no pricing impact then). Re-verified 2026-09-02: DeepSeek introduced peak/off-peak pricing effective 2026-08-16 (announced api-docs.deepseek.com/updates changelog entry dated 2026-08-13, filed under the V4 Pro GA release), realizing the price increase flagged as forthcoming in the 2026-08-11 refresh. Values recorded here are PEAK rates (list price): input (cache miss) $0.44, output $1.32, cache-hit input $0.014 per MTok — confirmed against the raw pricing table at api-docs.deepseek.com/quick_start/pricing. Off-peak rates are exactly half: cache hit $0.007, cache miss $0.22, output $0.66 per MTok, applying during peak-labeled hours' complement — i.e. off-peak is all hours EXCEPT 01:00-04:00 and 06:00-10:00 UTC, Monday-Friday (peak hours are the pricier window despite the off-peak/peak naming referring to demand, not price; peak hours are ~21% of the week). Schema has no time-of-day pricing axis, so off-peak is not separately encoded; use peak here as the conservative list price. Cache-hit input is now ~1/31 of cache-miss input ($0.014 / $0.44), not the previous 1/50 ratio. Cross-checked api-docs.deepseek.com/news/news260821 and api-docs.deepseek.com/updates: same-day (2026-08-21) launch of deepseek-v4-flash-vision-exp (new sibling row, vision-only variant); no other changes to this row's model_id, snapshot, context window, or concurrency limit (2500).
- **DeepSeek DeepSeek V4 Pro** — Pro-tier sibling to deepseek-v4-flash, positioned for higher-quality responses at lower concurrency (500 vs Flash's 2500). The permanent post-promo rate ($0.435 / $0.87 / $0.003625 cache hit) held from 2026-05-22 through the 2026-08-11 refresh. Re-verified 2026-09-02: DeepSeek-V4-Pro reached GA on 2026-08-13 (snapshot DeepSeek-V4-Pro-0813, previously unrecorded) and introduced peak/off-peak pricing effective 2026-08-16, per the changelog at api-docs.deepseek.com/updates. Values recorded here are PEAK rates (list price): input (cache miss) $1.32, output $3.96, cache-hit input $0.044 per MTok — confirmed against the raw pricing table at api-docs.deepseek.com/quick_start/pricing; exactly 3x the prior permanent rate. Off-peak rates are half: cache hit $0.022, cache miss $0.66, output $1.98 per MTok, applying all hours EXCEPT 01:00-04:00 and 06:00-10:00 UTC Monday-Friday (peak hours, ~21% of the week, are the pricier window). Schema has no time-of-day pricing axis, so off-peak is not separately encoded; peak used here as the conservative list price. Cache-hit input is now 1/30 of cache-miss input, versus 1/120 previously. GA also brought native Responses API support (previously flagged as coming early August, now live) and three thinking-effort levels (low/high/max), no further pricing impact. Cache miss continues to be billed at the base input rate; thinking output billed at the output rate, so reasoning_tokens_billed: true. Cross-checked api-docs.deepseek.com/news/news260821, no changes specific to this row (that entry covers the separate deepseek-v4-flash-vision-exp launch).
- **DeepSeek DeepSeek V3** — Canonical API model_id for the DeepSeek V3 lineage (V3 launched 2024-12-26 as deepseek-chat; upgraded through V3-0324, V3.1, V3.1-Terminus, V3.2 by 2025-12-01). Deprecated 2026-04-24 when V4 launched; still callable until scheduled discontinuation 2026-07-24, currently routing to deepseek-v4-flash non-thinking mode (prices captured here reflect that routing, pre-2026-08-16 peak/off-peak rate change). DeepSeek's pricing page no longer publishes V3-era historical rates; standalone deepseek-v3 model_id was never exposed by the API. Re-verified 2026-07-02: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged. Re-verified 2026-07-13: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged; now 11 days out, expect this row to flip from deprecated to fully retired at the next refresh. Re-verified 2026-07-19: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged against api-docs.deepseek.com/quick_start/pricing; now 5 days out. Re-verified 2026-07-27: model past discontinuation date (2026-07-24 15:59 UTC); no longer available via API. Kept for backward compatibility and referential resolution (replaced_by_model_id: deepseek-v4-flash). Re-verified 2026-08-11: still absent from api-docs.deepseek.com/quick_start/pricing (fully retired, confirms 2026-07-24 discontinuation held); no reversal. Re-verified 2026-09-02: still absent from api-docs.deepseek.com/quick_start/pricing and api-docs.deepseek.com/news/news260821; no reactivation. Prices frozen here at the pre-retirement snapshot, not restated for the 2026-08-16 peak/off-peak change since the model_id is no longer callable.
- **DeepSeek DeepSeek R1** — Canonical API model_id for the DeepSeek R1 reasoning lineage (R1 launched 2025-01-20 as deepseek-reasoner; R1-0528 update 2025-05-28). Deprecated 2026-04-24 when V4 launched; still callable until scheduled discontinuation 2026-07-24, currently routing to deepseek-v4-flash thinking mode (prices captured here reflect that routing, pre-2026-08-16 peak/off-peak rate change). Reasoning output is billed at the standard output rate (reasoning_tokens_billed: true). DeepSeek's pricing page no longer publishes R1-era historical rates; standalone deepseek-r1 model_id was never exposed by the API. Re-verified 2026-07-02: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged. Re-verified 2026-07-13: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged; now 11 days out, expect this row to flip from deprecated to fully retired at the next refresh. Re-verified 2026-07-19: discontinuation date (2026-07-24 15:59 UTC) and routing/pricing unchanged against api-docs.deepseek.com/quick_start/pricing; now 5 days out. Re-verified 2026-07-27: model past discontinuation date (2026-07-24 15:59 UTC); no longer available via API. Kept for backward compatibility and referential resolution (replaced_by_model_id: deepseek-v4-flash). Re-verified 2026-08-11: still absent from api-docs.deepseek.com/quick_start/pricing (fully retired, confirms 2026-07-24 discontinuation held); no reversal. Re-verified 2026-09-02: still absent from api-docs.deepseek.com/quick_start/pricing and api-docs.deepseek.com/news/news260821; no reactivation. Prices frozen here at the pre-retirement snapshot, not restated for the 2026-08-16 peak/off-peak change since the model_id is no longer callable.
- **DeepSeek DeepSeek V4 Flash Vision (Experimental)** — New row, added at the 2026-09-02 refresh. Launched 2026-08-21 (model version DeepSeek-V4-Flash-Vision-Exp) as an experimental vision-capable sibling to deepseek-v4-flash; matches V4 Flash on text capabilities (agents, reasoning, world knowledge) while adding multimodal input, targeted at agent benchmarks requiring visual understanding. Same peak-rate pricing as deepseek-v4-flash (identical MoE architecture): input (cache miss) $0.44, output $1.32, cache-hit input $0.014 per MTok; off-peak rates are half (cache hit $0.007, cache miss $0.22, output $0.66), applying all hours except 01:00-04:00 and 06:00-10:00 UTC Monday-Friday (peak hours, ~21% of the week) — schema has no time-of-day pricing axis so only the peak/list rate is encoded. Images are tokenized for billing at up to 384 tokens each per the launch announcement, using the standard V4 Flash input rate (no separate image-pricing field in schema; folded into input_per_mtok_usd). Supports both non-thinking and thinking (default) modes; thinking output billed at the output rate, so reasoning_tokens_billed: true. FIM completion is not supported (unlike plain deepseek-v4-flash, which supports it non-thinking-mode only). Concurrency limit 2500, matching deepseek-v4-flash. Works with Chat Completions, Anthropic-format, and Responses APIs; accepts base64, URL, or Files-API image references. Two-source verification: primary pricing table at api-docs.deepseek.com/quick_start/pricing plus the launch announcement at api-docs.deepseek.com/news/news260821 and the changelog at api-docs.deepseek.com/updates (both dated 2026-08-21).
- **Anthropic Claude Haiku 4.5** — Anthropic's fastest model with near-frontier intelligence; positioned for high-volume agentic workloads. Cache hit is 0.1x base input ($0.10/MTok); 5-minute cache write is 1.25x ($1.25/MTok); 1-hour cache write is 2x ($2/MTok). Batch API discounts both input and output by 50%. Supports extended thinking; thinking output tokens are billed at the output rate. Reliable knowledge cutoff Feb 2025; training data cutoff Jul 2025.
- **Anthropic Claude Sonnet 4.5** — Legacy listing in Anthropic's models overview but still active. Pricing identical to Sonnet 4.6, but 200k context window (vs 1M on Sonnet 4.6). Cache hit is 0.1x base input ($0.30/MTok); 5-minute cache write is 1.25x ($3.75/MTok); 1-hour cache write is 2x ($6/MTok). Batch API discounts both input and output by 50%. Supports extended thinking; thinking output tokens are billed at the output rate. Reliable knowledge cutoff Jan 2025.
- **Anthropic Claude Opus 4.5** — Legacy listing in Anthropic's models overview but still active. Pricing identical to Opus 4.6/4.7, but 200k context window (vs 1M on 4.6/4.7) and 64k max output (vs 128k). Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. Supports extended thinking; thinking output tokens are billed at the output rate. Reliable knowledge cutoff May 2025.
- **Anthropic Claude Sonnet 3.7** — RETIRED on the Claude API on 2026-02-19; still available on Amazon Bedrock and Google Vertex AI under partner retirement schedules. Anthropic's first reasoning model with extended thinking; can output up to 64k tokens in thinking mode (128k with the output-128k-2025-02-19 beta header). Prices are no longer listed on Anthropic's current pricing page; values sourced from OpenRouter and pricepertoken.com (confidence: medium). Cache and batch pricing inferred from Anthropic's standard multipliers (1.25x 5-min write, 0.1x cache read, 0.5x batch).
- **Anthropic Claude Haiku 3.5** — RETIRED on the Claude API on 2026-02-19; still listed on Anthropic's pricing page as available on Amazon Bedrock and Google Vertex AI only. No extended-thinking support. max_output_tokens=8192 sourced from Anthropic legacy model card and OpenRouter (confidence: medium — not present in current docs). Cache hit is 0.1x base input ($0.08/MTok); 5-minute cache write is 1.25x ($1/MTok); 1-hour cache write is 2x ($1.60/MTok). Batch API discounts both input and output by 50%.
- **OpenAI GPT-5 mini** — Faster, more cost-efficient GPT-5 variant for low-latency, high-volume workloads. Cached input at 10% of base ($0.025/MTok). Batch API at flat 50% off input and output. Reasoning model with adjustable reasoning_effort; reasoning tokens are billed at the output rate. PDF input via the Files API; image input native. Knowledge cutoff May 2024. Deprecation announced 2026-06-11: dated snapshot gpt-5-mini-2025-08-07 scheduled for API removal 2026-12-11, recommended replacement gpt-5.4-mini (per developers.openai.com/api/docs/deprecations); bare 'gpt-5-mini' still serving as of 2026-07-19, prices unchanged. 2026-07-13 re-verify: model card now also points new low-latency workloads to GPT-5.6 Terra, but bare gpt-5-mini remains active. 2026-07-19 re-verify: model card live with prices/specs unchanged; no longer listed on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices/specs unchanged against the model card; the deprecations table now redirects gpt-5-mini-2025-08-07 to gpt-5.6-terra (GPT-5.6 family's mid tier) instead of gpt-5.4-mini — replaced_by_model_id updated to match.
- **OpenAI GPT-5 nano** — Smallest GPT-5 variant; OpenAI model card lists 'Reasoning model: No' with 'Average' reasoning capability — reasoning_tokens_billed set to false on that basis (confidence: medium because other GPT-5 family members are reasoning models). Cached input at 10% of base ($0.005/MTok). Batch API at flat 50% off input and output. Knowledge cutoff May 2024. Deprecation announced 2026-06-11: dated snapshot gpt-5-nano-2025-08-07 scheduled for API removal 2026-12-11, recommended replacement gpt-5.4-nano (per developers.openai.com/api/docs/deprecations); bare 'gpt-5-nano' still serving as of 2026-07-19, prices unchanged. 2026-07-13 re-verify: model card now also points new speed/cost-sensitive workloads to GPT-5.6 Luna, but bare gpt-5-nano remains active. 2026-07-19 re-verify: model card live with prices/specs unchanged; no longer listed on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices/specs unchanged against the model card; the deprecations table now redirects gpt-5-nano-2025-08-07 to gpt-5.6-luna (GPT-5.6 family's cheapest tier) instead of gpt-5.4-nano — replaced_by_model_id updated to match.
- **OpenAI GPT-4.1** — Non-reasoning flagship with a ~1M-token context window (1,047,576). Cached input at 25% of base ($0.50/MTok). Batch API at flat 50% off input and output. PDF input via the Files API; image input native. Knowledge cutoff June 2024. Not on the June 2026 deprecation announcement; remains active alongside the new GPT-5.4/5.5 family. Re-verified 2026-07-13: prices unchanged, still not on the deprecations table; remains active alongside the new GPT-5.6 family. Re-verified 2026-07-19: model card live with prices/specs unchanged, still not on the deprecations table; no longer listed on the main pricing page (prices verified against the model card). Re-verified 2026-08-11: prices unchanged against the main pricing table ($2.00/$0.50/$8.00, batch $1.00/$4.00), still not on the deprecations table.
- **OpenAI GPT-4o** — Deprecated 2026-04-22 (dated alias gpt-4o-2024-05-13 scheduled for shutdown 2026-10-23 per OpenAI deprecations page); still serving on the API as of 2026-07-19. 2026-07-19 re-verify: model card live with prices/specs unchanged, deprecations table unchanged (gpt-4o-2024-05-13 -> gpt-5.5); no longer listed on the main pricing page (prices verified against the model card). Audio input/output are NOT supported on this model_id — they live on a sibling gpt-4o-audio-preview model card with separate pricing (audio input $40/MTok, audio output $80/MTok); text-mode prices captured here. Cached input at 50% of base ($1.25/MTok). Batch API at flat 50% off input and output. Knowledge cutoff Oct 2023. 2026-07-13 re-verify: the deprecations table now redirects the dated gpt-4o-2024-05-13 snapshot to gpt-5.5 (previously gpt-4.1); replaced_by_model_id updated to match the current table verbatim ('gpt-4o-2024-05-13 | gpt-5.5'). 2026-08-11 re-verify: prices/specs unchanged against the main pricing table; the deprecations table now redirects gpt-4o-2024-05-13 to gpt-5.6-sol instead of gpt-5.5 — replaced_by_model_id updated to match.
- **OpenAI GPT-4o mini** — Not on the April 2026 deprecation list; remains active. Re-verified 2026-07-19: model card live with prices/specs unchanged, still not on the deprecations table; no longer listed on the main pricing page (prices verified against the model card). Audio input/output are NOT supported on this model_id — they live on a sibling gpt-4o-mini-audio-preview model card; text-mode prices captured here. Cached input at 50% of base ($0.075/MTok). Batch API at flat 50% off input and output. Knowledge cutoff Oct 2023. Not on the June 2026 deprecation announcement either; remains active alongside the new GPT-5.4/5.5 family. Re-verified 2026-07-13: prices unchanged, listed as 'Default' tier, still not on the deprecations table. Re-verified 2026-08-11: prices unchanged against the main pricing table ($0.15/$0.075/$0.60, batch $0.075/$0.30), still not on the deprecations table.
- **OpenAI OpenAI o3** — Reasoning model for complex tasks; reasoning tokens are billed at the output rate. Cached input at 25% of base ($0.50/MTok). Batch API at flat 50% off input and output. Knowledge cutoff June 2024. Deprecation announced 2026-06-11: dated snapshot o3-2025-04-16 scheduled for API removal 2026-12-11, recommended replacement gpt-5.5 (per developers.openai.com/api/docs/deprecations); bare 'o3' still serving as of 2026-07-19, prices unchanged. 2026-07-19 re-verify: model card live with prices (incl. batch $1.00/$4.00) and specs unchanged, deprecations table unchanged; no longer listed on the main pricing page (prices verified against the model card). 2026-08-11 re-verify: prices unchanged against the main pricing table; the deprecations table now redirects o3-2025-04-16 to gpt-5.6-sol instead of gpt-5.5 — replaced_by_model_id updated to match.
- **OpenAI OpenAI o4-mini** — Fast, cost-efficient reasoning model; reasoning tokens are billed at the output rate. Deprecated 2026-04-22 (dated alias o4-mini-2025-04-16 scheduled for shutdown 2026-10-23 per OpenAI deprecations page); still serving as of 2026-07-19. 2026-07-19 re-verify: model card live with prices/specs unchanged, deprecations table unchanged (o4-mini-2025-04-16 -> gpt-5.4-mini); no longer listed on the main pricing page (prices verified against the model card). OpenAI's model card prose notes 'succeeded by GPT-5 mini'. Cached input at 25% of base ($0.275/MTok). Batch API at flat 50% off input and output. Knowledge cutoff June 2024. 2026-07-13 re-verify: the structured deprecations table lists the dated o4-mini-2025-04-16 snapshot's recommended replacement as gpt-5.4-mini (not gpt-5-mini); replaced_by_model_id updated to match that table verbatim. 2026-08-11 re-verify: prices unchanged; the deprecations table now redirects o4-mini-2025-04-16 to gpt-5.6-terra instead of gpt-5.4-mini — replaced_by_model_id updated to match.
- **OpenAI OpenAI o1** — First-generation reasoning model; reasoning tokens are billed at the output rate. Deprecated 2026-04-22 (dated alias o1-2024-12-17 scheduled for shutdown 2026-10-23 per OpenAI deprecations page); still serving as of 2026-07-19. 2026-07-19 re-verify: model card live with prices/specs unchanged, deprecations table unchanged (o1-2024-12-17 -> gpt-5.5); no longer listed on the main pricing page (prices verified against the model card). OpenAI's recommended replacement is gpt-5.5, which is now in-file; replaced_by_model_id updated from the interim 'o3' successor (o3 itself was deprecated 2026-06-11). Cached input at 50% of base ($7.50/MTok). Batch API at flat 50% off input and output. Knowledge cutoff Oct 2023. Re-verified 2026-07-13: deprecations table confirms shutdown date 2026-10-23 and replacement gpt-5.5 unchanged. 2026-08-11 re-verify: shutdown date 2026-10-23 unchanged, prices unchanged; deprecations table now redirects o1-2024-12-17 to gpt-5.6-sol instead of gpt-5.5 — replaced_by_model_id updated to match.
- **OpenAI GPT-5.5** — Re-verified 2026-07-19 against the main pricing page: standard ($5.00/$0.50/$30.00) and batch ($2.50/$15.00) prices unchanged. OpenAI frontier reasoning model, released 2025-12-01; replaces gpt-5 and o3 as the recommended default per the 2026-06-11 deprecations announcement. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Context window 1,050,000 (Azure Foundry lists 922k input / 128k output split within that total). Cached input at 10% of base ($0.50/MTok). Batch API at flat 50% off input and output. Prompts over 272k input tokens incur a 2x input / 1.5x output surcharge (not modeled as a separate pricing tier here). Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: unlike gpt-5's page, the modality table on this model's docs page lists only Text/Image/Audio/Video rows with no PDF/file callout — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($5.00/$0.50/$30.00) and batch ($2.50/$15.00) prices unchanged against the main pricing table; no longer OpenAI's top recommendation on the models overview page (superseded there by the GPT-5.6 family) but still active with no deprecation entry.
- **OpenAI GPT-5.4** — Re-verified 2026-07-19 against the main pricing page: standard ($2.50/$0.25/$15.00) and batch ($1.25/$7.50) prices unchanged. Default GPT-5.4-class frontier model for professional work, snapshot gpt-5.4-2026-03-05. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Cached input at 10% of base ($0.25/MTok). Batch API at flat 50% off input and output. Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: this model's modality table explicitly lists Audio and Video as not supported with no PDF/file row — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($2.50/$0.25/$15.00) and batch ($1.25/$7.50) prices unchanged against the main pricing table; remains active alongside GPT-5.6, no deprecation entry.
- **OpenAI GPT-5.4 mini** — Re-verified 2026-07-19 against the main pricing page: standard ($0.75/$0.075/$4.50) and batch ($0.375/$2.25) prices unchanged. Default mini model for low-latency, high-volume workloads; replaces gpt-5-mini per the 2026-06-11 deprecations announcement. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Cached input at 10% of base ($0.075/MTok). Batch API at flat 50% off input and output. Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: modality table lists no PDF/file row — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($0.75/$0.075/$4.50) and batch ($0.375/$2.25) prices unchanged against the main pricing table; remains active alongside GPT-5.6, no deprecation entry.
- **OpenAI GPT-5.4 nano** — Re-verified 2026-07-19 against the main pricing page: standard ($0.20/$0.02/$1.25) and batch ($0.10/$0.625) prices unchanged. Cheapest GPT-5.4-class model for simple, high-volume tasks (classification, extraction, ranking, sub-agents); replaces gpt-5-nano per the 2026-06-11 deprecations announcement. Reasoning model with adjustable reasoning_effort (none/low/medium/high/xhigh); reasoning tokens billed at the output rate. Cached input at 10% of base ($0.02/MTok). Batch API at flat 50% off input and output. Azure availability confirmed via learn.microsoft.com/azure/ai-foundry model list (context/output match exactly). supports_pdf set to false: modality table lists no PDF/file row — treating conservatively pending explicit confirmation. Re-verified 2026-08-11: standard ($0.20/$0.02/$1.25) and batch ($0.10/$0.625) prices unchanged against the main pricing table; remains active alongside GPT-5.6, no deprecation entry.
- **OpenAI GPT-5.6 Sol** — Frontier-tier model of the GPT-5.6 family; model_id and price cross-checked across the OpenAI pricing page, the API models overview page, and this model's own docs page. Reasoning model with adjustable reasoning_effort; model card now explicitly lists reasoning token support with billing (confirmed 2026-07-19). Context window 1,050,000, max output 128k, knowledge cutoff 2026-02-16. 2026-07-19 re-verify: batch pricing now published on the main pricing page, flat 50% off. Alias 'gpt-5.6' added — now confirmed by two sources (models overview page and this model card: 'gpt-5.6 routes to GPT-5.6 Sol'). supports_pdf false and deployment_options limited to native (no Azure availability confirmed) — still treated conservatively. Re-verified 2026-08-11: standard ($5.00/$0.50/$30.00) and batch ($2.50/$15.00) prices unchanged; confirmed as the top-recommended flagship on the models overview page, and the recommended replacement target for most older deprecated OpenAI models per the deprecations table. PRICE CHANGE 2026-09-02: standard price cut from $5.00/$0.50/$30.00 to $4.00/$0.40/$20.00 (cache still 10% of base); batch cut correspondingly from $2.50/$15.00 to $2.00/$10.00 (still flat 50% off). Confirmed by two sources: the main pricing page and this model's own docs page, which now states promotional pricing ('20% input and 33% output reductions versus prior generations') available at least through 2026-11-21. Batch rate still not restated verbatim on the model card itself, so confidence stays medium. Deprecations table redirects unchanged (gpt-5-2025-08-07, gpt-4o-2024-05-13, o3-2025-04-16, o1-2024-12-17 all still -> gpt-5.6-sol).
- **OpenAI GPT-5.6 Terra** — Mid-tier model of the GPT-5.6 family, described as balancing intelligence and cost; model_id and price cross-checked across the OpenAI pricing page, the API models overview page, and this model's own docs page. Reasoning model with adjustable reasoning_effort; model card lists reasoning token support (confirmed 2026-07-19). Context window 1,050,000, max output 128k, knowledge cutoff 2026-02-16. Cached input at 10% of base ($0.25/MTok). 2026-07-19 re-verify: base prices unchanged; batch pricing now published on the main pricing page (input $1.25 / output $7.50, flat 50% off) — added; batch rates appear on the pricing page only (model card quotes no batch rate), so confidence stays medium. supports_pdf false and deployment_options limited to native — still treated conservatively. PRICE CHANGE 2026-08-11: standard price cut from $2.50/$0.25/$15.00 to $2.00/$0.20/$12.00 (cache still 10% of base); batch cut correspondingly from $1.25/$7.50 to $1.00/$6.00 (still flat 50% off standard). Confirmed by two sources: the main pricing page and this model's own docs page both show the new $2/$0.2/$12 base rate. Batch rate still not stated on the model card itself (only implied by the flat-50%-off pattern and the pricing page's explicit $1.00/$6.00 batch column), so confidence stays medium.
- **OpenAI GPT-5.6 Luna** — Cheapest tier of the GPT-5.6 family, designed for cost-sensitive, high-volume workloads; model_id and price cross-checked across the OpenAI pricing page, the API models overview page, and this model's own docs page. Reasoning model with adjustable reasoning_effort; model card lists reasoning token support (confirmed 2026-07-19). Context window 1,050,000, max output 128k, knowledge cutoff 2026-02-16. Cached input at 10% of base ($0.10/MTok). 2026-07-19 re-verify: base prices unchanged; batch pricing now published on the main pricing page (input $0.50 / output $3.00, flat 50% off) — added; batch rates appear on the pricing page only (model card quotes no batch rate), so confidence stays medium. supports_pdf false and deployment_options limited to native — still treated conservatively. PRICE CHANGE 2026-08-11: standard price cut from $1.00/$0.10/$6.00 to $0.20/$0.02/$1.20 (cache still 10% of base, an 80% reduction); batch cut correspondingly from $0.50/$3.00 to $0.10/$0.60 (still flat 50% off standard). Confirmed by two sources: the main pricing page and this model's own docs page both show the new $0.2/$0.02/$1.2 base rate — this now undercuts even gpt-5.4-nano's $0.20/$0.02/$1.25. Batch rate still not stated on the model card itself (only implied by the flat-50%-off pattern and the pricing page's explicit $0.10/$0.60 batch column), so confidence stays medium.
- **OpenAI GPT-5.5 Pro** — 2026-07-19 re-verify: base prices unchanged; SOURCE CONFLICT on batch — the main pricing page now shows Batch input $15.00 / output $90.00 (50% off), but this model's docs page still states Batch runs at the standard $30/$180 rate; per the two-source rule the batch fields are left at $30/$180 (matching the model card and the prior verification) pending agreement between sources — re-check next refresh. Highest-effort reasoning tier of GPT-5.5, 'uses more compute to think harder'; model_id and price ($30/$180 per MTok) cross-checked across the OpenAI pricing page and this model's own docs page. No cache_read_per_mtok_usd field: the docs page states this model 'does not offer a cached input discount'. Batch API documented as the same rate as standard (no batch discount for this tier) — unusual versus sibling rows, recorded as given rather than assumed. Prompts over 272k input tokens incur a surcharge per the shared GPT-5.5 pricing note (not modeled as a separate pricing tier here, consistent with the gpt-5.5 row). supports_pdf false and deployment_options limited to native — treated conservatively pending explicit confirmation. confidence: medium because tool/structured-output flags were read from an AI-summarized fetch rather than raw page content. 2026-08-11 re-verify: base prices ($30/$180) unchanged, confirmed on both sources; SOURCE CONFLICT UNRESOLVED — pricing page still shows Batch $15.00/$90.00 while the model card still says batch runs at the standard $30/$180 rate; batch fields left unchanged at $30/$180 per the two-source rule, matching the model card — re-check next refresh.
- **OpenAI GPT-5.4 Pro** — 2026-07-19 re-verify: base prices unchanged; SOURCE CONFLICT on batch — the main pricing page now shows Batch input $15.00 / output $90.00 (50% off), but this model's docs page still quotes Batch at the standard $30/$180 rate; per the two-source rule the batch fields are left at $30/$180 (matching the model card and the prior verification) pending agreement between sources — re-check next refresh. Highest-effort reasoning tier of GPT-5.4, 'uses more compute to think harder'; model_id and price ($30/$180 per MTok) cross-checked across the OpenAI pricing page and this model's own docs page. structured_output recorded as false: re-confirmed 2026-07-19 — the docs page still explicitly states structured outputs are 'Not supported'; genuine outlier, no longer treated as a fetch artifact. No cache_read_per_mtok_usd field: no cache discount documented for this tier. Batch API documented as the same rate as standard (no batch discount). Prompts over 272k input tokens incur a surcharge per the shared GPT-5.4 pricing note (not modeled as a separate pricing tier here, consistent with the gpt-5.4 row). supports_pdf false and deployment_options limited to native — treated conservatively pending explicit confirmation. confidence: medium. 2026-08-11 re-verify: base prices ($30/$180) unchanged, confirmed on both sources; SOURCE CONFLICT UNRESOLVED — pricing page still shows Batch $15.00/$90.00 while the model card still quotes the standard $30/$180 rate; batch fields left unchanged at $30/$180 per the two-source rule — re-check next refresh.
- **OpenAI GPT-5.3-Codex** — New row this refresh — agentic coding model ('the most capable agentic coding model to date'), optimized for Codex environments; model_id and price ($1.75 input / $0.175 cached / $14.00 output per MTok) cross-checked across the OpenAI main pricing page and this model's own docs page (two-source rule met). Reasoning model with low/medium/high/xhigh reasoning_effort; reasoning tokens billed at the output rate. Context window 400,000, max output 128k, knowledge cutoff 2025-08-31. Cached input at 10% of base. Batch endpoint supported per the model card but no batch rate published on either source — batch fields omitted rather than assumed. Fine-tuning and predicted outputs not supported. supports_pdf false and deployment_options limited to native — treated conservatively pending explicit confirmation, matching how sibling rows were first added. Re-verified 2026-08-11: still active with prices ($1.75/$0.175/$14.00) unchanged, confirmed directly via its own model card (not present in the main pricing table's flagship-tier rows this pass, so the model card served as primary confirmation); no deprecation notice; model card notes GPT-5.3-Codex and GPT-5.2-Codex share identical pricing, no successor guidance given.
- **Google Gemini 2.5 Flash** — Hybrid reasoning model with dynamic thinking on by default; thinking can be disabled via thinkingBudget=0. When thinking is on, response pricing is the sum of output and thinking tokens (both billed at the output rate). Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.10/MTok vs $0.03/MTok for text/image/video. Batch Mode at flat 50% off; audio batch input is $0.50/MTok. No long-context tier. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-2.5-flash. Knowledge cutoff January 2025. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified prompt/completion price against openrouter.ai/google/gemini-2.5-flash. Re-verified 2026-07-02: prices unchanged. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 2.5 Flash-Lite** — Hybrid reasoning model; thinking is OFF by default (unlike 2.5 Flash/Pro) but can be enabled by setting thinkingBudget. When thinking is enabled, response pricing is the sum of output and thinking tokens at the output rate, so reasoning_tokens_billed is true. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.03/MTok vs $0.01/MTok for text/image/video. Batch Mode at flat 50% off; audio batch input is $0.15/MTok. No long-context tier. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-2.5-flash-lite. Knowledge cutoff January 2025. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified prompt/completion price against openrouter.ai/google/gemini-2.5-flash-lite. Re-verified 2026-07-02: prices unchanged. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 2.0 Flash** — Deprecated; shut down 2026-06-01 per the model card's verbatim notice: "Gemini 2.0 Flash is deprecated and has been shut down June 1, 2026. Migrate to Gemini 3.5 Flash to avoid service disruption." replaced_by_model_id updated 2026-07-02 to gemini-3.5-flash (Google's documented migration target), now that it is in-dataset; previously pointed at gemini-2.5-flash as a placeholder in-file successor. Standard production 2.0 Flash does not support thinking (thinking exists only on Gemini 2.5+ and 3 series per ai.google.dev/gemini-api/docs/thinking); reasoning_tokens_billed=false. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.175/MTok vs $0.025/MTok for text/image/video. Batch Mode at flat 50% off. No long-context tier. supports_pdf=false since the 2.0 Flash model card lists supported inputs as audio/images/video/text (PDF not enumerated). Confidence medium because Vertex AI's published pricing for the same model name differs ($0.15 input / $0.60 output) from AI Studio's $0.10/$0.40; AI Studio primary value retained per spec, and OpenRouter (openrouter.ai/google/gemini-2.0-flash-001) cross-confirms $0.10/$0.40. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Prices unchanged as of 2026-07-02 re-verification. Re-verified 2026-07-13: deprecation notice and shutdown date still confirmed on ai.google.dev/pricing; no new information. Re-verified 2026-07-19: deprecation notice (shut down June 1, 2026) still shown on ai.google.dev/pricing; model marked shut down on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: deprecation notice (shut down June 1, 2026) still shown on ai.google.dev/pricing; model remains shut down per ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 3.5 Flash** — New GA/stable row added 2026-07-02; this is Google's documented migration target for the retired Gemini 2.0 Flash (see that row's notes). Thinking is on by default (medium level) and configurable (minimal/low/medium/high); when on, response pricing is the sum of output and thinking tokens billed at the output rate, so reasoning_tokens_billed=true. No >200k-token pricing tier is published for this model (unlike Gemini 3.1 Pro Preview). Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.5-flash. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Explicit context caching storage is captured in cache_storage_per_mtok_per_hour_usd. Google Search grounding is billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); not captured as a structured field since sibling 2.5-series rows in this file omit the same grounding fee for consistency. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Pricing page now also lists Flex (same rates as Batch) and Priority ($2.70/$16.20) service tiers; no schema fields for those, standard/batch rates unchanged. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 3.1 Pro Preview** — New preview row added 2026-07-02; not GA. Optimized for software engineering and agentic workflows; a specialized gemini-3.1-pro-preview-customtools variant exists but is not tracked as a separate row. Input/output/cache-read pricing tiers at >200k tokens (>200k rate captured in pricing_tiers[0]); batch pricing also tiers ($2.00/$9.00 input/output above 200k) but batch has no tiered field in this schema, so only the <=200k batch rate is captured. Thinking is on by default at high level (configurable low/medium/high); thinking tokens billed at the output rate, so reasoning_tokens_billed=true. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Google Search grounding billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); omitted as a structured field for consistency with sibling rows. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices, tiers, and status (still preview, not GA) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices, tiers, and status (still preview) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 3 Flash Preview** — New preview row added 2026-07-02; not GA. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.10/MTok vs $0.05/MTok for text/image/video. No >200k-token pricing tier published. Thinking is on by default at high level (configurable minimal/low/medium/high); thinking tokens billed at the output rate, so reasoning_tokens_billed=true. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3-flash-preview. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Google Search grounding billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); omitted as a structured field for consistency with sibling rows. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices and status (still preview, not GA) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices and status (still preview) unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 3.1 Flash-Lite** — New GA/stable row added 2026-07-02. Audio input is priced separately (captured in audio_input_per_mtok_usd); cached audio input is $0.05/MTok vs $0.025/MTok for text/image/video. No >200k-token pricing tier published. Thinking is minimal by default (options: minimal or high); thinking tokens billed at the output rate, so reasoning_tokens_billed=true. Context window and max_output_tokens confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite. Knowledge cutoff January 2025, unchanged from the Gemini 2.5 family. Google Search grounding billed separately (5,000 prompts/month free, shared across Gemini 3 models, then $14/1,000 queries); omitted as a structured field for consistency with sibling rows. Vertex AI availability not confirmed as of this pass; deployment_options limited to native pending verification. Free tier (AI Studio) is published as per-minute RPM/TPM only, not per-day; free_tier omitted. Re-verified 2026-07-13: prices unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-07-19: prices unchanged against ai.google.dev/pricing; still listed stable on ai.google.dev/gemini-api/docs/models. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models.
- **Meta Llama 4 Maverick** — Multi-host pricing re-verified 2026-07-19: Together still $0.27/$0.85 per MTok (meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8) per its model page (together.ai/models/llama-4-maverick), the row's structured price as the sole confirmed direct-host rate; unchanged since 2026-05-18. Confidence lowered to medium this pass: Together's model page is the only live source for the serverless rate — together.ai/pricing and the docs.together.ai serverless catalog no longer list Maverick (fine-tuning tables only), and OpenRouter's endpoint list routes no Together endpoint. Groq formally deprecated Maverick (announced 2026-02-20, shutdown 2026-03-09 per console.groq.com/docs/deprecations); Fireworks still does not offer Maverick on serverless (on-demand deployments only). Bedrock pricing tables did not render on this pass; omitted per the single-primary-source rule for deployment_options[]. OpenRouter aggregator still routes at $0.20/$0.80 (informational only; structured input/output stay at the lowest direct-host price per the PR4 convention). Context window is Meta's published 1M (1048576 tokens); OpenRouter advertises 1.05M but the HuggingFace model card spec is 1M. max_output_tokens not published on the model card; defaulted to 8192. 17B activated / 400B total MoE with 128 experts. Re-verified 2026-08-11: Together's model page still the sole live source at $0.27/$0.85, unchanged; together.ai/pricing still omits Maverick from the serverless table (fine-tuning only). Fireworks' own model page confirms serverless still not supported (on-demand only). huggingface.co/meta-llama org page shows no new Llama 4 Maverick variant. Noted for the record: Meta Superintelligence Labs is shipping a separate closed-weights "Muse" line (Muse Spark et al.) via its own Meta Model API — unrelated to the open-weights Llama family this row tracks; out of scope here, would need its own provider row and primary source if ever added. Re-verified 2026-09-02: Together's model page (together.ai/models/llama-4-maverick) still the sole live source at $0.27/$0.85, unchanged; together.ai/pricing still omits Maverick from the serverless table. Fireworks' model page reconfirms serverless still not supported. console.groq.com/docs/deprecations still shows Maverick deprecated 2026-03-09 in favor of openai/gpt-oss-120b, no change. huggingface.co/meta-llama org page shows no new Llama 4 Maverick variant.
- **Meta Llama 4 Scout** — Price changed 2026-07-19: Groq retired Llama 4 Scout (announced 2026-06-17, shutdown 2026-07-17 per console.groq.com/docs/deprecations; removed from groq.com/pricing and console.groq.com/docs/models), so its $0.11/$0.34 rate no longer exists. Fireworks' model page now states serverless is not supported for accounts/fireworks/models/llama4-scout-instruct-basic, so Fireworks is also removed from deployment_options. Together is now the sole confirmed direct host at $0.18/$0.59 per MTok (meta-llama/Llama-4-Scout-17B-16E-Instruct, per together.ai/models/llama-4-scout; same Together rate as verified 2026-05-18 and 2026-07-13), which becomes the row's structured price. Confidence lowered to medium: Together's model page is the only live source for that rate — together.ai/pricing and the docs.together.ai serverless catalog no longer list Scout, and OpenRouter routes no Together endpoint (its Groq endpoint still listed at $0.11/$0.34 is stale post-shutdown). Bedrock pricing tables did not render on this pass; omitted per the single-primary-source rule for deployment_options[]. OpenRouter aggregator still routes at $0.10/$0.30 (informational only; structured input/output stay at the lowest direct-host price per the PR4 convention). Retired-host aliases (Groq, Fireworks) retained for lookup. Context window is Meta's published 10M (10485760 tokens); hosts cap below Meta's spec. max_output_tokens not published on the model card; defaulted to 8192. 17B activated / 109B total MoE with 16 experts. Re-verified 2026-08-11: Together's model page still the sole live source at $0.18/$0.59, unchanged; Fireworks' own model page reconfirms serverless still not supported for llama4-scout-instruct-basic (on-demand only). No new Scout variant on huggingface.co/meta-llama. Re-verified 2026-09-02: Together's model page still the sole live source at $0.18/$0.59, unchanged; Fireworks' model page reconfirms serverless still not supported. console.groq.com/docs/deprecations still shows Scout deprecated 2026-07-17 in favor of openai/gpt-oss-120b or qwen/qwen3.6-27b, no change. No new Scout variant on huggingface.co/meta-llama.
- **Meta Llama 3.3 70B Instruct** — Multi-host pricing re-verified 2026-07-19: Together still $1.04/$1.04 per MTok for Llama 3.3 70B (meta-llama/Llama-3.3-70B-Instruct-Turbo, confirmed on both together.ai/pricing and the docs.together.ai serverless catalog; unchanged since the 2026-07-02 jump from $0.88/$0.88); Groq still $0.59/$0.79 (llama-3.3-70b-versatile, confirmed on both groq.com/pricing and console.groq.com/docs/models) remains the lowest direct-host rate and stays the row's structured price. Fireworks still publishes $0.90 input on accounts/fireworks/models/llama-v3p3-70b-instruct and its own model page states serverless is not supported for this model, so Fireworks remains omitted from deployment_options. OpenRouter aggregator still routes at $0.10/$0.32 (informational only; structured input/output stay at the lowest direct-host price per the PR4 convention) and is structured as aggregators["openrouter"]. Context window 128K per Meta's spec (131072 tokens). Text-only; no vision. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff December 2023 per model card. Re-verified 2026-08-11: Groq (console.groq.com/docs/model/llama-3.3-70b-versatile) still $0.59/$0.79 and still the cheapest direct host, no deprecation announced per console.groq.com/docs/deprecations; Together (docs.together.ai serverless catalog) still $1.04/$1.04; Fireworks' own model page reconfirms serverless still not supported for llama-v3p3-70b-instruct. No new 3.3-family variant found. Price changed 2026-09-02: Groq moved llama-3.3-70b-versatile to Enterprise-only. console.groq.com/docs/models now lists its pricing and rate limits as "Contact Sales" (no public PAYG rate); console.groq.com/docs/deprecations confirms deprecation effective 2026-08-16, stating it "applies to free and developer-tier usage; enterprise customers with a committed-spend contract are not affected," and recommends migrating to openai/gpt-oss-120b or qwen/qwen3.6-27b. groq.com/pricing still redirects to the marketing homepage with no model pricing table. Groq removed from deployment_options accordingly. Together is now the sole confirmed direct host — $1.04/$1.04 per MTok, confirmed on both together.ai/pricing and together.ai/models/llama-3.3-70b, unchanged since 2026-07-02 — and becomes the row's structured price; the $0.59/$0.79 -> $1.04/$1.04 move is entirely the loss of the cheaper Groq host, not a Together repricing. Fireworks' model page (fireworks.ai/models/fireworks/llama-v3p3-70b-instruct) reconfirms serverless still not supported for llama-v3p3-70b-instruct (states "Serverless: Not supported" despite a $0.90 shared-endpoint figure elsewhere on the page). Retired-host alias llama-3.3-70b-versatile retained for lookup. No new 3.3-family variant found on huggingface.co/meta-llama.
- **Mistral Mistral Large 2 (24.11)** — La Plateforme rates ($2.00 / $6.00 per MTok) are the row's structured price. Multi-host availability: Bedrock (mistral.mistral-large-2407-v1:0 in us-west-2), Vertex AI, Azure AI Foundry, IBM watsonx. Bedrock published the 24.07 build only, not 24.11. Deprecated on La Plateforme 2026-02-27; retirement 2026-05-31 per Mistral's legacy table has now passed (re-verified 2026-07-13 via docs.mistral.ai/getting-started/models/models_overview) so the model_id should be treated as fully retired/non-callable, not merely deprecated. `mistral-large-latest` moved to Mistral Large 3 (mistral-large-2512, added to this dataset) so the alias was removed from this row to avoid a duplicate. Mistral's own deprecation table lists Mistral Medium 3.5 (mistral-medium-3-5) as the recommended alternative, not Large 3; replaced_by_model_id follows the vendor's stated alternative. Text-only; no vision (Pixtral Large was the multimodal sibling, also retired). max_output_tokens not published on the model card; defaulted to 8192. Batch API is a 50% discount where available but per-model availability is not confirmed from a single primary source on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2026-02-27, retired 2026-05-31, alternative still Mistral Medium 3.5; no field changes.
- **Mistral Mistral Medium 3** — La Plateforme rates ($0.40 / $2.00 per MTok) are the row's structured price; unchanged this pass. Mistral's launch post (2025-05-07) lists La Plateforme and Amazon SageMaker at GA with IBM watsonx, NVIDIA NIM, Azure AI Foundry, and Google Cloud Vertex as forthcoming; SageMaker is not in the deployment_options enum and Bedrock has not been confirmed, so deployment_options is restricted to native. Optimized for agentic and coding use cases. Deprecated on La Plateforme 2026-05-22 per docs.mistral.ai's legacy table; retirement 2026-08-31 has now passed, so the model_id should be treated as fully retired/non-callable, not merely deprecated. Alternative is Mistral Medium 3.5 (mistral-medium-3-5). `mistral-medium-latest` alias moved to the new Medium 3.5 row and was removed here to avoid a duplicate. max_output_tokens not published on the model card; defaulted to 8192. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2026-05-22, retirement 2026-08-31 has now passed (model no longer listed on mistral.ai/pricing/api), alternative still Mistral Medium 3.5; no price field changes.
- **Mistral Mistral Small 3** — La Plateforme rates ($0.10 / $0.30 per MTok) are the row's structured price; per Mistral's launch post, half the price of the previous mistral-small ($0.20 / $0.60). 24B-parameter latency-optimized model under Apache 2.0; text-only. Context window 32K per Mistral's spec (33000 tokens rounded; 32768 used here). Deprecated on La Plateforme 2025-11-06 and retired 2025-11-30 per Mistral's legacy table (both dates now well in the past; model_id should be treated as fully retired). Chain of intermediate successors (mistral-small-2503 / 3.1, mistral-small-2506 / 3.2, both also since deprecated) led to Mistral Small 4 (mistral-small-2603, added to this dataset this pass); replaced_by_model_id now resolves in-file. max_output_tokens not published on the model card; defaulted to 8192. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2025-11-06, retired 2025-11-30, alternative still Mistral Small 4; no field changes.
- **Mistral Codestral 25.08** — La Plateforme rates ($0.30 / $0.90 per MTok) are the row's structured price; unchanged this pass, cross-checked against mistral.ai/pricing/api. Code-specialized model optimized for fill-in-the-middle (FIM), code completion, code correction, and test generation; supports tool use and structured output per the 25.08 release. Not on Mistral's legacy/deprecation table on this date; still current. Context window corrected to 128K (131072 tokens) per the live docs.mistral.ai/models/model-cards/codestral-25-08 spec card, which lists 128k context; the original launch blog's 256K figure (262144, previously recorded here) does not match the current model card and is superseded by it. Also available on Google Cloud Vertex AI Model Garden as `codestral-2` under the `mistralai` publisher (Mistral Docs: Vertex AI cloud deployments page). max_output_tokens not published on the model card; defaulted to 8192. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.30/$0.90 on mistral.ai/pricing/api, still absent from docs.mistral.ai's deprecation table (still current); no field changes.
- **Mistral Pixtral Large** — La Plateforme rates ($2.00 / $6.00 per MTok) are the row's structured price; pricing parity with Mistral Large 2 since Pixtral Large is the multimodal 124B-parameter open-weight model built on top of Mistral Large 2. Vision-capable: handles documents, charts, and natural images alongside text. Context window 128K (131072 tokens). Bedrock publishes the 25.02 refresh (`mistral.pixtral-large-2502-v1:0`, also routed via `us.mistral.pixtral-large-2502-v1:0`), not the 24.11 build. Deprecated on La Plateforme 2026-02-27; retirement 2026-05-31 per Mistral's legacy table has now passed (re-verified 2026-07-13) so the model_id should be treated as fully retired/non-callable. Mistral's stated alternative is Mistral Medium 3.5 (mistral-medium-3-5, added to this dataset this pass); Pixtral as a standalone product line has been discontinued, its vision capability absorbed into Large 3 / Medium 3.5. `pixtral-large-latest` alias is likely non-functional post-retirement but is left on this row since nothing else claims it. max_output_tokens not published on the model card; defaulted to 8192. Vertex/Azure availability not confirmed for Pixtral Large on this date. Batch API discount per-model availability not confirmed on this date, so batch fields are unset. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02 via mistral.ai/pricing/api and docs.mistral.ai/getting-started/models/models_overview: still deprecated 2026-02-27, retired 2026-05-31, alternative still Mistral Medium 3.5; no field changes.
- **Mistral Mistral Large 3** — La Plateforme rates ($0.50 / $1.50 per MTok) confirmed on mistral.ai/pricing/api under alias `mistral-large-latest`; canonical dated model_id `mistral-large-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/mistral-large-3-25-12. Open-weight, general-purpose multimodal model (text + image input) with a Mixture-of-Experts architecture (41B active / 675B total parameters), released 2025-12-02 alongside the Ministral 3 family via the same announcement post. Successor to mistral-large-2411 per Mistral's own alternative-model recommendation is actually Mistral Medium 3.5, not this row, per the legacy table; Large 3 is nonetheless the direct version-number successor and is tracked here as a new, independently-priced row. Bedrock/Vertex/Azure availability not confirmed for this build on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.50/$1.50 on mistral.ai/pricing/api under `mistral-large-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes.
- **Mistral Mistral Medium 3.5** — La Plateforme rates ($1.50 / $7.50 per MTok) confirmed on mistral.ai/pricing/api under alias `mistral-medium-latest`; canonical model_id `mistral-medium-3-5` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04, released 2026-04-28. Frontier-class multimodal model (text + image input) optimized for agentic and coding use cases; released as open weights under a Modified MIT license. This is the model Mistral's own deprecation table names as the current alternative for mistral-large-2411, pixtral-large-2411, and mistral-medium-2505 (all now deprecated in this dataset), so replaced_by_model_id on those rows points here. Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $1.50/$7.50 on mistral.ai/pricing/api under `mistral-medium-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes.
- **Mistral Mistral Small 4** — La Plateforme rates ($0.15 / $0.60 per MTok) confirmed on mistral.ai/pricing/api under alias `mistral-small-latest`; canonical model_id `mistral-small-2603` and 256K (262144) context confirmed on the same docs.mistral.ai model card, released 2026-03-16. No dedicated mistral.ai/news announcement post was found for this release on this date, so the docs model card is used as source_url; the pricing page (mistral.ai/pricing/api) is the second confirming source, satisfying the two-source rule. Hybrid model unifying instruct, reasoning, and coding capabilities (119B parameters, 6.5B active); multimodal (text + image input). Named as the alternative to mistral-small-2501 (chain via 2503/2506) and to mistral-small-2506 in Mistral's legacy table. Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.15/$0.60 on mistral.ai/pricing/api under `mistral-small-latest`, still absent from docs.mistral.ai's deprecation table (still current; mistral-small-2506 is the deprecated row alternative-pointing here, not this row); no field changes.
- **Mistral Ministral 3 3B** — First Ministral-family row in this dataset. La Plateforme rates ($0.10 / $0.10 per MTok) confirmed on mistral.ai/pricing/api under alias `ministral-3b-latest`; canonical model_id `ministral-3b-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/ministral-3-3b-25-12, released 2025-12-02. Smallest/most efficient model in the Ministral 3 family; edge-deployment focused; multimodal (text + image input per the model card's 'robust language and vision capabilities'). Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.10/$0.10 on mistral.ai/pricing/api under `ministral-3b-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes.
- **Mistral Ministral 3 8B** — La Plateforme rates ($0.15 / $0.15 per MTok) confirmed on mistral.ai/pricing/api under alias `ministral-8b-latest`; canonical model_id `ministral-8b-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/ministral-3-8b-25-12, released 2025-12-02. Best-in-class text and vision capabilities for edge deployment; multimodal (text + image input). Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.15/$0.15 on mistral.ai/pricing/api under `ministral-8b-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes.
- **Mistral Ministral 3 14B** — La Plateforme rates ($0.20 / $0.20 per MTok) confirmed on mistral.ai/pricing/api under alias `ministral-14b-latest`; canonical model_id `ministral-14b-2512` and 256K (262144) context confirmed on docs.mistral.ai/models/model-cards/ministral-3-14b-25-12, released 2025-12-02. Largest model in the Ministral 3 family, performance comparable to the larger Mistral Small 3.2 24B; multimodal (text + image input), optimized for local deployment. Bedrock/Vertex/Azure availability not confirmed on this date, so deployment_options is restricted to native. max_output_tokens not published on the model card; defaulted to 8192. Knowledge cutoff not published by Mistral. Re-verified 2026-09-02: still $0.20/$0.20 on mistral.ai/pricing/api under `ministral-14b-latest`, still absent from docs.mistral.ai's deprecation table (still current); no field changes.
- **Cohere Command A** — Cohere's flagship 111B-parameter model: 256K context, text-only, optimized for tool use, RAG, agents, and 23-language multilingual workloads. Price ($2.50 / $10.00 per MTok) per artificialanalysis.ai citing Cohere's API (re-confirmed on its Command A model page 2026-07-19); Command A is still not listed on cohere.com/pricing as of 2026-09-02 (page only publishes a "Legacy Model Pricing" FAQ table for Command/Command-light/Command R/Command R+ 04-2024/Command R+ 08-2024, plus Model Vault dedicated-instance rates; newer generative models including Command A still route to "Get in touch for custom enterprise pricing"), so confidence remains medium. docs.cohere.com/docs/models confirms `command-a-03-2025` is still Live (256K context, 8K max output, text-only) and Cohere's deprecations page (docs.cohere.com/docs/deprecations, re-checked 2026-09-02) does not list it, so status and price are unchanged since the 2026-08-11 pass. Cohere's model catalog still lists newer generative models (command-a-reasoning-08-2025, command-a-vision-07-2025, command-a-translate-08-2025, command-r-08-2024, and command-a-plus-05-2026 — Live on docs.cohere.com/docs/models with 128K context / 64K max output, text+image input) but none of the generative-model additions publish per-token input/output pricing on cohere.com/pricing or a corroborating secondary source, so they remain deferred as of 2026-09-02. Cohere docs still don't surface AWS Bedrock availability for this model (so `bedrock` remains omitted from deployment_options); Azure AI Foundry availability is published but uses per-deployment IDs, so no Azure alias is encoded. Oracle OCI exposes it as `cohere.command-a-03-2025` (kept as alias). Cache and batch pricing not published by Cohere. Knowledge cutoff not published on the Cohere model card.
- **Cohere Command R+** — Cohere Platform rates ($2.50 / $10.00 per MTok) are the row's structured price, listed on cohere.com/pricing as "Command R+ 08-2024" (page now files it under a "Legacy Model Pricing" FAQ table alongside Command/Command-light/Command R/Command R+ 04-2024, meaning it's held for existing customers rather than a headline SKU, but the price is unchanged and it remains orderable, re-confirmed 2026-09-02). 128K context, text-only, optimized for complex RAG and multi-step tool use. Cohere's deprecations page (docs.cohere.com/docs/deprecations, re-checked 2026-09-02) still sunsets only the predecessor `command-r-plus-04-2024` on 2025-09-15 and names this 08-2024 build as the recommended replacement, so it is active on Cohere Platform; docs.cohere.com/docs/models also lists `command-r-plus-08-2024` as Live. Bedrock SKU `cohere.command-r-plus-v1:0` launched Aug 2024 with a Mar 2024 knowledge cutoff (per Bedrock model card), matching this row; as of the prior pass Bedrock had marked the model "Legacy" with an EOL of 2026-08-19, which has now passed — this refresh did not re-check the Bedrock console directly, so confirm the Bedrock-SKU status next pass (this is a Bedrock-side lifecycle marker, not a Cohere platform deprecation, so the row itself stays active on Cohere Platform regardless). Azure AI Foundry availability published by Cohere; per-deployment IDs there, so no Azure alias is encoded. Cache and batch pricing not published by Cohere. Cohere's docs/models catalog also now lists a plain `command-r-08-2024` (non-plus) SKU, but it has no published per-token pricing on cohere.com/pricing or a corroborating secondary source, so it is deferred rather than added this pass.
- **Cohere Aya Expanse 32B** — 32B-parameter multilingual research/generative model (23 languages), 128K context, text-only. Cohere's "Legacy Model Pricing" FAQ table on cohere.com/pricing still lists the same single bundled rate for "Aya Expanse (8B and 32B): $0.50 / $1.50 per MTok" (not broken out per size) as of 2026-09-02, so confidence remains medium. docs.cohere.com/docs/models lists `c4ai-aya-expanse-32b` as Live, 128K context, 4K max output, text-only, unchanged, and it still does not appear on Cohere's deprecations page (re-checked 2026-09-02). The 8B sibling `c4ai-aya-expanse-8b` reached its scheduled retirement/shutdown date of 2026-04-04 per docs.cohere.com/docs/deprecations (replacement: `command-r7b-12-2024`, `command-a-03-2025`, or `command-a-reasoning-08-2025`) and remains intentionally excluded as a row (never tracked, so no deprecated_at entry needed here). Tool use, structured output, and knowledge cutoff not published by Cohere for this model; booleans defaulted to false pending confirmation. Cache and batch pricing not published.
- **xAI Grok 4** — xAI native rates ($3.00 / $15.00 per MTok, $0.75 cached input) are the row's structured price; prompts above 128K total tokens are billed at the higher pricing_tiers rate ($6.00 / $30.00) per xAI's documented long-context tiering. Grok 4 (snapshot `grok-4-0709`, released 2025-07-09) was xAI's flagship reasoning model: reasoning is always on (thinking tokens billed at the output rate, hence reasoning_tokens_billed: true), parallel tool calling and structured outputs supported, accepts text and image inputs. max_output_tokens of 256000 reflects xAI's documented "up to 256K tokens of output" within the shared 256K prompt+response context. Retired from the xAI API on 2026-05-15 12:00 PM PT alongside seven other legacy slugs; requests to `grok-4-0709` and `grok-4` continue to resolve but are now redirected to `grok-4.3` with `low` reasoning effort and billed at grok-4.3 rates. Successor is `grok-4.3`, captured in this dataset and referenced via `replaced_by_model_id`. xAI's API is native-only (not on Bedrock/Vertex/Azure). Batch API not published for this model. Re-verified 2026-07-02, 2026-07-13, 2026-07-19, and 2026-08-11: `grok-4-0709` absent from both docs.x.ai/docs/models and docs.x.ai/docs/pricing, and still listed on docs.x.ai/developers/migration/may-15-retirement with redirect target `grok-4.3`, confirming retirement status is unchanged.
- **xAI Grok 3** — xAI native rates ($3.00 / $15.00 per MTok, $0.75 cached input) are the row's structured price (xAI's pricing page and mem0/pricepertoken aggregator both report $3/$15; artificialanalysis.ai reports a higher $4/$20 — choosing xAI-aligned figures). Grok 3 (released 2025-02-19) was xAI's flagship non-reasoning chat model; text-only inputs, function calling and structured outputs supported, 131,072-token combined prompt+response context window. Not a reasoning model (direct responses, no extended chain-of-thought; the reasoning sibling was `grok-3-mini`, not in this dataset). max_output_tokens defaulted to the documented context cap; xAI does not publish a separate max-output limit beyond the shared 131,072-token window. Retired from the xAI API on 2026-05-15 12:00 PM PT; requests to `grok-3` continue to resolve but are now redirected to `grok-4.3` with `none` reasoning effort and billed at grok-4.3 rates. Successor is `grok-4.3`, captured in this dataset and referenced via `replaced_by_model_id`. xAI's API is native-only. Batch API not published for this model. Re-verified 2026-07-02, 2026-07-13, 2026-07-19, and 2026-08-11: `grok-3` absent from both docs.x.ai/docs/models and docs.x.ai/docs/pricing, and still listed on docs.x.ai/developers/migration/may-15-retirement with redirect target `grok-4.3`, confirming retirement status is unchanged.
- **xAI Grok Code Fast 1** — xAI native rates ($0.20 / $1.50 per MTok, $0.02 cached input) are the row's structured price. Grok Code Fast 1 (released 2025-08-26) was xAI's speedy, economical coding-specialized reasoning model: 314B-parameter MoE architecture, 256K combined prompt+response context, agentic coding focus, visible reasoning traces (`reasoning_content` field in streaming responses), function calling and structured outputs supported, text-only. Reasoning is enabled by default so reasoning tokens are billed at the output rate. max_output_tokens defaulted to the documented 256K context cap; xAI does not publish a separate max-output limit beyond the shared window. Retired from the xAI API on 2026-05-15 12:00 PM PT; requests to `grok-code-fast-1` continue to resolve. CORRECTION on 2026-07-13 re-verification: docs.x.ai/developers/migration/may-15-retirement now states the redirect target is `grok-build-0.1` ("After May 15, requests to `grok-code-fast-1` are routed to `grok-build-0.1`"), not `grok-4.3` as previously recorded on 2026-07-02 — `replaced_by_model_id` updated accordingly; this is a referential correction, not a price change to this row, so `last_changed_at` is not bumped. `grok-build-0.1` is now captured in this dataset. xAI's API is native-only. Batch API not published for this model. Knowledge cutoff not published by xAI. `grok-code-fast-1` remains absent from both docs.x.ai/docs/models and docs.x.ai/docs/pricing, confirming retirement status is unchanged. Re-verified 2026-07-19 and 2026-08-11: still absent from both vendor pages, and docs.x.ai/developers/migration/may-15-retirement still lists redirect target `grok-build-0.1`.
- **xAI Grok 4.3** — xAI native rates ($1.25 / $2.50 per MTok, $0.20 cached input) are the row's structured price, listed on docs.x.ai/docs/models and docs.x.ai/docs/pricing; base $1.25/$2.50 unchanged since 2026-05-15. PRICE STRUCTURE UPDATE 2026-07-19: xAI now publishes a cached-input rate ($0.20/MTok) and long-context tiering for this model — prompts ≥200K tokens bill at $2.50 input / $5.00 output / $0.40 cached input — captured in cache_read_per_mtok_usd and pricing_tiers; cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-4.3: prompt $0.00000125/token = $1.25/MTok, completion $0.0000025/token = $2.50/MTok, input_cache_read $0.0000002/token = $0.20/MTok) — satisfies the two-source rule; last_changed_at bumped. 1M-token combined prompt+response context window (a 4x expansion over the 256K window on Grok 4 / Grok Code Fast 1). Successor to `grok-4-0709`, `grok-3`, and (via corrected redirect) predecessor line of `grok-build-0.1`, retired 2026-05-15 12:00 PM PT. Thinking mode is exposed via the `reasoning_effort` parameter (`none` / `low` / `medium` / `high`); when reasoning is on, thinking tokens are billed at the output rate, hence reasoning_tokens_billed: true. Accepts text and image inputs, text output; parallel tool calling and structured outputs supported. supports_pdf flipped to true on 2026-07-19: OpenRouter's architecture for x-ai/grok-4.3 now lists a `file` input modality — the same derivation already used for the grok-4.5 and grok-build-0.1 rows; capability fix, not a price change. max_output_tokens reflects the documented 1M-token shared window — xAI does not publish a separate max-output limit. Knowledge cutoff and release date still not published by xAI; omitted rather than guessed. xAI's API is native-only (not on Bedrock / Vertex / Azure). Batch API not published for this model. Remains listed on both vendor pages with no deprecation banner as of 2026-07-19; xAI's top-line recommendation remains `grok-4.5`. The `grok-4.20-0309-*` beta variants deferred on 2026-07-02/2026-07-13 were added as rows on 2026-07-19 after sources converged — see those rows. Re-verified 2026-08-11: price, tiering, and context window unchanged on both docs.x.ai/docs/models and docs.x.ai/docs/pricing.
- **xAI Grok 4.5** — Grok 4.5 launched 2026-07-08 as xAI's top-line recommended model ("the most intelligent and fastest model we've built"; docs.x.ai/docs/models: "For everything else, including code, use Grok 4.5"), positioned for coding, agentic, and knowledge-work use. xAI native rates ($2.00 / $6.00 per MTok, $0.30 cached input) are the row's structured price, confirmed on both docs.x.ai/docs/models and docs.x.ai/docs/pricing, cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-4.5: prompt $0.000002/token = $2.00/MTok, completion $0.000006/token = $6.00/MTok, input_cache_read $0.0000003/token = $0.30/MTok cache read) — satisfies the two-source rule. PRICE CHANGE 2026-07-19: cached-input rate dropped from $0.50 to $0.30/MTok (both vendor pages and OpenRouter agree; $0.50 was the verified rate on 2026-07-13), and xAI now publishes long-context tiering — prompts ≥200K tokens bill at $4.00 input / $12.00 output / $0.60 cached input — captured in pricing_tiers (OpenRouter does not model per-tier rates; tier figures are per docs.x.ai/docs/models and docs.x.ai/docs/pricing, which agree); last_changed_at bumped. Context window 500,000 tokens (smaller than grok-4.3's 1M) per docs.x.ai and OpenRouter agreement; knowledge cutoff "February 1, 2026" per docs.x.ai/docs/models verbatim. max_output_tokens defaulted to the 500K context cap; xAI does not publish a separate max-output limit (OpenRouter reports max_completion_tokens: null, consistent with no separate cap). Input modalities text + image (OpenRouter architecture also lists a `file` input modality, taken here as supports_pdf: true; images up to 20MiB, jpg/jpeg/png); output text-only. OpenRouter's supported_parameters list for this model includes `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), and `reasoning`/`reasoning_effort`/`include_reasoning` (reasoning_tokens_billed: true, consistent with the `reasoning_effort` parameter pattern on grok-4.3). Audio input/output not supported. Batch pricing and deployment options beyond native not published by xAI as of last_verified; xAI's API remains native-only (not on Bedrock/Vertex/Azure). Does not affect the existing `grok-4.3` row — see that row's notes. Re-verified 2026-08-11: price, tiering, and context window unchanged on both docs.x.ai/docs/models and docs.x.ai/docs/pricing.
- **xAI Grok 4.6** — New row, added 2026-09-02. Grok 4.6 launched as xAI's new top-line recommended model (docs.x.ai/docs/models: "For everything else, including code, use Grok 4.6. It is the most intelligent and fastest model we've built." — the same positioning language previously used for grok-4.5), for coding, agentic, and knowledge-work use. xAI native rates ($2.00 / $6.00 per MTok, $0.50 cached input) confirmed on both docs.x.ai/docs/models and docs.x.ai/docs/pricing, cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-4.6: prompt $0.000002/token = $2.00/MTok, completion $0.000006/token = $6.00/MTok, input_cache_read $0.0000005/token = $0.50/MTok) — satisfies the two-source rule. Long-context tiering: prompts ≥200K tokens bill at $4.00 input / $12.00 output / $1.00 cached input, per docs.x.ai/docs/pricing (OpenRouter does not model per-tier rates; tier figures are vendor-only). Context window 500,000 tokens, unchanged from grok-4.5; knowledge cutoff "February 1, 2026" per docs.x.ai/docs/models verbatim. max_output_tokens defaulted to the 500K context cap; xAI does not publish a separate max-output limit (OpenRouter reports no separate completion cap). Input modalities text + image (OpenRouter architecture also lists a `file` input modality, taken here as supports_pdf: true, same derivation used for sibling xAI rows); output text-only. OpenRouter's supported_parameters for x-ai/grok-4.6 include `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), and `reasoning`/`reasoning_effort`/`include_reasoning` (reasoning_tokens_billed: true). Audio input/output not supported. Batch pricing and deployment options beyond native not published; xAI's API remains native-only (not on Bedrock/Vertex/Azure). The predecessor `grok-4.5` row remains listed on both vendor pages with no deprecation banner as of 2026-09-02 and unchanged pricing ($2.00/$6.00/$0.30 cached, tiering $4.00/$12.00/$0.60); not marked deprecated in this dataset since xAI has only repositioned its top-line recommendation, not retired the model — grok-4.5's `featured: true` flag is preserved per the runbook's never-remove-featured rule. `featured: true` was NOT added to this new grok-4.6 row in this pass: that requires a `pnpm costs:featured --write` lockfile regen (`web/lib/featured.lock.json`), which is outside this dispatch's scope (costs/llm.json only, per workflow constraints). Flagging grok-4.6 as a featured candidate for the controller to evaluate in a follow-up pass.
- **xAI Grok Build 0.1** — Grok Build 0.1 (public beta, first opened via API 2026-05-20) is xAI's fast, economical agentic-coding model that powers the Grok Build CLI; it is the documented redirect target for the retired `grok-code-fast-1` slug (docs.x.ai/developers/migration/may-15-retirement: "After May 15, requests to `grok-code-fast-1` are routed to `grok-build-0.1`") — see that row's `replaced_by_model_id`. xAI native rates ($1.00 / $2.00 per MTok) confirmed on docs.x.ai/docs/models and docs.x.ai/docs/pricing, cross-checked against OpenRouter's public models API (openrouter.ai/api/v1/models, x-ai/grok-build-0.1: prompt $0.000001/token = $1.00/MTok, completion $0.000002/token = $2.00/MTok, input_cache_read $0.0000002/token = $0.20/MTok cache read) — satisfies the two-source rule; base and cached rates re-confirmed unchanged on 2026-07-19. PRICE STRUCTURE UPDATE 2026-07-19: xAI now publishes long-context tiering for this model — prompts ≥200K tokens bill at $2.00 input / $4.00 output / $0.40 cached input — captured in pricing_tiers (tier figures per docs.x.ai/docs/models and docs.x.ai/docs/pricing, which agree; OpenRouter does not model per-tier rates); last_changed_at bumped. Context window 256,000 tokens per both sources; max_output_tokens defaulted to the 256K context cap (OpenRouter reports max_completion_tokens: null, no separate limit published). Input modalities text + image (OpenRouter architecture lists a `file` input modality too, taken here as supports_pdf: true); output text-only. OpenRouter's supported_parameters for this model include `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), and `reasoning`/`include_reasoning` (reasoning_tokens_billed: true — consistent with predecessor `grok-code-fast-1`'s visible reasoning traces). Audio input/output not supported. Knowledge cutoff and release date beyond the 2026-05-20 API-opening not published by xAI; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure); limited regional availability noted by third-party docs (us-east-1, us-west-2) but not encoded as a separate deployment_option since it isn't a distinct hosted-API SKU distinction xAI itself publishes. Re-verified 2026-08-11: price, tiering, and context window unchanged on both docs.x.ai/docs/models and docs.x.ai/docs/pricing.
- **xAI Grok 4.20 Reasoning** — New row, added on the 2026-07-19 refresh; deferred on 2026-07-02 and 2026-07-13 because a third-party writeup (buildfastwithai.com, 2026-03-14) contradicted xAI's published specs ($2.00/$6.00 per MTok, 2M context). Sources have now converged on price: docs.x.ai/docs/models and docs.x.ai/docs/pricing both list `grok-4.20-0309-reasoning` at $1.25 / $2.50 per MTok with $0.20 cached input and long-context tiering (prompts ≥200K tokens: $2.50 / $5.00, $0.40 cached input), and OpenRouter's public models API corroborates the family pricing (openrouter.ai/api/v1/models, x-ai/grok-4.20: prompt $0.00000125/token = $1.25/MTok, completion $0.0000025/token = $2.50/MTok, input_cache_read $0.0000002/token = $0.20/MTok). confidence: medium because the secondary-source match is at family level, not exact slug — OpenRouter lists the family as `x-ai/grok-4.20` (no `-0309-` snapshot suffix) and reports a 2M-token context, while xAI's own models and pricing pages state 1M for this API model_id; the vendor figure (1,000,000) is used as authoritative for the native API. Grok 4.20 beta reasoning variant (snapshot suffix 0309, consistent with a 2026-03-09 snapshot; xAI does not publish a release date). Reasoning tokens billed at the output rate (reasoning_tokens_billed: true). Capabilities per OpenRouter x-ai/grok-4.20 supported_parameters: `tools`/`tool_choice` (supports_tool_use: true), `response_format`/`structured_outputs` (structured_output: true), `reasoning`/`include_reasoning`. Input modalities text + image (OpenRouter architecture also lists `file`, taken as supports_pdf: true, same derivation as sibling xAI rows); output text-only; audio not supported. max_output_tokens defaulted to the 1M context cap — no separate max-output limit published. Knowledge cutoff not published; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure). Priced identically to `grok-4.3`. Sibling rows: `grok-4.20-0309-non-reasoning`, `grok-4.20-multi-agent-0309`. Re-verified 2026-08-11: price, tiering, and context window unchanged on docs.x.ai/docs/models and docs.x.ai/docs/pricing; confidence left at medium pending re-check of the OpenRouter family-level slug/context discrepancy.
- **xAI Grok 4.20 Non-Reasoning** — New row, added on the 2026-07-19 refresh; deferred on 2026-07-02 and 2026-07-13 because a third-party writeup (buildfastwithai.com, 2026-03-14) contradicted xAI's published specs. Sources have now converged on price: docs.x.ai/docs/models and docs.x.ai/docs/pricing both list `grok-4.20-0309-non-reasoning` at $1.25 / $2.50 per MTok with $0.20 cached input and long-context tiering (prompts ≥200K tokens: $2.50 / $5.00, $0.40 cached input), and OpenRouter's public models API corroborates the family pricing under `x-ai/grok-4.20` (prompt $0.00000125/token, completion $0.0000025/token, input_cache_read $0.0000002/token). confidence: medium for the same reasons as the sibling `grok-4.20-0309-reasoning` row: family-level (not exact-slug) secondary corroboration, and a context-window discrepancy (OpenRouter reports 2M; xAI's models and pricing pages state 1M — vendor figure 1,000,000 used as authoritative for the native API). Grok 4.20 beta non-reasoning variant: direct responses without extended chain-of-thought, so reasoning_tokens_billed: false (analogous to the retired `grok-3` and to grok-4.3's `reasoning_effort: none` mode). Function calling and structured outputs supported per OpenRouter x-ai/grok-4.20 supported_parameters (`tools`/`tool_choice`, `response_format`/`structured_outputs`). Input modalities text + image (OpenRouter architecture also lists `file`, taken as supports_pdf: true, same derivation as sibling xAI rows); output text-only; audio not supported. max_output_tokens defaulted to the 1M context cap — no separate max-output limit published. Knowledge cutoff and release date not published; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure). Priced identically to `grok-4.3`. Sibling rows: `grok-4.20-0309-reasoning`, `grok-4.20-multi-agent-0309`. Re-verified 2026-08-11: price, tiering, and context window unchanged on docs.x.ai/docs/models and docs.x.ai/docs/pricing; confidence left at medium pending re-check of the OpenRouter family-level slug/context discrepancy.
- **xAI Grok 4.20 Multi-Agent** — New row, added on the 2026-07-19 refresh; deferred on 2026-07-02 and 2026-07-13 (a third-party writeup then reported the multi-agent API as 'coming soon' with conflicting prices). Sources have now converged: docs.x.ai/docs/models and docs.x.ai/docs/pricing both list `grok-4.20-multi-agent-0309` at $1.25 / $2.50 per MTok with $0.20 cached input and long-context tiering (prompts ≥200K tokens: $2.50 / $5.00, $0.40 cached input), and OpenRouter carries a live listing `x-ai/grok-4.20-multi-agent` with matching pricing (prompt $0.00000125/token, completion $0.0000025/token, input_cache_read $0.0000002/token) — the API is live, no longer 'coming soon'. confidence: medium: OpenRouter's slug lacks the `-0309` snapshot suffix and reports a 2M-token context while xAI's own pages state 1M — vendor figure (1,000,000) used as authoritative for the native API. Grok 4.20 beta multi-agent variant: multiple agents run in parallel to conduct deep research, coordinate tool use internally, and synthesize a final answer (per OpenRouter's model description). supports_tool_use omitted rather than guessed: OpenRouter's supported_parameters for x-ai/grok-4.20-multi-agent do not include `tools`/`tool_choice` (client-side function calling), unlike the sibling variants — tool use appears to be internal to the agent swarm; xAI does not publish a definitive statement. `response_format`/`structured_outputs` supported (structured_output: true); `reasoning`/`reasoning_effort` supported, thinking tokens billed at the output rate (reasoning_tokens_billed: true). Input modalities text + image (OpenRouter architecture also lists `file`, taken as supports_pdf: true); output text-only; audio not supported. max_output_tokens defaulted to the 1M context cap — no separate max-output limit published. Knowledge cutoff and release date not published; omitted rather than guessed. Batch pricing not published. xAI's API is native-only (not on Bedrock/Vertex/Azure). Priced identically to `grok-4.3`. Sibling rows: `grok-4.20-0309-reasoning`, `grok-4.20-0309-non-reasoning`. Re-verified 2026-08-11: price, tiering, and context window unchanged on docs.x.ai/docs/models and docs.x.ai/docs/pricing; confidence left at medium pending re-check of the OpenRouter family-level slug/context discrepancy.
- **Alibaba Qwen3-Max** — Alibaba DashScope (International) tiered pricing by input-token bucket: 0-32K = $1.20 input / $6.00 output per MTok (base row rate); 32K-128K = $2.40 / $12.00; 128K-252K = $3.00 / $15.00 (captured in pricing_tiers) — unchanged at last_verified. Context window 262,144 tokens (DashScope publishes 252K as the top-tier pricing ceiling; Qwen team and OpenRouter publish the full 262,144 model context). Max output 32,768 tokens. `qwen3-max` is a rolling alias: as of 2026-07-02 it resolves to `qwen3-max-2026-01-23` (previously `qwen3-max-2025-09-23`, still separately listed); pricing identical across both dated snapshots, so this is a metadata-only pointer change. Knowledge cutoff 2025-06-30 (unverified against the new snapshot; vendor page does not publish per-snapshot cutoffs, carried over from prior verification). Hybrid thinking model: thinking mode disabled by default but available via `/think` (and disabled via `/no_think`); when enabled, thinking tokens are billed at the output rate, hence reasoning_tokens_billed: true. Text-only inputs and outputs (the Qwen3-VL family is a separate set of model_ids). Tool calling and structured outputs supported via the DashScope and OpenAI-compatible endpoints. Explicit context cache discounts cached input tokens to 10% of the standard rate, but DashScope does not publish a single cache_read figure across the tiered input rates, so cache_read_per_mtok_usd is omitted rather than guessed. Deployment via DashScope (Model Studio) only at last_verified; not on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. CHANGE 2026-07-02: batch_input_per_mtok_usd/batch_output_per_mtok_usd removed — the vendor pricing page's International/Singapore "Qwen-Max" table for qwen3-max no longer carries a batch-inference-discount badge (verified against raw page HTML: only the "context caching discount" annotation remains). The 50% batch discount badge is now only shown on the Chinese-mainland deployment row for this model family, which this dataset does not price (DashScope International is canonical). Prior batch figures ($0.60/$3.00) were the International 50%-off rate and are no longer offered there as of this verification. Re-verified 2026-07-13: prices, tiers, and batch/cache posture unchanged against alibabacloud.com/help/en/model-studio/billing; still no batch badge on the International row (China-mainland row still carries it), model_id still active (not deprecated), no newer dated snapshot beyond qwen3-max-2026-01-23. Re-verified 2026-07-19: International (Singapore) list prices and tiers unchanged ($1.20/$6.00, $2.40/$12.00, $3.00/$15.00); alias still resolves to qwen3-max-2026-01-23; no deprecation notice. New since last pass: a 'Limited-time 50% off' promo label now appears on the International qwen3-max row — promo only, list prices (recorded here) unchanged, so last_changed_at not bumped. Re-verified 2026-08-11: International tiers unchanged ($1.20/$6.00, $2.40/$12.00, $3.00/$15.00 across 0-32K/32K-128K/128K-256K); alias still qwen3-max-2026-01-23, no newer dated snapshot; model_id still listed and active on the billing page (not in the 'recommended models' table alongside qwen3.7-max/qwen3.7-plus/qwen3.6-flash per the models page, but still priced with no deprecation banner); no deprecation notice. Re-verified 2026-09-02: International tiers unchanged ($1.20/$6.00, $2.40/$12.00, $3.00/$15.00 across 0-32K/32K-128K/128K-256K); alias still resolves to qwen3-max-2026-01-23 with no newer dated snapshot; row still priced with no deprecation banner on the billing page. A new `qwen3.8-max` flagship generation launched this pass (added as a separate row below) alongside `qwen3.7-max`; qwen3-max remains a distinct, separately-priced, still-active SKU (not superseded/replaced) so no `deprecated_at`/`replaced_by_model_id` set here.
- **Alibaba Qwen3-Coder-Plus** — Alibaba DashScope (International) tiered pricing by input-token bucket: 0-32K = $1.00 / $5.00 per MTok (base row rate); 32K-128K = $1.80 / $9.00; 128K-256K = $3.00 / $15.00; 256K-1M = $6.00 / $60.00 (captured in pricing_tiers) — unchanged at last_verified. 1,000,000-token context window with 65,536 max output tokens. `qwen3-coder-plus` still resolves to `qwen3-coder-plus-2025-09-23` (no newer dated snapshot published; an older `qwen3-coder-plus-2025-07-22` snapshot remains listed separately at identical pricing). Built on the Qwen3-Coder 480B-A35B MoE base; positioned for agentic coding (robust tool calling and environment interaction). Not a thinking/reasoning SKU (no chain-of-thought billing semantics), so reasoning_tokens_billed is false. Text-only modalities. Explicit context cache discounts cached input to 10% of the standard rate; implicit cache to 20%; DashScope does not publish a single cache_read figure across tiered input rates, so cache_read_per_mtok_usd is omitted rather than guessed. Knowledge cutoff not published. Deployment via DashScope (Model Studio) only at last_verified. CHANGE 2026-07-02: batch_input_per_mtok_usd/batch_output_per_mtok_usd removed — the vendor's "Qwen-Coder" pricing section intro no longer mentions batch-call pricing at all (unlike the "Qwen-Max" section, which still documents a batch discount), and neither the International nor Chinese-mainland qwen3-coder-plus table rows carry a batch badge (verified against raw page HTML across all 4 occurrences of this model_id on the page). Batch inference discount is no longer offered for this model as of this verification; prior figures ($0.50/$2.50) are stale. Re-verified 2026-07-13: prices, tiers, and no-batch posture unchanged against alibabacloud.com/help/en/model-studio/billing; still resolves to qwen3-coder-plus-2025-09-23 with no newer dated snapshot; model_id still active. Re-verified 2026-07-19: International (Singapore) tiered prices unchanged ($1.00/$5.00, $1.80/$9.00, $3.00/$15.00, $6.00/$60.00); alias still qwen3-coder-plus-2025-09-23; context caching discount still present, no batch badge; no deprecation notice. Re-verified 2026-08-11: International tiered prices unchanged ($1.00/$5.00, $1.80/$9.00, $3.00/$15.00, $6.00/$60.00); alias still qwen3-coder-plus-2025-09-23, no newer dated snapshot; model_id still active on the billing page; no deprecation notice. Re-verified 2026-09-02 (raw-HTML scrape of the Singapore table): International tiered prices unchanged ($1.00/$5.00, $1.80/$9.00, $3.00/$15.00, $6.00/$60.00); alias still qwen3-coder-plus-2025-09-23; still no batch-discount badge on the International row (context-caching-discount blockquote only); no deprecation notice. Not in scope of this pass's requested 4-row list but re-verified alongside the other Alibaba rows for consistency since it was already pulled from the same source fetch.
- **Alibaba Qwen3.7-Max** — New flagship generation added this pass. `qwen3.7-max` is a rolling alias, currently equivalent to `qwen3.7-max-2026-05-20` per the vendor pricing page (updated Jun 26, 2026); a newer `qwen3.7-max-2026-06-08` dated snapshot is also listed at identical pricing but is not yet the default alias. Flat (non-tiered) DashScope International rate of $2.50 input / $7.50 output per MTok across the full 0-1M token range (no pricing_tiers needed, unlike qwen3-max/qwen3-coder-plus). Cross-verified against OpenRouter (openrouter.ai/qwen/qwen3.7-max), which lists the same $2.50/$7.50 standard rate (its displayed $1.25/$3.75 is an explicitly-labeled 50%-off launch promotion, not the standard price). No batch-inference-discount badge on the International row (Chinese-mainland deployment shows one; not priced here, DashScope International is canonical). Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted. 1M-token context window confirmed on the vendor pricing page; max_output_tokens (65,536) is not published on the vendor pricing/model pages directly and is sourced from third-party trackers (OpenRouter, llm-stats.com) that agree on the figure — confidence: medium reflects this one field, not the price. Text-only modalities (MarkTechPost coverage of the 2026-05-20 launch explicitly notes no image input); native tool/function calling supported (reported over 1,000 sequential tool calls in agentic testing). Hybrid thinking model (Non-Thinking and Thinking modes both billed at the same rate), hence reasoning_tokens_billed: true. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only; not yet confirmed on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. Re-verified 2026-07-13: flat $2.50/$7.50 rate, 1M context, and no-batch-on-International posture unchanged against alibabacloud.com/help/en/model-studio/billing; cross-checked openrouter.ai/api/v1/models (qwen/qwen3.7-max: context_length 1,000,000, top_provider.max_completion_tokens 65,536) — corroborates max_output_tokens beyond the prior third-party-tracker citation. Still no qwen4-generation flagship listed by the vendor as of this pass; qwen3.7-max remains current. Re-verified 2026-07-19: flat $2.50/$7.50 list rate over 0-1M unchanged on the International row; the 'Limited-time 50% off' promo label now appears on the vendor billing page itself (previously observed only on OpenRouter) — promo only, list price recorded here, so last_changed_at not bumped. Alias still resolves to qwen3.7-max-2026-05-20; the qwen3.7-max-2026-06-08 dated snapshot remains listed at identical pricing but is still not the default alias. No deprecation notice; still no qwen4/qwen3.8 flagship listed. Re-verified 2026-08-11: flat $2.50/$7.50 list rate over 0-1M unchanged; alias still resolves to qwen3.7-max-2026-05-20 with qwen3.7-max-2026-06-08 still listed at identical pricing and still not the default; no deprecation notice; still no qwen4/qwen3.8 flagship. Re-verified 2026-09-02 (raw-HTML scrape): flat $2.50/$7.50 list rate over 0-1M unchanged on the International row; alias still resolves to qwen3.7-max-2026-05-20; no deprecation notice. A new `qwen3.8-max` flagship generation launched this pass at $2.00/$6.00 flat (added as a separate row below); qwen3.7-max remains a distinct, separately-priced, still-active SKU on the billing page — not superseded, so `deprecated_at`/`replaced_by_model_id` are not set here.
- **Alibaba Qwen3.7-Plus** — New row, added on the 2026-07-13 refresh. Qwen3.7-Plus is Alibaba's cost-effective multimodal agent model in the Qwen3.7 series (released 2026-05-31/06-02), a step down from qwen3.7-max. `qwen3.7-plus` is a rolling alias, currently equivalent to `qwen3.7-plus-2026-05-26`. DashScope International tiered pricing by input-token bucket: 0-256K = $0.40 input / $1.60 output per MTok (base row rate); 256K-1M = $1.20 input / $4.80 output (captured in pricing_tiers). CHANGE 2026-07-19: 256K-1M tier output rate corrected/moved from $1.60 to $4.80 per MTok — the vendor billing page now unambiguously lists $4.8 output for the 256K-1M International tier (confirmed in two separate fetches of the page), and OpenRouter's models API corroborates via its min_prompt_tokens=256000 pricing override (completion $3.84/MTok = exactly 20% off the $4.80 list, matching the vendor's limited-time 20%-off promo; base-tier override figures likewise match $0.40/$1.60/$1.20 list at 20% off). Structured prices here remain list prices, not promo prices. The 20%-off limited-time promotion now appears on the International row itself (both tiers), not only China-mainland. Context window 1,048,576 tokens (1M) per vendor billing page. max_output_tokens (65,536) not published on the vendor billing page directly; sourced from OpenRouter's public models API (openrouter.ai/api/v1/models, qwen/qwen3.7-plus: top_provider.max_completion_tokens 65536, context_length 1,000,000) and cross-checked against llm-stats.com/api/models/qwen3.7-plus ("up to 65,536 output tokens", 1M context, multimodal: true) — satisfies the two-source rule; confidence: medium reflects this field, not the price. Input modalities text + image ("multimodal interactive hybrid agent": perceives scenes, reads screens/GUIs, writes code from visual references) per OpenRouter architecture (text+image->text) and llm-stats.com; output text-only. Vendor billing page itself does not call out modality explicitly (no separate image-token pricing tier), so supports_vision is sourced from the two secondary trackers, not the primary. Tool calling and structured outputs supported per OpenRouter's supported_parameters (tools/tool_choice, response_format/structured_outputs). Always-on/hybrid thinking (llm-stats.com tags thinking: true); thinking tokens billed at the standard output rate, hence reasoning_tokens_billed: true. Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted rather than guessed. No batch-inference-discount badge found for the International row. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only; not yet confirmed on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. Re-verified 2026-08-11: International tiered prices unchanged ($0.40/$1.60 for 0-256K; $1.20/$4.80 for 256K-1M); alias still resolves to qwen3.7-plus-2026-05-26; no deprecation notice. Re-verified 2026-09-02 (raw-HTML scrape): International tiered list prices unchanged ($0.40/$1.60 for 0-256K; $1.20/$4.80 for 256K-1M, still shown with the limited-time 20%-off promo label — list prices recorded here); alias still resolves to qwen3.7-plus-2026-05-26; no deprecation notice; row still in the vendor's models-page 'recommended' set alongside qwen3.8-max and qwen3.8-flash.
- **Alibaba Qwen3.6-Flash** — New row, added on the 2026-07-13 refresh. Qwen3.6-Flash is Alibaba's fast/efficient tier in the Qwen3.6 series (released 2026-04-16), positioned below the Qwen3.7 generation on price but supporting multimodal input. `qwen3.6-flash` is a rolling alias, currently equivalent to `qwen3.6-flash-2026-04-16`. DashScope International tiered pricing by input-token bucket: 0-256K = $0.25 input / $1.50 output per MTok (base row rate); 256K-1M = $1.00 / $4.00 (captured in pricing_tiers). Context window 1,048,576 tokens (1M) per vendor billing page. max_output_tokens (65,536) not published on the vendor billing page directly; sourced from OpenRouter's public models API (openrouter.ai/api/v1/models, qwen/qwen3.6-flash: top_provider.max_completion_tokens 65536, context_length 1,000,000, architecture text+image+video->text) — satisfies the two-source rule alongside the vendor's own pricing table; confidence: medium reflects this field and modality, not the price. Input modalities text + image + video per OpenRouter; output text-only. Unlike qwen3-max/qwen3.7-max, this SKU carries a 50% batch-inference-discount badge on the International/Singapore row itself (confirmed on the vendor billing page) rather than only the China-mainland row — batch_input_per_mtok_usd/batch_output_per_mtok_usd are nonetheless omitted here because the discount applies per input-token tier and the schema's batch fields are flat (non-tiered), so a single number would misrepresent the 256K/1M split; noted here instead of guessed. Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted. Tool calling and structured outputs supported per OpenRouter's supported_parameters. Hybrid thinking (OpenRouter reasoning.mandatory: false); thinking tokens billed at the standard output rate, hence reasoning_tokens_billed: true. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only; not yet confirmed on Bedrock, Vertex, Together, Fireworks, or Groq as a first-party offering. Re-verified 2026-07-19: International tiered prices unchanged ($0.25/$1.50 for 0-256K; $1.00/$4.00 for 256K-1M); 50% batch-discount badge still on the International row (batch fields still omitted for the tiered-vs-flat reason above); alias still resolves to qwen3.6-flash-2026-04-16; no deprecation notice. Re-verified 2026-08-11: International tiered prices unchanged ($0.25/$1.50 for 0-256K; $1.00/$4.00 for 256K-1M); 50% batch-discount badge still on the International row; alias still qwen3.6-flash-2026-04-16; no deprecation notice. New-launch check this pass surfaced qwen3.5-flash, qwen-flash (rolling family alias), and qwen3-omni-flash on the billing page, none currently in this dataset — deferred: the models catalog page does not expose their context window / max output tokens / knowledge cutoff for a two-source-compliant add in this pass; revisit next refresh. Re-verified 2026-09-02 (raw-HTML scrape of the Singapore table): International tiered prices unchanged ($0.25/$1.50 for 0-256K; $1.00/$4.00 for 256K-1M); 50% batch-discount badge still on the International row; alias still resolves to qwen3.6-flash-2026-04-16; no deprecation notice; qwen3.6-flash is absent from the vendor's models-page 'recommended' set (now qwen3.8-max/qwen3.7-plus/qwen3.8-flash) but remains separately priced with no deprecation banner on the billing page, so not marked deprecated. New-launch check this pass found four verifiable new rows added below: `qwen3.8-max` (flagship successor), `qwen3.8-flash` (fast-tier successor), and the two Groq-cross-check models `qwen3.6-27b`/`qwen3.8-27b` (both genuinely on DashScope International, priced there as primary per the two-source rule; Groq is a second direct host, not the canonical price). Also surfaced but deferred this pass (out of the requested scope and thinly documented): `qwen3.7-flash` (billing page shows tiered $0.03/$0.13 + $0.10/$0.40, but not in the models-page recommended set and possibly already superseded by qwen3.8-flash), `qwen3.6-35b-a3b` (open-weight sibling of qwen3.6-27b, $0.375/$2.25 flat on the Singapore table), and `qwen3.8-2.4t-a95b` (open-weight sibling of qwen3.8-max, Singapore $2.00/$6.00 flat with context caching; China-Beijing $1.65/$4.951) — revisit next refresh.
- **Alibaba Qwen3.8-Max** — New row, added on the 2026-09-02 refresh. Qwen3.8-Max is the new flagship generation, general-availability successor line to qwen3-max/qwen3.7-max (both remain separately priced and active, so not marked deprecated/replaced_by). `qwen3.8-max` is a rolling alias currently equivalent to the dated snapshot `qwen3.8-max-0902`, also listed separately at identical pricing (both confirmed via raw-HTML scrape of the vendor billing page). Flat (non-tiered) DashScope International rate of $2.00 input / $6.00 output per MTok across the full 0-1M token range; no batch-inference-discount badge on the International row (context-caching-discount blockquote only; the China-Beijing row carries both a discounted rate of $1.65/$4.951 and a batch badge — not priced here, DashScope International is canonical). Cross-verified against OpenRouter (openrouter.ai/api/v1/models, qwen/qwen3.8-max: prompt $0.000002/completion $0.000006 per token = exactly $2.00/$6.00, matching the DashScope list price exactly) and against llm-stats.com/api/models/qwen3.8-max, whose per-host provider rows (fireworks, novita) also list $2/$6 at max_output_tokens 131,072 (deepinfra and together show slightly different context/output caps and $1.65-$2.50 rates — those are third-party host repricing, not the DashScope figure used here). Context window and max_output_tokens are not published on the vendor billing/models pages directly; sourced from OpenRouter (context_length 1,000,000, top_provider.max_completion_tokens 131,072) and corroborated by llm-stats.com's fireworks/novita provider rows (131,072 max output on both) — satisfies the two-source rule; confidence: medium reflects these two fields and the modality below, not the price. Input modalities text + image + video per OpenRouter's architecture tag (text+image+video->text) and llm-stats.com (multimodal: true, vision/video tags true) — a departure from the text-only qwen3-max/qwen3.7-max lineage; OpenRouter's per-model architecture tags were spot-checked against qwen3.7-max (correctly tagged text->text) to confirm the tags are model-specific, not a generic template. `reasoning`/`reasoning_effort` present in OpenRouter's supported_parameters confirms hybrid thinking support; the vendor billing table lists a single "Non-Thinking and Thinking modes" price column (same rate for both), hence reasoning_tokens_billed: true. Tool calling and structured outputs supported per OpenRouter's supported_parameters (tools/tool_choice, response_format/structured_outputs). Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted rather than guessed (OpenRouter's own cache figures, e.g. $0.25 read, do not match a clean fraction of the DashScope list price and are not used). Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only at last_verified; also listed on Fireworks, Novita, Together, and DeepInfra as third-party hosts (per llm-stats.com) but not added to deployment_options here — only vendor-direct plus hosts explicitly cross-checked this pass (Groq, on the 27b rows below) are recorded. Not marked `featured` in this pass (flagship candidate, but flipping the flag requires regenerating web/lib/featured.lock.json, which is out of scope for a costs/llm.json-only edit — flagged for the controller as a featured candidate).
- **Alibaba Qwen3.8-Flash** — New row, added on the 2026-09-02 refresh. Qwen3.8-Flash is the fast/efficient tier of the Qwen3.8 series, on the vendor's models-page 'recommended models' list alongside qwen3.8-max and qwen3.7-plus. Flat (non-tiered) DashScope International rate of $0.15 input / $0.47 output per MTok across the full 0-1M token range (confirmed via raw-HTML scrape of the vendor billing page); context-caching-discount blockquote present, no batch-inference-discount badge on the International row. Cross-verified against OpenRouter (openrouter.ai/api/v1/models, qwen/qwen3.8-flash: prompt $0.00000015/completion $0.00000047 per token = exactly $0.15/$0.47, matching the DashScope list price exactly) and against llm-stats.com/api/models/qwen3.8-flash's novita provider row (also $0.15/$0.47 exactly, max_output_tokens 131,072). Context window and max_output_tokens not published on the vendor billing/models pages directly; sourced from OpenRouter (context_length 1,000,000, top_provider.max_completion_tokens 131,072) and llm-stats.com's novita row (131,072) and native_context_length tag (1,000,000) — satisfies the two-source rule; llm-stats.com's own summary tag lists a slightly different 128,000 max-output figure, so confidence: medium reflects this field (and the modality below), not the price. Input modalities text + image + video per OpenRouter's architecture tag and llm-stats.com (multimodal: true, video/vision tags true); output text-only. The vendor billing table's price row for this model lacks the explicit "Non-Thinking mode / Thinking mode" two-column split seen on the qwen3.6-flash/qwen3.7-plus/qwen3.8-max sibling tables (single output-price column instead), but OpenRouter's supported_parameters for this model include `reasoning`/`include_reasoning` (llm-stats.com's description also calls it 'a multimodal reasoning model' with a 256K thinking budget tag), so this is treated as a hybrid-thinking model with reasoning_tokens_billed: true, billed at the standard output rate; the single-column table format is read as a formatting choice rather than evidence of no-thinking support. Tool calling and structured outputs supported per OpenRouter's supported_parameters. Context caching discount available but DashScope does not publish a flat cache_read rate, so cache_read_per_mtok_usd is omitted rather than guessed. Knowledge cutoff not published by the vendor; omitted rather than guessed. Deployment via DashScope (Model Studio) only at last_verified; also listed on Novita as a third-party host per llm-stats.com but not added to deployment_options here, consistent with how this dataset prices Alibaba rows at the DashScope/vendor-direct rate.
- **Alibaba Qwen3.6-27B** — New row, added on the 2026-09-02 refresh, triggered by a cross-check against a sibling provider pass that found `qwen/qwen3.6-27b` hosted on Groq (console.groq.com/docs/models) at $0.60/$3.00 per MTok. Confirmed this is a genuine Alibaba release with its own DashScope International listing, not a Groq-exclusive open-weight drop: `qwen3.6-27b` (a dense 27B open-weight model, Apache 2.0, released 2026-04-21 per llm-stats.com) is priced on the Singapore table at flat $0.60 input / $3.60 output per MTok, 0<Token≤256K, with the Non-Thinking and Thinking-mode columns both showing the identical $3.60 output rate (confirmed via raw-HTML scrape of the vendor billing page; no context-caching or batch-discount blockquote present on this row unlike the flagship SKUs). Per the two-source rule and the runbook's 'DashScope is canonical' convention, this DashScope price ($0.60/$3.60) is used as the primary/structured price, not Groq's ($0.60/$3.00 — input matches, output is 16.7% cheaper on Groq). Groq is recorded as a second direct host in deployment_options rather than as an aggregator (it serves its own inference, distinct from OpenRouter's routing). Cross-verified against OpenRouter (openrouter.ai/api/v1/models, qwen/qwen3.6-27b: prompt $0.0000006/completion $0.0000036 per token = exactly $0.60/$3.60, matching the DashScope list price exactly, not Groq's) and against llm-stats.com/api/models/qwen3.6-27b's novita provider row (input/output $0.6/$3.6 exactly, max_output_tokens 65,536, max_input_tokens 262,144) — three independent sources (DashScope, OpenRouter, llm-stats/Novita) agree on both price and context window. max_output_tokens: llm-stats.com's Novita provider row gives a clean 65,536 figure; OpenRouter's top_provider.max_completion_tokens reports 235,929 (= 0.9 x 262,144, a Groq/OpenRouter platform-side completion cap rather than a model spec — the sibling qwen3.6-35b-a3b entry shows the identical 235,929 figure, confirming it is not model-specific) — 65,536 is used as the more credible model-level figure; confidence: medium reflects this reconciliation and the modality below, not the price. Input modalities text + image + video per OpenRouter's architecture tag and llm-stats.com (multimodal: true, vision/thinking tags true; video tag present on OpenRouter, though llm-stats.com's Novita provider row shows video: false for that specific host — output modality text-only on all sources). Groq's own listing caps context at 131,072 tokens and max completion at 16,384 tokens — these are Groq's deployment-specific limits, not the model's native 262,144-token DashScope context, and are not used for context_window/max_output_tokens here (DashScope/OpenRouter/llm-stats are canonical). Tool calling and structured outputs supported per OpenRouter's supported_parameters (tools/tool_choice, response_format/structured_outputs, reasoning/reasoning_effort). Context caching discount not advertised on this row's DashScope table entry; cache_read_per_mtok_usd omitted. Knowledge cutoff not published by the vendor or the secondary sources (llm-stats.com reports null); omitted rather than guessed. Open-weight release date 2026-04-21 per llm-stats.com.
- **Alibaba Qwen3.8-27B** — New row, added on the 2026-09-02 refresh, triggered by the same Groq cross-check as qwen3.6-27b: a sibling provider pass found `qwen/qwen3.8-27b` hosted on Groq at $0.80/$4.00 per MTok. Confirmed this is a genuine Alibaba release with its own DashScope International listing: `qwen3.8-27b` (an open-weight dense vision-language model, per OpenRouter's description) is priced on the Singapore table at flat $0.50 input / $3.00 output per MTok, 0<Token≤1M, Non-Thinking and Thinking modes at the same rate, with a context-caching-discount blockquote (confirmed via raw-HTML scrape of the vendor billing page; no batch-discount badge on the International row). Per the two-source rule and the runbook's 'DashScope is canonical' convention, this DashScope price ($0.50/$3.00) is used as primary — notably cheaper than Groq's ($0.80/$4.00), the opposite direction from the usual pattern where a US inference host undercuts the vendor-direct rate. Groq is recorded as a second direct host in deployment_options (its own inference, not OpenRouter-routed). OpenRouter (openrouter.ai/api/v1/models, qwen/qwen3.8-27b) lists prompt $0.000000425/completion $0.00000255 per token = $0.425/$2.55, a clean 15%-off pass-through of the $0.50/$3.00 DashScope list price (same technique seen on the qwen3.7-plus row's promo-pricing note) — corroborates the DashScope figure rather than contradicting it; OpenRouter's price is not itself used as the structured price. Context window and max_output_tokens: OpenRouter reports context_length 1,000,000 / top_provider.max_completion_tokens 131,072, matching the DashScope billing table's '0<Token≤1M' pricing ceiling; llm-stats.com's friendli provider row instead shows max_input_tokens 262,144 / max_output_tokens 131,072 — the max_output figure agrees across both secondary sources (131,072, used here), but the context-window figure conflicts (262,144 on friendli vs 1,000,000 on OpenRouter and the DashScope pricing ceiling); 1,048,576 is used as context_window since it agrees with both the primary DashScope source and OpenRouter, and matches the dataset's convention of representing '1M' as 1,048,576 on sibling rows (qwen3.7-max, qwen3.7-plus, qwen3.6-flash); confidence: medium reflects this reconciliation and the modality below, not the price. Input modalities text + image + video per OpenRouter's architecture tag and llm-stats.com (multimodal: true, vision/video/thinking tags true); output text-only. Tool calling and structured outputs supported per OpenRouter's supported_parameters (tools/tool_choice, response_format/structured_outputs, reasoning/reasoning_effort). Knowledge cutoff not published by the vendor or the secondary sources; omitted rather than guessed.
- **Perplexity Sonar** — Perplexity native rates: $1.00 input / $1.00 output per MTok. Web search is built into the API as a first-class capability rather than a user-defined tool, so total cost per query = token costs + a per-request fee. per_request_usd captures the low-context tier ($5 / 1,000 requests = $0.005); medium and high search-context tiers add $8 / $12 per 1,000 requests respectively (not captured structurally — single-value field). Perplexity does not publish a separate per-search fee for this SKU (per-search metering applies to `sonar-deep-research`), so per_search_usd is omitted. 128,000-token context window; max_output_tokens 8,000 per Perplexity's documented Sonar limits. Sonar (released January 2025) is built on a fine-tuned Llama 3.3 70B base optimized for web-grounded question answering with inline citations. Knowledge cutoff is intentionally omitted: Sonar fetches the live web at query time, so a static cutoff date does not meaningfully describe its answer space. Not a reasoning SKU (no chain-of-thought tokens), hence reasoning_tokens_billed is false. supports_tool_use is set conservatively to false because web search — the model's primary capability — is exposed as a built-in feature of the Sonar endpoint, not a user-defined tool; Perplexity recommends its separate Agent API for production tool-using agents. Structured outputs supported via response_format. Perplexity API is native-only (no Bedrock / Vertex / Azure / Together / Fireworks / Groq first-party deployment). Cache and batch APIs are not published. Re-verified 2026-08-11: $1/$1 per MTok and $5/$8/$12 per-1K-request search-context tiers unchanged on the vendor pricing page; model still listed as active on the Sonar models page. Re-verified 2026-09-02: $1/$1 per MTok and $5/$8/$12 per-1K-request search-context tiers unchanged; model still active on both the pricing page and the Sonar models page. Heads-up for the next pass: both pages now carry the banner "Sonar Chat Completions is now Agent API. Sonar will be supported until September 27, 2026" — not treated as deprecated here (still priced, still listed as active, no successor model_id given), but the sunset date is close enough to the next refresh cycle that it should be re-checked then; do not set deprecated_at speculatively ahead of an actual retirement.
- **Perplexity Sonar Pro** — Perplexity native rates: $3.00 input / $15.00 output per MTok. Web search is built into the API, not exposed as a user-defined tool; total cost per query = token costs + a per-request fee. per_request_usd captures the low-context tier ($6 / 1,000 requests = $0.006); medium and high search-context tiers add $10 / $14 per 1,000 requests respectively (not captured structurally — single-value field). 200,000-token context window (largest of the Sonar family); max_output_tokens 8,000 per Perplexity's documented Sonar limits. Positioned as Perplexity's advanced search SKU for complex multi-source queries and follow-ups. Knowledge cutoff intentionally omitted: Sonar Pro fetches the live web at query time. Not a reasoning SKU, hence reasoning_tokens_billed is false. supports_tool_use set conservatively to false because Perplexity exposes web search as the built-in capability and recommends the separate Agent API for production tool-using agents. Structured outputs supported via response_format. Perplexity API is native-only (no Bedrock / Vertex / Azure / Together / Fireworks / Groq first-party deployment). Cache and batch APIs are not published. Re-verified 2026-08-11: $3/$15 per MTok and $6/$10/$14 per-1K-request search-context tiers unchanged on the vendor pricing page; model still listed as active on the Sonar models page. Re-verified 2026-09-02: $3/$15 per MTok and $6/$10/$14 per-1K-request search-context tiers unchanged; model still active on both the pricing page and the Sonar models page. Heads-up for the next pass: both pages now carry the banner "Sonar Chat Completions is now Agent API. Sonar will be supported until September 27, 2026" — not treated as deprecated here (still priced, still listed as active, no successor model_id given), but the sunset date is close enough to the next refresh cycle that it should be re-checked then; do not set deprecated_at speculatively ahead of an actual retirement.
- **Perplexity Sonar Reasoning Pro** — Perplexity native rates: $2.00 input / $8.00 output per MTok. Web search is built into the API; total cost per query = token costs + a per-request fee. per_request_usd captures the low-context tier ($6 / 1,000 requests = $0.006); medium and high search-context tiers add $10 / $14 per 1,000 requests respectively (not captured structurally — single-value field). 128,000-token context window; max_output_tokens 8,000 per Perplexity's documented Sonar limits. Reasoning SKU built on DeepSeek R1 with Chain-of-Thought; responses include a leading `<think>` reasoning block followed by the answer, and those reasoning tokens are billed at the output rate, hence reasoning_tokens_billed: true. Knowledge cutoff intentionally omitted: Sonar Reasoning Pro fetches the live web at query time. supports_tool_use set conservatively to false because Perplexity exposes web search as the built-in capability and recommends the separate Agent API for production tool-using agents. Structured outputs supported via response_format. Perplexity API is native-only (no Bedrock / Vertex / Azure / Together / Fireworks / Groq first-party deployment). Cache and batch APIs are not published. Note: the older `sonar-reasoning` SKU is no longer listed in Perplexity's current model lineup at last_verified. Re-verified 2026-08-11: $2/$8 per MTok and $6/$10/$14 per-1K-request search-context tiers unchanged on the vendor pricing page; model still listed as active on the Sonar models page. `sonar-deep-research` remains the only other SKU in the current lineup; not added as a row because its schema-required max_output_tokens is not published by Perplexity (model card and OpenRouter listing both omit it) — deferred rather than guessed. Re-verified 2026-09-02: $2/$8 per MTok and $6/$10/$14 per-1K-request search-context tiers unchanged; model still active on both the pricing page and the Sonar models page. `sonar-deep-research` still lacks a published max_output_tokens (pricing page now shows input/output/citation/reasoning/search-query rates for it but no token-limit field), so it remains deferred. Heads-up for the next pass: both pages now carry the banner "Sonar Chat Completions is now Agent API. Sonar will be supported until September 27, 2026" — not treated as deprecated here (still priced, still listed as active, no successor model_id given), but the sunset date is close enough to the next refresh cycle that it should be re-checked then; do not set deprecated_at speculatively ahead of an actual retirement.
- **Anthropic Claude Opus 4.6** — New row this refresh: fully active legacy listing that was present on Anthropic's pricing page and models overview but previously missing from this dataset (sits between Opus 4.5 and 4.7). Pricing identical to Opus 4.5/4.7/4.8. Cache hit is 0.1x base input ($0.50/MTok); 5-minute cache write is 1.25x ($6.25/MTok); 1-hour cache write is 2x ($10/MTok). Batch API discounts both input and output by 50%. 1M context window at standard pricing, 128k max output; uses the pre-4.7 tokenizer (~750k words per 1M tokens). Supports both extended thinking and adaptive thinking; thinking output tokens billed at the output rate. Fast mode is no longer available on this model as of 2026-06-29 (speed:'fast' requests run and bill at standard rates). Reliable knowledge cutoff May 2025; training data cutoff Aug 2025. last_changed_at is the launch date 2026-02-05, INFERRED from Anthropic's tentative retirement date 'not sooner than February 5, 2027' (launch + 1 year pattern confirmed against Fable 5, Opus 4.8, and Sonnet 5); prices verified as current, not historical. Bedrock ID recorded as shown in the models overview ('anthropic.claude-opus-4-6-v1').
- **Anthropic Claude Opus 4.1** — RETIRED on the Claude API (and Claude Platform on AWS) on 2026-08-05 as scheduled; requests to claude-opus-4-1-20250805 on those surfaces now fail. Still available on Amazon Bedrock (status: deprecated, not retired, per the Bedrock legacy model table) and Google Cloud Vertex AI under their own partner retirement schedules — deployment_options narrowed to bedrock/vertex only, 'native' removed. Recommended replacement claude-opus-4-8. Pricing unchanged and still listed for reference on Anthropic's pricing page: $15/$75 base; cache hit $1.50/MTok (0.1x); 5-minute cache write $18.75/MTok (1.25x); 1-hour cache write $30/MTok (2x); batch $7.50/$37.50 (50% off). 200k context window, 32k max output. Supports extended thinking (not adaptive); thinking output tokens billed at the output rate. Reliable knowledge cutoff Jan 2025; training data cutoff Mar 2025. Cross-verified against platform.claude.com/docs/en/about-claude/pricing and platform.claude.com/docs/en/about-claude/model-deprecations, and Amazon Bedrock's legacy model table, on 2026-09-02.
- **Google Gemini 3.6 Flash** — New model added 2026-07-27. Context window and max_output_tokens (65,536) confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.6-flash. Thinking is supported; thinking tokens billed at the output rate, hence reasoning_tokens_billed: true. Multimodal input (text, image, audio, video per the model documentation); text-only output. Tool use and structured outputs supported. Caching and batch APIs available. Deployment via native API only at this verification. Knowledge cutoff and release date not published by Google; omitted rather than guessed. No long-context tiering published (unlike gemini-2.5-pro). Free tier published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified against secondary source: ai.google.dev/gemini-api/docs/models/gemini-3.6-flash for context, output token limit, modalities, and capabilities. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: ai.google.dev/pricing now lists an introductory-discount price of $0.75 input / $3.75 output / $0.075 cache-read / $0.50 cache-storage-per-hour / $0.375 batch-input / $1.875 batch-output through 2026-12-31, reverting to the previous rate ($1.50/$7.50/$0.15/$1.00/$0.75/$3.75) on 2027-01-01; input/output/cache/batch fields updated to the currently-billed discounted rate and last_changed_at bumped accordingly. Cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 3.5 Flash-Lite** — New model added 2026-07-27. Prices confirmed from ai.google.dev/pricing: $0.30 input / $2.50 output per MTok; batch API at 50% off ($0.15/$1.25). Context window and max_output_tokens (65,536) confirmed via ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite. Thinking is supported; thinking tokens billed at the output rate, hence reasoning_tokens_billed: true. Multimodal input (text, image, audio, video per the model documentation); text-only output. Tool use and structured outputs supported. Batch API available. Context caching not available for this model (per ai.google.dev/pricing, which omits cache pricing for this SKU unlike gemini-3.6-flash). Deployment via native API only at this verification. Knowledge cutoff and release date not published by Google; omitted rather than guessed. No long-context tiering published. Free tier published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified against secondary source: ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite for context, output token limit, modalities, caching availability, and capabilities. Pricing is the primary-source figure pending broader aggregator publication. Re-verified 2026-08-11: prices, tiers, context window, and modalities unchanged against ai.google.dev/pricing; cross-checked ai.google.dev/gemini-api/docs/models. Re-verified 2026-09-02: prices unchanged against ai.google.dev/pricing (explicitly re-confirmed no context-caching price is published for this SKU, unlike gemini-3.6-flash/gemini-3.7-flash); audio input confirmed billed at the same standard rate as text/image/video (no separate audio premium); cross-checked ai.google.dev/gemini-api/docs/models.
- **Google Gemini 3.7 Flash** — New GA model, added 2026-09-02; released August 2026 as "the next iteration in the Gemini 3 series" and described by Google as the latest and most capable Flash model, superseding gemini-3.6-flash without deprecating it. Standard pricing carries an introductory discount through 2026-12-31 ($0.75 input / $3.75 output / $0.075 cache-read / $0.50 cache-storage-per-hour / $0.375 batch-input / $1.875 batch-output per ai.google.dev/pricing), reverting to $1.50/$7.50/$0.15/$1.00/$0.75/$3.75 on 2027-01-01; fields capture the currently-billed discounted rate, identical in structure to sibling gemini-3.6-flash. Thinking supported at low/medium/high (minimal not offered); thinking tokens billed at the output rate, hence reasoning_tokens_billed=true. Multimodal input (text, image, audio, video, PDF); text-only output. Tool use, structured outputs, and context caching supported. Deployment via native API only at this verification; Vertex AI availability not confirmed. No long-context (>200k token) pricing tier published. Knowledge cutoff not published by Google; omitted rather than guessed. Free tier (AI Studio) published as per-minute RPM/TPM only, not per-day; free_tier omitted. Cross-verified against secondary source: ai.google.dev/gemini-api/docs/models/gemini-3.7-flash for context window, max output tokens, modalities, tool-use/structured-output support, and GA status.


## Speech-to-text

| Provider | Model | $/min | $/min batch | Streaming | Realtime | Languages | Diarization | Verified | Source |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Deepgram | Nova-3 Monolingual (`nova-3-monolingual`) | $0.004800 | $0.004300 | yes | yes | en | extra-cost | 2026-09-02 | [link](https://deepgram.com/pricing) |
| Deepgram | Nova-3 Multilingual (`nova-3-multilingual`) | $0.005800 | $0.005200 | yes | yes | 61+ | extra-cost | 2026-09-02 | [link](https://deepgram.com/pricing) |
| Deepgram | Nova-3 Medical (`nova-3-medical`) | $0.004800 | $0.004300 | yes | yes | en, en-US, en-AU, en-CA, en-GB, en-IE, en-IN, en-NZ | extra-cost | 2026-09-02 | [link](https://deepgram.com/learn/introducing-nova-3-medical-speech-to-text-api) |
| Deepgram | Flux (English) (`flux-general-en`) | $0.006500 | — | yes | yes | en | — | 2026-09-02 | [link](https://deepgram.com/pricing) |
| Deepgram | Flux (Multilingual) (`flux-general-multi`) | $0.007800 | — | yes | yes | en, es, fr, de, hi, ru, pt, ja, it, nl | — | 2026-09-02 | [link](https://deepgram.com/pricing) |
| Deepgram | Whisper Cloud (Large) (`whisper-large`) | $0.004800 | — | no | no | 99+ | — | 2026-09-02 | [link](https://deepgram.com/pricing) |
| AssemblyAI | Universal-2 (`universal-2`) | $0.002500 | — | no | no | 99+ | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-Streaming (`universal-streaming`) | $0.002500 | — | yes | yes | en | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-Streaming English (`universal-streaming-english`) | $0.002500 | — | yes | yes | en | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-3.5 Pro (`universal-3-pro`) | $0.003500 | — | no | no | en, es, fr, de, it, pt, ar, da, nl, fi, he, hi, ja, zh, no, sv, tr, vi | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-3.5 Pro (`universal-3-5-pro`) | $0.003500 | — | no | no | en, es, fr, de, it, pt, ar, da, nl, fi, he, hi, ja, zh, no, sv, tr, vi | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-3 Pro Streaming (`universal-3-pro-streaming`) | $0.007500 | — | yes | yes | en, es, fr, de, it, pt, ar, da, nl, fi, he, hi, ja, zh, no, sv, tr, vi | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-3.5 Pro Realtime (`u3-rt-pro`) | $0.007500 | — | yes | yes | en, es, fr, de, it, pt, ar, da, nl, fi, he, hi, ja, zh, no, sv, tr, vi | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-3.5 Pro Realtime (`universal-3-5-pro`) | $0.007500 | — | yes | yes | en, es, fr, de, it, pt, ar, da, nl, fi, he, hi, ja, zh, no, sv, tr, vi | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Universal-Streaming Multilingual (`universal-streaming-multilingual`) | $0.002500 | — | yes | yes | en, es, pt, de, fr, it | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| AssemblyAI | Whisper-Streaming (`whisper-streaming`) | $0.005000 | — | yes | yes | 99+ | extra-cost | 2026-09-02 | [link](https://www.assemblyai.com/pricing) |
| Cartesia | Ink Whisper (`ink-whisper`) | $0.003000 | $0.001500 | yes | yes | 99+ | unsupported | 2026-09-02 | [link](https://cartesia.ai/pricing) |
| Cartesia | Ink 2 (`ink-2`) | $0.009000 | — | yes | yes | en | — | 2026-09-02 | [link](https://cartesia.ai/pricing) |
| OpenAI | Whisper (`whisper-1`) | $0.006000 | — | no | no | 99+ | unsupported | 2026-09-02 | [link](https://developers.openai.com/api/docs/pricing) |
| OpenAI | GPT-4o Transcribe (`gpt-4o-transcribe`) | $0.006000 | — | no | no | 99+ | unsupported | 2026-09-02 | [link](https://developers.openai.com/api/docs/pricing) |
| OpenAI | GPT-4o mini Transcribe (`gpt-4o-mini-transcribe`) | $0.003000 | — | no | no | 99+ | unsupported | 2026-09-02 | [link](https://developers.openai.com/api/docs/pricing) |
| OpenAI | GPT-4o Transcribe Diarize (`gpt-4o-transcribe-diarize`) | $0.006000 | — | no | no | 99+ | included | 2026-09-02 | [link](https://developers.openai.com/api/docs/pricing) |
| OpenAI | GPT Realtime Whisper (`gpt-realtime-whisper`) | $0.017000 | — | yes | yes | 99+ | unsupported | 2026-09-02 | [link](https://developers.openai.com/api/docs/pricing) |
| OpenAI | GPT Transcribe (`gpt-transcribe`) | $0.004500 | — | no | no | 99+ | unsupported | 2026-09-02 | [link](https://developers.openai.com/api/docs/pricing) |
| OpenAI | GPT Live Transcribe (`gpt-live-transcribe`) | $0.017000 | — | yes | yes | 99+ | unsupported | 2026-09-02 | [link](https://developers.openai.com/api/docs/pricing) |
| Groq | Whisper V3 Large (Groq) (`whisper-large-v3`) | $0.001850 | — | no | no | 99+ | unsupported | 2026-09-02 | [link](https://console.groq.com/docs/models) |
| Groq | Whisper Large v3 Turbo (Groq) (`whisper-large-v3-turbo`) | $0.000667 | — | no | no | 99+ | unsupported | 2026-09-02 | [link](https://console.groq.com/docs/models) |
| Microsoft Azure | Azure Speech (Real-time) (`azure-speech-realtime`) | $0.016667 | — | yes | yes | 100+ | extra-cost | 2026-09-02 | [link](https://azure.microsoft.com/en-us/pricing/details/speech/) |
| Microsoft Azure | Azure Speech (Batch) (`azure-speech-batch`) | $0.003000 | — | no | no | 100+ | included | 2026-09-02 | [link](https://azure.microsoft.com/en-us/pricing/details/speech/) |
| Microsoft Azure | Azure Speech (Fast) (`azure-speech-fast`) | $0.006000 | — | no | no | 100+ | included | 2026-09-02 | [link](https://azure.microsoft.com/en-us/pricing/details/speech/) |
| Google | Chirp 2 (`chirp_2`) | $0.016000 | — | yes | yes | 20+ | included | 2026-09-02 | [link](https://cloud.google.com/speech-to-text/pricing) |
| Google | Chirp 3 (`chirp_3`) | $0.016000 | — | yes | yes | 98+ | included | 2026-09-02 | [link](https://cloud.google.com/speech-to-text/pricing) |
| Speechmatics | Speechmatics Enhanced (`enhanced`) | $0.002150 | — | yes | yes | 56+ | included | 2026-09-02 | [link](https://www.speechmatics.com/pricing) |
| Speechmatics | Speechmatics Melia 1 (`melia-1`) | $0.002150 | — | no | no | 68+ | included | 2026-09-02 | [link](https://www.speechmatics.com/pricing) |
| Speechmatics | Speechmatics Standard (`standard`) | $0.002150 | — | yes | yes | 56+ | included | 2026-09-02 | [link](https://www.speechmatics.com/pricing) |
| Rev.ai | Rev.ai Whisper Fusion (`whisper-fusion`) | $0.005000 | — | yes | yes | en | included | 2026-09-02 | [link](https://www.rev.ai/pricing) |
| Rev.ai | Rev.ai Reverb (`reverb`) | $0.003300 | — | no | no | en | included | 2026-09-02 | [link](https://www.rev.ai/pricing) |
| Gladia | Gladia Solaria-1 (`solaria-1`) | $0.012500 | $0.010170 | yes | yes | 100+ | included | 2026-09-02 | [link](https://www.gladia.io/pricing) |
| Gladia | Gladia Solaria-3 (`solaria-3`) | $0.010170 | — | no | no | en, fr, de, es, it | included | 2026-09-02 | [link](https://www.gladia.io/pricing) |
| Soniox | Soniox STT Real-time v4 (`stt-rt-v4`) | $0.002000 | — | yes | yes | 60+ | included | 2026-09-02 | [link](https://soniox.com/pricing) |
| Soniox | Soniox STT Real-time v5 (`stt-rt-v5`) | $0.002000 | — | yes | yes | 60+ | included | 2026-09-02 | [link](https://soniox.com/pricing) |
| Soniox | Soniox STT Async v4 (`stt-async-v4`) | $0.001670 | — | no | no | 60+ | included | 2026-09-02 | [link](https://soniox.com/pricing) |
| Soniox | Soniox STT Async v5 (`stt-async-v5`) | $0.001670 | — | no | no | 60+ | included | 2026-09-02 | [link](https://soniox.com/pricing) |

**Notes:**

- **Deepgram Nova-3 Monolingual** — Pay-as-you-go tier (English). Growth tier (volume commitment) is $0.0042/min streaming, $0.0036/min pre-recorded. Diarization add-on is captured in diarization_per_minute_usd. Correction (2026-09-02): price_per_minute_batch_usd restored to $0.0043/min, matching the current deepgram.com/pricing structured data (schema.org Offer entries) and visible pricing table. The 2026-07-02 refresh had mistakenly recorded $0.0077/min here -- that figure is actually the struck-through, non-promotional REGULAR STREAMING rate (labeled 'Regular price' next to the promotional $0.0048/min streaming rate), not a pre-recorded price. Corroborated by third-party aggregators (e.g. convertaudiototext.com, llmreference.com) quoting $0.0043/min batch for Nova-3 mono.
- **Deepgram Nova-3 Multilingual** — Pay-as-you-go tier (multilingual, ~61 languages). Growth tier is $0.0050/min streaming, $0.0043/min pre-recorded. Diarization add-on is captured in diarization_per_minute_usd. Correction (2026-09-02): price_per_minute_batch_usd restored to $0.0052/min, matching the current deepgram.com/pricing structured data (schema.org Offer entries) and visible pricing table. The 2026-07-02 refresh had mistakenly recorded $0.0092/min here -- that figure is actually the struck-through, non-promotional REGULAR STREAMING rate (labeled 'Regular price' next to the promotional $0.0058/min streaming rate), not a pre-recorded price.
- **Deepgram Nova-3 Medical** — Medical-tuned Nova-3 for clinical transcription. English variants only (en, en-US, en-AU, en-CA, en-GB, en-IE, en-IN, en-NZ). Invoked via `model=nova-3-medical` in the Deepgram API. Pricing is not separately listed on the public pricing page; Deepgram's launch announcement quotes $0.0043/min pre-recorded, which matches the Nova-3 Monolingual batch rate (now confirmed current -- see nova-3-monolingual row). Streaming rate assumed equal to Nova-3 Monolingual ($0.0048/min PAYG); verify with Deepgram sales for production commitments.
- **Deepgram Flux (English)** — New conversational speech recognition (CSR) model built for voice agents; adds turn-detection/end-of-turn modeling on top of transcription. Launched free for October 2025 promo; announcement post (https://deepgram.com/learn/introducing-flux-conversational-speech-recognition) does not quote a per-minute price, so pricing is single-sourced from deepgram.com/pricing (Pay-As-You-Go streaming $0.0065/min; Growth tier $0.0057/min), corroborated by a third-party pricing aggregator quoting the same figure. Correction (2026-09-02): the $0.0077/min figure is the struck-through, non-promotional REGULAR STREAMING rate shown next to the promotional $0.0065/min streaming price on deepgram.com/pricing -- it is not a pre-recorded/batch column. The structured pricing data (schema.org Offer entries) on the page lists no Pre-Recorded offer for Flux English at all, confirming price_per_minute_batch_usd should stay omitted here.
- **Deepgram Flux (Multilingual)** — Multilingual conversational speech recognition (CSR) variant of Flux, launched after the English-only October 2025 debut (see press release: https://deepgram.com/learn/deepgram-launches-flux-multilingual-press-release). Covers 10 languages via `language_hint` on the `flux-general-multi` model_id (en, es, fr, de, hi, ru, pt, ja, it, nl). Pay-As-You-Go streaming $0.0078/min; Growth tier $0.0068/min. No pre-recorded/batch price is listed on the pricing page for this row (confirmed again 2026-09-02: no Pre-Recorded Offer entry for Flux Multilingual in the page's structured pricing data).
- **Deepgram Whisper Cloud (Large)** — Deepgram-hosted OpenAI Whisper, pre-recorded only ('Live streaming is not available with Deepgram Whisper Cloud' per docs -- use Nova-3 for streaming). Invoked via `model=whisper-large` (defaults to large-v2); other sizes available (`whisper-tiny`, `whisper-base`, `whisper-small`, `whisper-medium`). Pricing page (schema.org Offer 'Pre-Recorded - Whisper Large') lists $0.0048/min flat across both Pay-As-You-Go and Growth tiers -- no volume discount, unlike the Nova/Flux rows. Officially supports 99 languages per developers.deepgram.com/docs/deepgram-whisper-cloud. Not available on the EU endpoint; docs note Whisper models are 'less scalable' than Deepgram's proprietary models. confidence: medium because the secondary source (developer docs) confirms the model_id and capabilities but not the price itself -- only the pricing page quotes a number.
- **AssemblyAI Universal-2** — Pre-recorded (file-based) only — broadest AssemblyAI language coverage. Published as $0.15/hr. Diarization add-on +$0.02/hr; the pricing page no longer lists the +$0.065/hr experimental diarization tier seen in a prior pass (llms/pricing.md, dated 2026-05-29, still shows it — treated as stale). For real-time use, see universal-3-5-pro (streaming) / universal-streaming-english / universal-streaming-multilingual.
- **AssemblyAI Universal-Streaming** — Superseded by explicit `universal-streaming-english` / `universal-streaming-multilingual` model_ids — the streaming AsyncAPI spec (docs.assemblyai.com) no longer enumerates a bare `universal-streaming` value. Price unchanged at $0.15/hr; treated as a rename, not a price change. confidence:medium because no explicit vendor deprecation notice was found, only enum absence.
- **AssemblyAI Universal-Streaming English** — English-only streaming model, explicit successor id to the former bare `universal-streaming`. Published as $0.15/hr. Higher-tier streaming (Universal-3.5 Pro Realtime / Streaming, model_id `universal-3-5-pro`) is $0.45/hr ($0.0075/min).
- **AssemblyAI Universal-3.5 Pro** — AssemblyAI's highest-accuracy pre-recorded model with native code-switching, documented at 18 languages per assemblyai.com/docs/getting-started/models and assemblyai.com/llms/models.md; unsupported languages fall back to Universal-2. Deprecated: the changelog entry "Deprecated: universal-3-pro and universal for New and Inactive Accounts" (2026-07-10) cut off new/inactive accounts; the entry "Default speech model changing to Universal-3.5 Pro on September 2, 2026" confirms requests pinned to `universal-3-pro` now return an error as of today (2026-09-02) instead of being auto-routed — hard cutover complete. Legacy request value `best` now routes to `universal-3-5-pro` (`nano` still routes to universal-2). Superseded by the `universal-3-5-pro` row at the same $0.21/hr price — a rename, not a price change. Featured flag kept per runbook policy (indexed compare pages stay live with a deprecation banner).
- **AssemblyAI Universal-3.5 Pro** — Canonical async model_id as of 2026-09-02, renamed from `universal-3-pro` (see that deprecated row). Confirmed via three current first-party sources: the docs/getting-started/models secondary source links "Universal-3.5 Pro" to /docs/pre-recorded-audio/universal-3-5-pro, whose code samples across Python/JS/TS/cURL all show `"speech_models": ["universal-3-5-pro"]`; the changelog entry "Default speech model changing to Universal-3.5 Pro on September 2, 2026" states `universal-3-pro` now hard-errors and this id is the new default when no model is pinned; and the pricing page (primary) confirms $0.21/hr unchanged. Same 18-language coverage as the prior `universal-3-pro` row; unsupported languages fall back to Universal-2. This exact model_id string is shared with the streaming row below (`universal-3-5-pro`, streaming:true, realtime:true, $0.45/hr) — AssemblyAI unified the identifier across the async (`/v2/transcript`) and streaming (`wss://streaming.assemblyai.com/v3/ws`) endpoints, confirmed on the dedicated streaming model-selection reference. The two products are billed separately and are distinguished in this dataset by the `streaming` field, not by model_id — first duplicate model_id in this file; flagging for review. A separate Sync API product also uses this id (≤2min/request, $0.45/hr via an `X-AAI-Model: universal-3-5-pro` header at sync.assemblyai.com) but is not represented here — a third row with this model_id and streaming:false at a different price is not representable in the current schema; deferred.
- **AssemblyAI Universal-3 Pro Streaming** — Real-time streaming tier for voice agents, now branded "Universal-3.5 Pro Realtime". Deprecated in this dataset 2026-07-19: for a second consecutive refresh (2026-07-13 and 2026-07-19) the pricing page and the assemblyai.com/llms/models.md quick reference both give `u3-rt-pro` (alias `u3-pro`) as the API identifier, and no vendor source uses the literal string `universal-3-pro-streaming`. This was the repeat confirmation the 2026-07-13 refresh required before renaming; superseded by the `u3-rt-pro` row at the same $0.45/hr price (a rename, not a price change). confidence:medium because this id was never published by the vendor, so there is no explicit vendor deprecation notice for it.
- **AssemblyAI Universal-3.5 Pro Realtime** — Canonical id for the streaming tier this dataset previously tracked as `universal-3-pro-streaming` (GA 2026-03-03 per assemblyai.com/llms/models.md; `u3-pro` kept as a backward-compatible alias that routes here). Deprecated: the current dedicated streaming model-selection reference (docs/streaming/select-the-speech-model) lists only three available streaming models — `universal-3-5-pro` (recommended, also the default), `universal-streaming-english`, `universal-streaming-multilingual` — `u3-rt-pro`/`u3-pro` no longer appear. The streaming migration guide "Universal Streaming to Universal-3.5 Pro Streaming" and the async-migration changelog entry ("u3-pro-rt (Universal-3 Pro Realtime) Redirected to universal-3-5-pro — no action required") corroborate the same target id. Superseded by the `universal-3-5-pro` row at the same $0.45/hr price — a rename, not a price change. confidence:medium because, unlike the async id's explicit hard-error notice, no changelog entry states `u3-rt-pro` itself now errors — only that it redirects, so treat with the same caution as the prior `universal-streaming` rename.
- **AssemblyAI Universal-3.5 Pro Realtime** — Canonical streaming model_id as of 2026-09-02 — unified with the async model. The identical string `universal-3-5-pro` is used both for the async endpoint (see the streaming:false row above, $0.21/hr) and for this streaming WebSocket product ($0.45/hr); this is the first duplicate model_id in this dataset, flagging for review since site tooling assumes model_id is a stable per-row key. Confirmed via the dedicated streaming model-selection reference (docs/streaming/select-the-speech-model: `"speech_model": "universal-3-5-pro"`, marked Recommended and used as the default when the parameter is omitted) and the streaming migration guide "Universal Streaming to Universal-3.5 Pro Streaming". Supersedes `u3-rt-pro` (alias `u3-pro`) — see that row's deprecation. Pricing page (primary) brands this "Universal-3.5 Pro Realtime" at $0.45/hr; current docs call it "Universal-3.5 Pro Streaming" — same product, same price, unchanged from the prior u3-rt-pro row. Real-time inline diarization add-on +$0.12/hr (`speaker_labels: true`), billed on session duration not audio duration.
- **AssemblyAI Universal-Streaming Multilingual** — Multilingual streaming variant covering EN/ES/PT/DE/FR/IT at the same $0.15/hr rate as the English-only universal-streaming-english. Good balance of cost and latency for voice agents.
- **AssemblyAI Whisper-Streaming** — OpenAI Whisper served via AssemblyAI's streaming infrastructure with 99+ language coverage. Published as $0.30/hr as of last verification, but no longer listed on the pricing page, docs/getting-started/models, the llms/models.md quick reference, or the changelog — still absent as of 2026-09-02, reconfirming silent discontinuation. No in-file replacement (no surviving streaming model offers 99+ language coverage); confidence:medium since no explicit vendor deprecation notice was found, only absence across all checked sources.
- **Cartesia Ink Whisper** — model_id corrected from the invented 'ink-1' to the actual API identifier 'ink-whisper' (docs.cartesia.ai STT AsyncAPI spec enumerates only ink-2 and ink-whisper; latest dated snapshot is ink-whisper-2025-06-04). cartesia.ai/pricing no longer publishes an explicit $/min table -- STT is now billed in credits (docs.cartesia.ai/pricing.md: ink-whisper = 1 credit/sec streaming, 1 credit per 2 sec via the batch endpoint) with credit value varying by plan tier. Converting at the Pro-tier rate ($5/mo for 100,000 credits = $0.00005/credit, the same rate implied by the prior $0.003/min figure) yields $0.003/min streaming (unchanged) and $0.0015/min batch (new price_per_minute_batch_usd field, derived not published, hence confidence: medium). Languages widened from '42+' to '99+' per docs.cartesia.ai/build-with-cartesia/stt/older-models. Still the same STT product as the Sonic TTS provider; supports both batch (/stt) and manual realtime websocket. 2026-07-13 re-verify: still active, no deprecation banner; docs.cartesia.ai/build-with-cartesia/stt/older-models confirms 99-language support and ink-whisper-2025-06-04 as latest snapshot unchanged; docs.cartesia.ai/pricing.md credit rates (1 credit/sec streaming, 1 credit/2sec batch) and cartesia.ai/pricing Pro-tier allotment ($5/mo, ~9h16m ink-2 hours implying $0.00005/credit) unchanged, so both prices confirmed unchanged. 2026-07-19 re-verify: docs.cartesia.ai/pricing.md still lists ink-whisper at 1 credit/sec streaming and 1 credit/2sec batch; cartesia.ai/pricing plan allotments unchanged ($5 Pro ~9h16m); no deprecation notice (only 'Older STT Models' placement); prices unchanged. 2026-07-27 re-verify: both sources confirm ink-whisper remains stable with credit rates (1 credit/sec streaming, 1 credit/2sec batch) and Pro-tier plan allotment (~9h16m for $5, implying $0.00005/credit) unchanged; no new Cartesia STT models found; prices unchanged. 2026-08-11 re-verify: docs.cartesia.ai/build-with-cartesia/stt/older-models confirms ink-whisper-2025-06-04 still Stable, 99-language support unchanged, still positioned under 'Older Models' recommending Ink 2 for English turn-detection use cases (no formal deprecation date set); docs.cartesia.ai/pricing.md credit rates (1 credit/sec streaming, 1 credit/2sec batch) unchanged; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; prices unchanged; no new Cartesia STT models found. 2026-09-02 re-verify: docs.cartesia.ai/build-with-cartesia/stt/older-models confirms ink-whisper-2025-06-04 still Stable, 99-language support unchanged; docs.cartesia.ai/pricing.md credit rates (1 credit/sec streaming, 1 credit/2sec batch) unchanged; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; prices unchanged; no new Cartesia STT models found; no deprecation notices.
- **Cartesia Ink 2** — New model launched 2026-05-22 (docs.cartesia.ai/build-with-cartesia/stt/latest): Cartesia's fastest streaming STT with built-in turn detection, replacing the need for separate VAD. English only ('en') as of this refresh; batch (/stt) support listed as 'coming soon', so only the streaming/auto-websocket endpoint is priced here (no price_per_minute_batch_usd yet). Billed at 3 credits/sec (docs.cartesia.ai/pricing.md); converted at the Pro-tier rate of $0.00005/credit ($5/mo for 100,000 credits) = $0.009/min, consistent with the ~9h16m Ink-2 allotment shown for the $5 Pro plan on cartesia.ai/pricing (9.27h => ~$0.009/min). Diarization support unconfirmed in docs, field omitted pending verification. Not a replacement for ink-whisper -- both remain listed as stable/active models. 2026-07-13 re-verify: still active/stable, still English-only, batch endpoint still listed as not yet available, turn-detection lifecycle events (turn.start/update/eager_end/resume/end) confirmed live on docs.cartesia.ai/build-with-cartesia/stt/latest; credit rate (3 credits/sec) and Pro-tier allotment (~9h16m for $5) unchanged, so $0.009/min confirmed unchanged. 2026-07-19 re-verify: still latest/stable on docs.cartesia.ai/build-with-cartesia/stt/latest, still English-only, batch (/stt) still not available per docs.cartesia.ai/pricing.md (3 credits/sec streaming unchanged); cartesia.ai/pricing Pro allotment ~9h16m/$5 unchanged; price unchanged; no new Cartesia STT models found on either source. 2026-07-27 re-verify: still latest/stable on docs.cartesia.ai/build-with-cartesia/stt/latest, still English-only, batch still not available per docs.cartesia.ai/pricing.md (3 credits/sec streaming unchanged); cartesia.ai/pricing Pro allotment ~9h16m/$5 unchanged; price unchanged; no new Cartesia STT models; no deprecation notices. 2026-08-11 re-verify: docs.cartesia.ai/build-with-cartesia/stt/latest confirms Ink 2 still Stable, still English-only, turn-detection lifecycle events unchanged; docs.cartesia.ai/pricing.md batch (/stt) still not available for ink-2, streaming rate still 3 credits/sec; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; price unchanged; no new Cartesia STT models found; no deprecation notices. 2026-09-02 re-verify: docs.cartesia.ai/build-with-cartesia/stt/latest confirms Ink 2 still Stable, still English-only, released 2026-05-22, turn-detection lifecycle events unchanged; docs.cartesia.ai/pricing.md batch (/stt) still not available for ink-2, streaming rate still 3 credits/sec; cartesia.ai/pricing Pro-tier allotment (~9h16m for $5) unchanged; price unchanged; no new Cartesia STT models found; no deprecation notices.
- **OpenAI Whisper** — Billed to the nearest second. Pre-recorded only. Price unchanged at $0.006/min. openai.com/api/pricing/ still 403s and platform.openai.com/docs/pricing 301-redirects to developers.openai.com/api/docs/pricing (developers.openai.com/api/docs/models likewise 301s from platform.openai.com/docs/models). 2026-08-11 re-verify: as of that pass whisper-1 no longer appeared on the /api/docs/models catalog listing at all — rate re-confirmed directly on the model detail page developers.openai.com/api/docs/models/whisper-1 ('Default snapshot: whisper-1', $0.006/min, no deprecation banner), cross-checked against /api/docs/deprecations (whisper-1 not present at that time). 2026-09-02 re-verify: price still $0.006/min (unchanged) on developers.openai.com/api/docs/pricing and the model detail page. BUT /api/docs/deprecations now (as of the 2026-08-26 entry) lists whisper-1 under 'Transcription models': "On August 26, 2026, we notified developers using whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize of their deprecation and removal from the API on February 26, 2027", table row 'Feb 26, 2027 | whisper-1 | gpt-live-transcribe or gpt-transcribe'. deprecated_at set to the 2026-08-26 announcement date (matching this dataset's convention, e.g. the llm.json gpt-5 row); shutdown is 2027-02-26. replaced_by_model_id set to gpt-transcribe (same non-realtime /v1/audio/transcriptions file-upload shape and flat per-minute pricing as this row) — vendor also names gpt-live-transcribe as an acceptable replacement for realtime use cases.
- **OpenAI GPT-4o Transcribe** — GPT-4o-powered transcription model served via the same /v1/audio/transcriptions endpoint as whisper-1; higher quality with prompting support for improved accuracy. Billed per-token ($2.50/1M input audio tokens, $10.00/1M output tokens); developers.openai.com/api/docs/pricing publishes this as an 'Estimated cost' of $0.006/minute (confidence: medium — the real bill tracks token usage, not a flat per-minute rate like whisper-1). 25 MB file-size limit. Supports `stream=true` for incremental text output on an already-uploaded file, which is not the same as live/realtime audio-in. 2026-08-11 re-verify: still active on both sources, price unchanged. 2026-09-02 re-verify: price still $0.006/min estimated ($2.50/$10.00 per 1M tokens), unchanged, on developers.openai.com/api/docs/pricing. Now listed on /api/docs/deprecations 'Transcription models' entry (announced 2026-08-26): removal from the API on 2027-02-26, vendor-recommended replacement gpt-live-transcribe or gpt-transcribe. deprecated_at set to the 2026-08-26 announcement date; replaced_by_model_id set to gpt-transcribe (same file-upload /v1/audio/transcriptions shape as this row).
- **OpenAI GPT-4o mini Transcribe** — Cost-efficient sibling of gpt-4o-transcribe with the same prompting and (non-realtime) streaming-of-output capabilities; same /v1/audio/transcriptions endpoint and 25 MB file-size limit. Billed per-token ($1.25/1M input audio tokens, $5.00/1M output tokens); developers.openai.com/api/docs/pricing publishes this as an 'Estimated cost' of $0.003/minute (confidence: medium, same token-vs-flat-rate caveat as gpt-4o-transcribe). 2026-08-11 re-verify: base model_id `gpt-4o-mini-transcribe` still active, price unchanged; only a snapshot-rotation deprecation was on file at that point (`gpt-4o-mini-transcribe-2025-03-20` -> `gpt-4o-mini-transcribe-2025-12-15`, shutdown 2027-01-20), which did not touch this row's base-alias model_id. 2026-09-02 re-verify: price still $0.003/min estimated ($1.25/$5.00 per 1M tokens), unchanged. /api/docs/deprecations now carries a base-model entry too (announced 2026-08-26, distinct from the earlier snapshot-rotation entry): removal from the API on 2027-02-26, vendor-recommended replacement gpt-live-transcribe or gpt-transcribe. deprecated_at set to the 2026-08-26 announcement date; replaced_by_model_id set to gpt-transcribe (same file-upload /v1/audio/transcriptions shape as this row).
- **OpenAI GPT-4o Transcribe Diarize** — Speaker-aware transcription variant for meeting recordings / multi-speaker audio: emits `transcript.text.segment` events with speaker labels (up to 4 known-speaker reference clips accepted), no separate diarization fee — priced identically to gpt-4o-transcribe at $2.50/$10.00 per 1M audio tokens, 'Estimated cost' $0.006/minute (confidence: medium, same token-vs-flat-rate caveat). Hidden/collapsed on the main pricing table. Requires `chunking_strategy` (auto or VAD config) for audio longer than 30 seconds; does not support the `prompt` or `logprobs` parameters. 25 MB file-size limit. 2026-08-11 re-verify: still active, price unchanged, not on /api/docs/deprecations at that time. 2026-09-02 re-verify: price still $0.006/min estimated, unchanged, on developers.openai.com/api/docs/pricing. Now listed on /api/docs/deprecations 'Transcription models' entry (announced 2026-08-26): removal from the API on 2027-02-26, vendor-recommended replacement gpt-live-transcribe or gpt-transcribe. deprecated_at set to the 2026-08-26 announcement date; replaced_by_model_id set to gpt-transcribe (no diarize successor announced — replacement lineup drops the diarization feature; noting this as a functional regression for any consumer relying on diarization: included).
- **OpenAI GPT Realtime Whisper** — New row this refresh (deferred 2026-07-13 pending STT-vs-realtime categorization; vendor now lists it plainly as 'Streaming speech-to-text model for realtime transcription', so it belongs here). Low-latency streaming transcription for live audio (microphone/call/media stream) via /v1/realtime and /v1/realtime/transcription_sessions. Flat $0.017 per minute of audio duration — not token-billed; price confirmed on both developers.openai.com/api/docs/pricing and the model detail page developers.openai.com/api/docs/models/gpt-realtime-whisper. confidence: medium because the supported-language list is not explicitly published for this model — '99+' is inherited from the Whisper family (speech-to-text guide's language section covers the file-based endpoints only). No function calling / structured outputs / fine-tuning. 2026-08-11 re-verify: still active, price unchanged ($0.017/min on both sources); note a new `gpt-live-transcribe` model launched this pass at the identical $0.017/min live-transcription rate (added as its own row — distinct model_id, not a rename of this one). 2026-09-02 re-verify: still active, price unchanged ($0.017/min). Explicitly named on /api/docs/deprecations as a recommended *replacement* target for the newly-deprecated whisper-1 / gpt-4o-transcribe family (not itself deprecated).
- **OpenAI GPT Transcribe** — New row this refresh: flat-rate file transcription model on the same /v1/audio/transcriptions endpoint as gpt-4o-transcribe, billed at a straight $0.0045/minute of audio duration rather than per-token (contrast gpt-4o-transcribe's token-billed 'Estimated cost' of $0.006/min). Price confirmed on both developers.openai.com/api/docs/pricing and the model detail page developers.openai.com/api/docs/models/gpt-transcribe ('Default snapshot: gpt-transcribe', active). Streaming/realtime modelled as false to match this dataset's convention for gpt-4o-transcribe: the model supports `stream=true` incremental text output on an already-uploaded file and can be used from `/v1/realtime/transcription_sessions`, but that is not continuous live audio-in the way `gpt-live-transcribe`/`gpt-realtime-whisper` are. confidence: medium because the supported-language list is not explicitly broken out for this model on either source — '99+' inherited from the rest of the OpenAI transcription family pending a dedicated languages page. 2026-09-02 re-verify: still active, price unchanged ($0.0045/min). Now the primary vendor-recommended replacement on /api/docs/deprecations for the newly-deprecated whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize rows (see their replaced_by_model_id).
- **OpenAI GPT Live Transcribe** — New row this refresh: low-latency streaming transcription model for live audio via /v1/realtime/transcription_sessions, flat $0.017/minute of realtime audio duration — not token-billed. Price confirmed on both developers.openai.com/api/docs/pricing and the model detail page developers.openai.com/api/docs/models/gpt-live-transcribe ('Default snapshot: gpt-live-transcribe', active). Priced identically to the existing gpt-realtime-whisper row at the same $0.017/min; kept as a separate row since it is a distinct model_id on the vendor's models catalog, not documented as a rename or successor of gpt-realtime-whisper (both remain live and undeprecated as of this pass). confidence: medium because the supported-language list is not explicitly published for this model — '99+' inherited from the rest of the OpenAI transcription family. 2026-09-02 re-verify: still active, price unchanged ($0.017/min). Explicitly named on /api/docs/deprecations as a recommended replacement target for the newly-deprecated whisper-1 / gpt-4o-transcribe family (not itself deprecated).
- **Groq Whisper V3 Large (Groq)** — Whisper V3 Large hosted on Groq. Published as $0.111/hr. Speed factor 217x realtime. Re-verified 2026-09-02 against console.groq.com/docs/model/whisper-large-v3 and console.groq.com/docs/models (listed as a Production Model): model_id active, price ($0.111/hr = $0.00185/min) and 10s minimum billing unchanged. Not on console.groq.com/docs/deprecations. groq.com/pricing still 308-redirects to the marketing homepage (dead as a source_url) — use console.groq.com.
- **Groq Whisper Large v3 Turbo (Groq)** — Whisper Large v3 Turbo hosted on Groq — faster variant. Published as $0.04/hr. Speed factor 228x realtime. Re-verified 2026-09-02 against console.groq.com/docs/model/whisper-large-v3-turbo and console.groq.com/docs/models (listed as a Production Model): model_id active, price ($0.04/hr = $0.000667/min) and 10s minimum billing unchanged. console.groq.com/docs/deprecations confirms whisper-large-v3-turbo active and lists it as the replacement for the already-deprecated distil-whisper-large-v3-en (shutdown 2025-08-23, not present in this dataset). groq.com/pricing still 308-redirects to the marketing homepage (dead as a source_url) — use console.groq.com.
- **Microsoft Azure Azure Speech (Real-time)** — Azure Speech Standard (S0) real-time speech-to-text. Published as $1.00/hr pay-as-you-go (= $0.016667/min), unchanged; custom real-time endpoint is $1.20/hr ($0.02/min). Commitment tiers reduce effective rate (2,000 hrs/mo $0.80/hr; 10,000 hrs/mo $0.65/hr; 50,000 hrs/mo $0.50/hr; 100,000 hrs/mo $0.40/hr — all confirmed unchanged via Retail Prices API). Enhanced add-on features on real-time (diarization, continuous language identification, pronunciation assessment) cost $0.30/hr per feature (= $0.005/min). Real-time diarization limited to 240 min/session. Languages: 100+ (139 locales per learn.microsoft.com language-support, up from 137 previously recorded). 2026-09-02 pass: azure.microsoft.com/en-us/pricing/details/speech/ fetched via WebFetch (per-hour figures still render client-side, page shows placeholder '$-' values), exact rates re-confirmed via Azure Retail Prices API (prices.azure.com, eastus USD: 'S1 Speech To Text' $1.00/hr, 'S1 Custom Speech To Text' $1.20/hr, 'S1 Speech to Text Enhanced Feature Audio' $0.30/hr, all commitment-tier meters unchanged including newly-checked 100K tier at $0.40/hr), cross-verified via learn.microsoft.com/en-us/azure/ai-services/speech-service/releasenotes (no STT pricing changes; 2026 changes are SDK feature releases — multichannel audio, source-language autodetection — plus retirement of ConversationTranslator/MeetingTranscriber/Intent Recognition, none of which affect this row) and language-support docs. No new base STT pricing tiers found in the Retail Prices API catalog for productName 'Azure Speech'. All figures unchanged from 2026-08-11 pass; confidence high.
- **Microsoft Azure Azure Speech (Batch)** — Azure Speech Standard (S0) batch transcription. Published as $0.18/hr (= $0.003/min); custom batch endpoint is $0.225/hr ($0.00375/min). Batch enhanced add-ons (diarization, continuous language identification) are included at no extra charge; diarization up to 240 min/file. Fast transcription is modelled as its own row (azure-speech-fast). 2026-09-02 pass: azure.microsoft.com/en-us/pricing/details/speech/ fetched via WebFetch (figures still render client-side, page shows placeholder '$-' values), exact rates re-confirmed via Azure Retail Prices API (prices.azure.com, eastus USD: 'S1 Speech to Text Batch' $0.18/hr, 'S1 Custom Speech to Text Batch' $0.225/hr, unchanged), cross-verified via learn.microsoft.com/en-us/azure/ai-services/speech-service/releasenotes (no STT pricing changes or deprecations in 2026 releases; only SDK feature/retirement notes unrelated to batch pricing). No new batch STT tiers found in the Retail Prices API catalog. Figures unchanged from 2026-08-11 pass; confidence high.
- **Microsoft Azure Azure Speech (Fast)** — Azure Speech fast transcription: synchronous REST API (/speechtotext/transcriptions:transcribe) returning faster-than-real-time file transcription; not streaming. Published as $0.36/hr (= $0.006/min); custom fast transcription is $0.45/hr ($0.0075/min). Diarization included at no extra charge. Limits (S0): <500 MB and <5 hrs (300 min) per file, 600 requests/min. Supported across the 139-locale Azure STT locale set (fast transcription marked supported for the large majority; a handful of niche locales unsupported) per language-support docs. model_id is a Hail-coined tier slug — Azure exposes no stable model identifier. 2026-09-02 pass: rates re-confirmed via Azure Retail Prices API (prices.azure.com, eastus USD: 'Fast Transcription Speech To Text' $0.36/hr, 'Custom - Fast Transcription' $0.45/hr, unchanged), cross-verified via learn.microsoft.com/en-us/azure/ai-services/speech-service/releasenotes (no fast-transcription pricing or limit changes in 2026 releases) and language-support docs. No new fast-transcription tiers found in the Retail Prices API catalog. Figures unchanged from 2026-08-11 pass; confidence high.
- **Google Chirp 2** — Google Cloud Speech-to-Text v2 multilingual model. Standard tier $0.016/min for both real-time and batch (down from v1's $0.024/min, per https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-speech-to-text-v2-api). Dynamic Batch tier (up to 24h SLA) listed at $0.003/min on the pricing page as of 2026-07-19 (sku 7700-6778-EF8E; previously noted $0.004/min '75% off'; page footnote enumerates the Standard dynamic-batch models as default/command_and_search/latest_short/latest_long/phone_call/video/chirp without naming chirp_2) — not modelled as price_per_minute_batch_usd because standard batch is the same as real-time. Standard volume tiers: $0.016/min (0-500k min/mo), $0.01 (500k-1M), $0.008 (1M-2M), $0.004 (2M+). Supports StreamingRecognize (~20 languages), Recognize, and BatchRecognize (broadest language coverage) per https://docs.cloud.google.com/speech-to-text/docs/models/chirp-2. GA in us-central1, europe-west4, asia-southeast1. 2026-07-19 re-verify: WebFetch of the pricing page still truncates (known issue), but a raw-HTML fetch of the primary succeeded this pass — V2 standard recognition confirmed $0.016/min (sku 3099-B70F-0949); cross-checked via https://docs.cloud.google.com/speech-to-text/v2/docs/transcription-model (secondary; chirp_2 still listed GA, streaming, not deprecated) and web-search snippets citing the Google V2 launch blog. No structured price change. Language-count sub-field left untouched — prior passes found inconsistent counts (72/90+/150+) from the truncating supported-languages table, so it does not meet the two-source bar for a non-price-field edit. 2026-08-11 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch of the primary succeeded and confirms Standard recognition tiers and the $0.003/min Dynamic Batch rate (sku 7700-6778-EF8E) unchanged, with no 'deprecated'/'retired'/'sunset' text anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_2 as GA with streaming and batch support. No new Google STT model_ids found (checked for a Chirp 4 launch; none). No structured price change. 2026-09-02 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch (decoded unicode-escaped JSON payload) confirms Standard Recognition tiers unchanged (sku 3099-B70F-0949: $0.016/$0.01/$0.008/$0.004 by volume) and Dynamic Batch Recognition unchanged (sku 7700-6778-EF8E: $0.003/min); zero hits for 'deprecat'/'retire'/'sunset'/'discontinu' anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_2 as not deprecated, streaming+batch supported. Web search confirms no Chirp 4 launch; Chirp 3 remains Google's latest-generation model as of 2026-09. No structured price change.
- **Google Chirp 3** — Google Cloud Speech-to-Text v2 latest-generation generative ASR model. Standard tier $0.016/min for both real-time and batch; Dynamic Batch tier (up to 24h SLA) listed at $0.003/min on the pricing page as of 2026-07-19 (sku 7700-6778-EF8E; previously noted $0.004/min '75% off'; page footnote enumerates the Standard dynamic-batch models without naming chirp_3) — not modelled as price_per_minute_batch_usd because standard batch is the same as real-time. Standard volume tiers: $0.016/min (0-500k min/mo), $0.01 (500k-1M), $0.008 (1M-2M), $0.004 (2M+). Adds automatic language detection and diarization vs Chirp 2 per https://docs.cloud.google.com/speech-to-text/docs/models/chirp-3. 98+ languages and locales (24 GA + 74 preview) per prior verification; supports StreamingRecognize and BatchRecognize. 2026-07-19 re-verify: WebFetch of the pricing page still truncates (known issue), but a raw-HTML fetch of the primary succeeded this pass — V2 standard recognition confirmed $0.016/min (sku 3099-B70F-0949); cross-checked via https://docs.cloud.google.com/speech-to-text/v2/docs/transcription-model (secondary; chirp_3 still listed GA as latest generation, streaming, not deprecated) and web-search snippets confirming $0.016/min standard for Chirp 3. No structured price change. Language-count sub-field left untouched — prior passes found inconsistent GA/preview splits and totals from truncating tables, and this pass's search snippets differ again (125+), so it does not meet the two-source bar for a non-price-field edit; existing 98+ figure retained pending a cleaner source. 2026-08-11 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch of the primary succeeded and confirms Standard recognition tiers and the $0.003/min Dynamic Batch rate (sku 7700-6778-EF8E) unchanged, with no 'deprecated'/'retired'/'sunset' text anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_3 as GA, latest generation, streaming, batch. No new Google STT model_ids found (checked for a Chirp 4 launch; none). No structured price change. 2026-09-02 re-verify: WebFetch of the primary truncated again after one retry; a raw-HTML fetch (decoded unicode-escaped JSON payload) confirms Standard Recognition tiers unchanged (sku 3099-B70F-0949: $0.016/$0.01/$0.008/$0.004 by volume) and Dynamic Batch Recognition unchanged (sku 7700-6778-EF8E: $0.003/min); zero hits for 'deprecat'/'retire'/'sunset'/'discontinu' anywhere on the page. Secondary (docs.cloud.google.com/speech-to-text/v2/docs/transcription-model) still lists chirp_3 as not deprecated, streaming+batch supported. Web search confirms no Chirp 4 launch; Chirp 3 remains Google's latest-generation model as of 2026-09. No structured price change.
- **Speechmatics Speechmatics Enhanced** — Speechmatics offers three operating points — `enhanced` (highest accuracy), `standard` (faster/cheaper), and the newer `melia-1` (multilingual, batch-only) — selected via the `operating_point` API parameter on both Batch and Real-time APIs (Melia 1 is batch-only). Pricing page lists Pro tier from $0.129/hr ($0.00215/min) on PAYG (price cut from $0.24/hr at last check) with an automatic 20% volume discount above 500 hrs/month; the same flat rate is used for both real-time and batch and is not broken out per operating point. Free plan now includes 3,000 minutes (50 hrs)/month, up from 480 min (not modelled as `free_tier` because schema expects per-day/per-token quotas). Confidence is medium because the pricing page exposes tier names rather than per-model SKU rates; verify against contract for production. Cross-verified product structure via https://docs.speechmatics.com/. 2026-07-19 re-verify: pricing page (primary) still shows Pro tier $0.129/hr ($0.00215/min) PAYG with the same 20% >500 hrs/month volume discount and 3,000 min (50 hrs)/month free tier; per-model rates still not broken out. Cross-checked https://docs.speechmatics.com/speech-to-text/models (secondary) — Enhanced still GA, batch+real-time, EU/US/AU; legacy `operating_point` parameter remains deprecated in favor of `model` (naming only, no effect on this row's model_id or price). Docs also mention a specialized Enhanced Medical variant with no separately published price. No price change. 2026-08-11 re-verify: pricing page (primary) still shows Pro tier $0.129/hr ($0.00215/min) PAYG with the same 20% >500 hrs/month volume discount (24,000 hrs/yr threshold for additional discounts); per-model rates still not broken out. Free plan has changed structure: it is now a one-time "$100 in credit, no card required" on signup rather than the previously-noted recurring 3,000 min (50 hrs)/month allotment — no recurring monthly free minutes are advertised on the page any longer (not modelled as `free_tier` field per prior note). Cross-checked https://docs.speechmatics.com/speech-to-text/models (secondary) — Enhanced still GA, batch+real-time, custom dictionaries/confidence scores/speaker ID/intelligence features intact. No price change. 2026-09-02 re-verify: pricing page (primary) still shows Pro tier $0.129/hr ($0.00215/min) PAYG, $100 signup credit, and the same 20% >500 hrs/month + 24,000 hrs/yr volume-discount structure; per-model rates still not broken out. Page now also lists an opt-in "model training discount" of 33% off STT rates for allowing Speechmatics to use submitted data for model improvement (off by default, reversible) — not modelled as a price field since it is conditional, not a published flat rate. Pricing page's general tier-features blurb currently rounds to "55+ languages" platform-wide (vs this row's "56+"); left unchanged since that figure is not published per-model and the secondary source gives no per-model count. Cross-checked https://docs.speechmatics.com/speech-to-text/models (secondary) — Enhanced still GA, batch+real-time, EU/US/AU, custom dictionaries/confidence scores/speaker ID intact; Medical domain variant still listed with no separate price. No price change; not deprecated.
- **Speechmatics Speechmatics Melia 1** — New `operating_point=melia-1` model added since last refresh: early-access, multilingual model that auto-detects and transcribes code-switched audio into one continuous transcript without language-pack selection, per https://docs.speechmatics.com/speech-to-text/models. Batch-only (no real-time API support) — Enhanced and Standard remain active and unaffected. Matches Standard's accuracy tier; lacks custom dictionaries, confidence scores, and speaker/intelligence features available on Enhanced/Standard, though diarization and word timings are supported. Available only in EU/US regions (not AU). `68+` languages is a derived count (the docs table lists 72 entries: ~68 individual languages plus 4 bilingual packs which Melia 1 does not use) — not an officially published per-model figure. Billed at the same flat Pro-tier rate as Enhanced/Standard per https://www.speechmatics.com/pricing — the pricing page does not publish a per-model rate, hence confidence: medium. 2026-07-19 re-verify: still early-access, batch-only, EU/US-only (accuracy on par with Standard; invoked via `model: melia-1` + `language: multi`) per https://docs.speechmatics.com/speech-to-text/models (secondary); flat Pro-tier rate ($0.129/hr / $0.00215/min) unchanged on the pricing page (primary). No price change; not deprecated. 2026-08-11 re-verify: still early-access, batch-only, EU/US-only per https://docs.speechmatics.com/speech-to-text/models (secondary) — no GA promotion, no region expansion. Flat Pro-tier rate ($0.129/hr / $0.00215/min) unchanged on the pricing page (primary); same rate applies across Enhanced/Standard/Melia 1 since Speechmatics still does not publish per-model SKU pricing. No price change; not deprecated. 2026-09-02 re-verify: still early-access, batch-only, EU/US-only per https://docs.speechmatics.com/speech-to-text/models (secondary) — no GA promotion, no region expansion, still lacks custom dictionaries/confidence scores/speaker ID. Flat Pro-tier rate ($0.129/hr / $0.00215/min) unchanged on the pricing page (primary), including the same 20% >500 hrs/month volume discount; pricing page's new opt-in 33% model-training discount applies platform-wide, not modelled as a separate field. No price change; not deprecated.
- **Speechmatics Speechmatics Standard** — Added 2026-09-02: `operating_point=standard` (now `model=standard`) is Speechmatics' third GA operating point — faster/cheaper than Enhanced, matched to Melia 1's accuracy tier per the Enhanced-vs-Standard-vs-Melia-1 comparison at https://docs.speechmatics.com/speech-to-text/models (secondary). This tier was described in the `enhanced` row's notes since the dataset's first Speechmatics pass but never modelled as its own row; added now for completeness since it is a distinct, independently-selectable, fully-documented operating point. GA (not early-access, unlike Melia 1), Batch + Realtime, available in EU/US/AU (same regional footprint as Enhanced) and carries the same custom dictionaries/confidence scores/speaker ID/diarization feature set as Enhanced. Billed at the same flat Pro-tier rate as Enhanced/Melia 1 per https://www.speechmatics.com/pricing (primary) — $0.129/hr ($0.00215/min) PAYG with the same 20% >500 hrs/month volume discount; Speechmatics does not publish a per-model SKU rate, hence confidence: medium (same reasoning as the other two rows). `56+` languages mirrors the Enhanced row's figure since both use the same language-pack model and no separate per-model count is published; not officially confirmed distinct from Enhanced's count. Not deprecated.
- **Rev.ai Rev.ai Whisper Fusion** — Rev.ai's streaming transcription product, branded `Whisper Fusion` on the pricing page at $0.005/min; the parallel Whisper Large streaming tier is also $0.005/min. Free credits equivalent to 5 hours of Reverb ASR (cross-applicable across products). Reverb (batch) is modelled separately; see https://docs.rev.ai/ for the full API surface. English-primary; foreign language support is a distinct Reverb Foreign Language product line. 2026-07-13 re-verify: still listed at $0.005/min on the pricing page (primary); no price change. Note for context: the async job API's `transcriber` enum uses `fusion`/`machine`/`low_cost` rather than these marketing names (per docs.rev.ai, secondary) — model_id here follows this dataset's pricing-page-branding convention, unchanged from prior refresh. 2026-07-19 re-verify: still listed at $0.005/min on the pricing page (primary); `fusion` still an active `transcriber` value in the async job API reference (docs.rev.ai, secondary); no price change, not deprecated. 2026-08-11 re-verify: still listed at $0.005/min on the pricing page (primary); no price change. Secondary (docs.rev.ai/api/streaming/transcribers/) now documents only `machine` (account default) and `machine_v2` (always routes to Reverb ASR) — `fusion`/`low_cost` are no longer named there (a docs restructuring), but `Whisper Fusion` remains listed and priced on the pricing page itself; not deprecated. 2026-09-02 re-verify: still listed at $0.005/min on the pricing page (primary); no price change. Secondary (docs.rev.ai/api/streaming/transcribers/) still documents only `machine`/`machine_v2`, no `fusion` value; not deprecated — pricing page remains the authoritative listing for this product.
- **Rev.ai Rev.ai Reverb** — Rev.ai's async/batch ASR model branded `Reverb` at $0.20/hr ($0.0033/min). A `Reverb Turbo` tier exists at $0.10/hr ($0.0017/min) — not modelled as a separate row since it's a latency/quality dial on the same product; `Reverb Foreign Language` ($0.30/hr, $0.005/min, 56+ languages) is also priced separately and could be added if needed. Free credits equivalent to 5 hours of Reverb ASR. See https://docs.rev.ai/ for the async transcription API. 2026-07-13 re-verify: still listed at $0.20/hr on the pricing page (primary); no price change. Reverb Turbo and Reverb Foreign Language remain distinct pricing tiers, still not modelled as separate rows (unchanged decision). 2026-07-19 re-verify: still listed at $0.20/hr on the pricing page (primary); `machine` (routing to Reverb) still the default `transcriber` value in the async job API reference (docs.rev.ai, secondary); no price change, not deprecated. 2026-08-11 re-verify: still listed at $0.20/hr on the pricing page (primary); no price change; Reverb Foreign Language language count on the pricing page now reads 56+ (previously noted as 57+ — wording/count drift, not a pricing change). Secondary (docs.rev.ai/api/asynchronous/transcribers/) now documents `machine` as routing to the Reverb ASR model (`machine_v2` on streaming always routes to Reverb); no deprecation. 2026-09-02 re-verify: still listed at $0.20/hr on the pricing page (primary); no price change. Secondary (docs.rev.ai/api/asynchronous/transcribers/) documents `machine` as the default, routed to the Reverb ASR model; `human` transcriber also documented separately ($1.99/min, not modelled — human transcription, not ASR). Not deprecated.
- **Gladia Gladia Solaria-1** — Gladia's first-generation universal STT model `solaria-1`, supporting 100+ languages with automatic language detection and code-switching. Default model when `model` param is omitted (confirmed via docs.gladia.io pre-recorded transcription reference). Starter (PAYG) pricing: real-time $0.75/hr ($0.0125/min), async $0.61/hr (~$0.01017/min). Growth (committed) plan lowers real-time to $0.25/hr ($0.0042/min) and async to $0.20/hr ($0.0033/min). Sub-300ms streaming latency claimed on the pricing page. Speaker diarization and word-level timestamps included on all tiers. Model name confirmed via https://docs.gladia.io/. 2026-07-19 re-verify: still listed at the same rates on the pricing page (primary) and the pre-recorded API reference on docs.gladia.io (secondary, still the default model); no price change. 2026-08-11 re-verify: still listed at the same rates on the pricing page (primary); pre-recorded API reference on docs.gladia.io (secondary) still lists `solaria-1` as the default `model` value with 100+ languages and code-switching support; no price change. Free-tier wording corrected: the pricing page now describes a one-time 50€ signup credit (~80+ async hrs or ~60+ real-time hrs, no monthly reset), not the '10 free hours per month' previously noted here — likely a prior misread rather than a policy change, since no monthly-reset credit is documented on the current page. 2026-09-02 re-verify: still listed at the same rates on the pricing page (primary); docs.gladia.io models comparison page and pre-recorded API reference (secondary) still list `solaria-1` as the default model, 100+ languages, code-switching supported, async+live; no price change; no new Gladia STT models found.
- **Gladia Gladia Solaria-3** — Gladia's higher-accuracy model for European real-world audio (English, French, German, Spanish, Italian); async/pre-recorded only, requires exactly one language in `language_config.languages` (no code-switching). Confirmed as a valid `model` param value via https://docs.gladia.io (pre-recorded transcription API reference). The pricing page does not break out per-model rates, so the Starter async rate ($0.61/hr ≈ $0.01017/min) is assumed to apply; `confidence: medium` reflects that this price is not explicitly stated per-model on the pricing page. No real-time/streaming support documented for this model. 2026-07-19 re-verify: still active and async-only per the pre-recorded API reference (secondary); pricing page (primary) still lists no per-model rate, so the Starter async assumption stands; no price change. 2026-08-11 re-verify: still active and async-only per the pre-recorded API reference (secondary, still "English, French, German, Spanish, Italian", still no code-switching); pricing page (primary) still has no per-model breakdown, so the Starter async assumption and `confidence: medium` stand; no price change. 2026-09-02 re-verify: still active and async-only per docs.gladia.io models comparison page and pre-recorded API reference (secondary, same 5 languages, no code-switching); pricing page (primary) still has no per-model breakdown, so the Starter async assumption and `confidence: medium` stand; no price change.
- **Soniox Soniox STT Real-time v4** — Soniox real-time STT model `stt-rt-v4`. Primary billing metric is input audio tokens at $2.00 per 1M tokens; vendor approximates ~$0.12/hour which we use as $0.002/min for comparability. Aliased from `stt-rt-v3` (deprecated 2026-02-05; removed 2026-02-28 per https://soniox.com/docs/stt/models). Removed 2026-06-30 per https://soniox.com/docs/stt/models: requests using `stt-rt-v4` now auto-route to `stt-rt-v5` with no service interruption. 60+ languages with automatic language detection. Confidence medium because per-minute is an approximation of token-based pricing; actual cost varies with audio content density. 2026-07-19 re-verify: still deprecated/auto-routing to stt-rt-v5 per soniox.com/docs/stt/models (secondary); pricing page (primary) still quotes $2.00/1M input audio tokens for real-time; no price change. 2026-08-11 re-verify: still deprecated/auto-routing to stt-rt-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $2.00/1M input audio tokens for real-time; no price change. 2026-09-02 re-verify: still deprecated/auto-routing to stt-rt-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $2.00/1M input audio tokens for real-time; no price change.
- **Soniox Soniox STT Real-time v5** — Soniox real-time STT model `stt-rt-v4` (deprecated 2026-06-30; auto-routes to `stt-rt-v5`) per https://soniox.com/docs/stt/models: `stt-rt-v5` released 2026-06-16 with reworked speaker separation and improved spoken-language identification. Soniox's pricing page bills real-time by input audio tokens at $2.00 per 1M tokens regardless of model version (~$0.12/hour, used here as $0.002/min for comparability); the docs page confirms pricing continuity across the v4→v5 auto-route ("no service interruption"). 60+ languages with automatic language detection. Confidence medium because (a) per-minute is an approximation of token-based pricing and (b) the pricing page does not itemize a v5-specific rate distinct from the general real-time rate. 2026-07-19 re-verify: still the current active real-time model per soniox.com/docs/stt/models (secondary), no newer version listed (no v6); pricing page (primary) still quotes $2.00/1M input audio tokens; no price change. 2026-08-11 re-verify: still the current active real-time model per soniox.com/docs/stt/models (secondary, still no v6, still "up to 5 hours" per request); pricing page (primary) still quotes $2.00/1M input audio tokens; no price change. 2026-09-02 re-verify: still the current active real-time model per soniox.com/docs/stt/models (secondary, still no v6, still "up to 5 hours" per request); pricing page (primary) still quotes $2.00/1M input audio tokens; no price change.
- **Soniox Soniox STT Async v4** — Soniox async (file) STT model `stt-async-v4`. Primary billing metric is input audio tokens at $1.50 per 1M tokens; vendor approximates ~$0.10/hour which we use as $0.00167/min for comparability. Aliased from `stt-async-v3` (deprecated 2026-02-05; removed 2026-02-28 per https://soniox.com/docs/stt/models). Removed 2026-06-30 per https://soniox.com/docs/stt/models: requests using `stt-async-v4` now auto-route to `stt-async-v5` with no service interruption. Supports up to 5 hours of audio per request. 60+ languages with automatic language detection. Confidence medium because per-minute is an approximation of token-based pricing. 2026-07-19 re-verify: still deprecated/auto-routing to stt-async-v5 per soniox.com/docs/stt/models (secondary); pricing page (primary) still quotes $1.50/1M input audio tokens for async; no price change. 2026-08-11 re-verify: still deprecated/auto-routing to stt-async-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $1.50/1M input audio tokens for async; no price change. 2026-09-02 re-verify: still deprecated/auto-routing to stt-async-v5 per soniox.com/docs/stt/models (secondary, no v6 listed); pricing page (primary) still quotes $1.50/1M input audio tokens for async; no price change.
- **Soniox Soniox STT Async v5** — Soniox async (file) STT model `stt-async-v4` (deprecated 2026-06-30; auto-routes to `stt-async-v5`) per https://soniox.com/docs/stt/models: `stt-async-v5` released 2026-06-11 with reengineered speaker separation and improved spoken-language identification. Soniox's pricing page bills async by input audio tokens at $1.50 per 1M tokens regardless of model version (~$0.10/hour, used here as $0.00167/min for comparability); the docs page confirms pricing continuity across the v4→v5 auto-route ("no service interruption"). Supports up to 5 hours of audio per request (carried over from v4; not independently re-stated for v5 in the docs excerpt). 60+ languages with automatic language detection. Confidence medium because (a) per-minute is an approximation of token-based pricing and (b) the max-audio-duration and pricing figures are inferred from vendor continuity statements rather than v5-specific line items. 2026-07-19 re-verify: still the current active async model per soniox.com/docs/stt/models (secondary), no newer version listed (no v6); pricing page (primary) still quotes $1.50/1M input audio tokens; no price change. 2026-08-11 re-verify: still the current active async model per soniox.com/docs/stt/models (secondary, still no v6, still "up to 5 hours" per request); pricing page (primary) still quotes $1.50/1M input audio tokens; no price change. 2026-09-02 re-verify: still the current active async model per soniox.com/docs/stt/models (secondary, still no v6, still "up to 5 hours" per request); pricing page (primary) still quotes $1.50/1M input audio tokens; no price change.


## Text-to-speech

| Provider | Model | $/1M chars | Quality | Cloning | Languages | TTFB | SSML | Verified | Source |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ElevenLabs | Eleven Flash v2.5 (`eleven_flash_v2_5`) | $50.00 | neural | yes | 32+ | 75ms | no | 2026-09-02 | [link](https://elevenlabs.io/pricing/api) |
| ElevenLabs | Eleven Multilingual v2 (`eleven_multilingual_v2`) | $100.00 | neural | yes | 29+ | — | no | 2026-09-02 | [link](https://elevenlabs.io/docs/overview/models) |
| ElevenLabs | Eleven v3 (`eleven_v3`) | $100.00 | neural | yes | 70+ | — | no | 2026-09-02 | [link](https://elevenlabs.io/docs/overview/models) |
| ElevenLabs | Eleven v3 Conversational (`eleven_v3_conversational`) | $50.00 | neural | — | 70+ | 280ms | no | 2026-09-02 | [link](https://elevenlabs.io/pricing/api) |
| ElevenLabs | Eleven Turbo v2.5 (`eleven_turbo_v2_5`) | $50.00 | neural | yes | 32+ | — | no | 2026-09-02 | [link](https://elevenlabs.io/docs/overview/models) |
| ElevenLabs | Eleven Flash v2 (`eleven_flash_v2`) | $50.00 | neural | yes | English | 75ms | no | 2026-09-02 | [link](https://elevenlabs.io/docs/overview/models) |
| OpenAI | TTS-1 (`tts-1`) | $15.00 | neural | no | en | — | no | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/tts-1) |
| OpenAI | TTS-1 HD (`tts-1-hd`) | $30.00 | neural | no | en | — | no | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/tts-1-hd) |
| OpenAI | GPT-4o mini TTS (`gpt-4o-mini-tts`) | $20.00 | neural | no | en | — | no | 2026-09-02 | [link](https://developers.openai.com/api/docs/models/gpt-4o-mini-tts) |
| Cartesia | Sonic 3.6 (`sonic-3.6`) | $50.00 | neural | yes | 44+ | 90ms | yes | 2026-09-02 | [link](https://cartesia.ai/pricing) |
| Cartesia | Sonic 3.5 (`sonic-3.5`) | $50.00 | neural | yes | 42+ | 90ms | yes | 2026-09-02 | [link](https://cartesia.ai/pricing) |
| Cartesia | Sonic 3 (`sonic-3`) | $50.00 | neural | yes | 44+ | — | yes | 2026-09-02 | [link](https://docs.cartesia.ai/build-with-cartesia/tts-models/older-models) |
| Cartesia | Sonic 2 (`sonic-2`) | $50.00 | neural | yes | en, fr, de, es, pt, zh, ja, ko | 90ms | no | 2026-09-02 | [link](https://docs.cartesia.ai/build-with-cartesia/tts-models/older-models) |
| Cartesia | Sonic Turbo (`sonic-turbo`) | $50.00 | neural | yes | en, fr, de, es, pt, zh, ja, hi, ko | 40ms | no | 2026-09-02 | [link](https://docs.cartesia.ai/build-with-cartesia/tts-models/older-models) |
| Groq | Canopy Labs Orpheus English (Groq) (`canopy-labs-orpheus-english`) | $22.00 | neural | — | en | — | no | 2026-09-02 | [link](https://console.groq.com/docs/models) |
| Groq | Canopy Labs Orpheus Arabic Saudi (Groq) (`canopy-labs-orpheus-arabic-saudi`) | $40.00 | neural | — | ar-SA | — | no | 2026-09-02 | [link](https://console.groq.com/docs/models) |
| Google | Google Cloud TTS — Studio (`google-tts-studio`) | $160.00 | neural | no | 40+ | — | yes | 2026-09-02 | [link](https://cloud.google.com/text-to-speech/pricing) |
| Google | Google Cloud TTS — Neural2 (`google-tts-neural2`) | $16.00 | neural | no | 40+ | — | yes | 2026-09-02 | [link](https://cloud.google.com/text-to-speech/pricing) |
| Google | Google Cloud TTS — WaveNet (`google-tts-wavenet`) | $4.00 | neural | no | 40+ | — | yes | 2026-09-02 | [link](https://cloud.google.com/text-to-speech/pricing) |
| Google | Google Cloud TTS — Chirp 3: HD (`google-tts-chirp-3-hd`) | $30.00 | neural | no | 30+ | — | no | 2026-09-02 | [link](https://cloud.google.com/text-to-speech/pricing) |
| Google | Google Cloud TTS — Chirp 3: Instant custom voice (`google-tts-instant-custom-voice`) | $60.00 | cloned | yes | 30+ | — | no | 2026-09-02 | [link](https://cloud.google.com/text-to-speech/pricing) |
| Microsoft Azure | Azure AI Speech — Neural (`azure-tts-neural`) | $15.00 | neural | no | 100+ | — | yes | 2026-09-02 | [link](https://azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/) |
| Microsoft Azure | Azure AI Speech — Neural HD (DragonHD) (`azure-tts-hd`) | $22.00 | neural | no | en-US, zh-CN, de-DE, es-ES, fr-FR, ja-JP | — | no | 2026-09-02 | [link](https://azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/) |
| Microsoft Azure | Azure AI Speech — Personal Voice (`azure-tts-personal-voice`) | $24.00 | neural | yes | 90+ | — | no | 2026-09-02 | [link](https://azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/) |
| Inworld | Inworld Realtime TTS-2 (`inworld-tts-2`) | $25.00 | neural | yes | 200+ | — | no | 2026-09-02 | [link](https://inworld.ai/pricing) |
| Inworld | Inworld Realtime TTS-2 Flash (`inworld-tts-2-flash`) | $15.00 | neural | yes | 200+ | 20ms | no | 2026-09-02 | [link](https://inworld.ai/pricing) |
| Inworld | Inworld Realtime TTS 1.5 Max (`inworld-tts-1.5-max`) | $35.00 | neural | yes | en, zh, ja, ko, ru, it, es, pt, fr, de, pl, nl, hi, he, ar | — | no | 2026-09-02 | [link](https://inworld.ai/pricing) |
| Inworld | Inworld Realtime TTS 1.5 Mini (`inworld-tts-1.5-mini`) | $15.00 | neural | yes | en, zh, ja, ko, ru, it, es, pt, fr, de, pl, nl, hi, he, ar | — | no | 2026-09-02 | [link](https://inworld.ai/pricing) |
| Smallest.ai | Lightning v3.1 (`lightning_v3.1`) | $17.50 | neural | yes | en, hi, es, mr, kn, ta, bn, gu, te, ml, pa, or | 200ms | no | 2026-09-02 | [link](https://smallest.ai/pricing/models) |
| Smallest.ai | Lightning v3.1 Pro (`lightning_v3.1_pro`) | $19.50 | neural | yes | en, hi, mr, ta, ml, te, kn, pa, bn, or, gu, ar, zh, id, ja, ko, ms, tr, vi, de, es, fr, it, nl, sv, pt, ru, el, fi, no, pl | 200ms | no | 2026-09-02 | [link](https://smallest.ai/pricing/models) |
| Rime | Rime Mist v3 (`mistv3`) | $30.00 | neural | no | en, fr, de, es | 37ms | no | 2026-09-02 | [link](https://rime.ai/pricing) |
| Rime | Rime Coda (`coda`) | $50.00 | neural | — | en, ar, fr, de, hi, it, ja, pt, es | — | no | 2026-09-02 | [link](https://rime.ai/pricing) |
| LMNT | LMNT Blizzard (`blizzard`) | $50.00 | neural | yes | 31+ | — | no | 2026-09-02 | [link](https://www.lmnt.com/pricing) |
| Deepgram | Deepgram Aura 2 (`aura-2`) | $30.00 | neural | no | en, es, de, fr, nl, it, ja | — | no | 2026-09-02 | [link](https://deepgram.com/pricing) |
| Resemble AI | Resemble Chatterbox Turbo (`chatterbox-turbo`) | — | neural | yes | en | — | no | 2026-09-02 | [link](https://www.resemble.ai/pricing) |

**Notes:**

- **ElevenLabs Eleven Flash v2.5** — Current low-latency flagship; eleven_turbo_v2_5 is deprecated and replaced by Flash v2.5. Pay-as-you-go rate $0.05/1K chars. ~10K voices available. Plain text input only (no SSML).
- **ElevenLabs Eleven Multilingual v2** — High-quality professional model for audiobooks, video narration, and rich emotional expression. 29 languages, max 10,000 chars per request. Pay-as-you-go billed at 1 credit per character; Flash/Turbo v2.5 are billed at 0.5 credits/char (hence 2x the Flash $50/1M-chars rate). Higher latency than Flash; not recommended for real-time agents.
- **ElevenLabs Eleven v3** — Most expressive ElevenLabs TTS model (GA after alpha). 70+ languages, max 5,000 chars per request. Supports inline audio tags ([whispers], [sighs], [laughs], [happily]) for emotion/delivery control instead of SSML. Higher latency than Flash/Turbo v2.5 — ElevenLabs explicitly recommends v2.5 Flash/Turbo for real-time use. Pay-as-you-go billed at 1 credit/char (same multiplier as Multilingual v2). PVCs (professional voice clones) not yet fully optimized for v3.
- **ElevenLabs Eleven v3 Conversational** — New: low-latency realtime variant of Eleven v3, tuned for voice agents/conversational use ("v3 Conversational" on pricing page; model_id eleven_v3_conversational confirmed via elevenlabs.io/docs/models). 70+ languages, ~280ms latency, exposed via Text to Dialogue WebSocket. Supports audio tags for emotion/delivery control (no SSML, consistent with rest of ElevenLabs lineup). Billed at the Flash tier rate (0.5 credits/char, $0.05/1K chars) vs. eleven_v3's 1 credit/char. Character limit, exact launch date, and voice-cloning support not stated on either source as of verification; not marked featured pending a controller decision. Added 2026-09-02 refresh.
- **ElevenLabs Eleven Turbo v2.5** — Deprecated per ElevenLabs models page — outclassed by and replaced by eleven_flash_v2_5. Still callable but not recommended for new applications. No official sunset date published; deprecated_at reflects verification date. Pay-as-you-go billed at 0.5 credits/char (same as Flash v2.5).
- **ElevenLabs Eleven Flash v2** — English-only predecessor to Flash v2.5, still listed as an active model. Same pay-as-you-go rate as Flash/Turbo family ($0.05/1K chars per pricing page, which groups Flash and Turbo variants together). ~75ms latency. Plain text input only (no SSML). Previously absent from this dataset; added on 2026-07-02 refresh after confirming it is still live via docs.elevenlabs.io/models.
- **OpenAI TTS-1** — Standard quality. Price unchanged at $15/1M chars as of 2026-09-02 refresh; not listed on the deprecations page; not in the main model-catalog listing (audio section lists only gpt-4o-mini-tts) but confirmed active and priced via the model detail page and the pricing page (openai.com/docs/models and openai.com/docs/pricing redirect to developers.openai.com).
- **OpenAI TTS-1 HD** — High-definition tier — 2x the price of tts-1 for higher quality output. Price unchanged at $30/1M chars as of 2026-09-02 refresh; not listed on the deprecations page; not in the main model-catalog listing (audio section lists only gpt-4o-mini-tts) but confirmed active and priced via the model detail page and the pricing page (openai.com/docs/models and openai.com/docs/pricing redirect to developers.openai.com).
- **OpenAI GPT-4o mini TTS** — OpenAI's newer GPT-4o-based TTS. Native pricing is token-based — $0.60/1M text-input tokens + $12/1M audio-output tokens — not per character. OpenAI's published estimate is ~$0.015 per minute of audio; converted to ~$20/1M chars assuming ~150 WPM (~750 chars/min) for consistency with the Cartesia row. Actual $/1M chars varies with speech rate and language. Supports voice steering via natural-language instructions (style/emotion) instead of SSML. Latest snapshot gpt-4o-mini-tts-2025-12-15; unchanged as of 2026-09-02 refresh (gpt-4o-mini-tts-2025-03-20 remains available as a pinned snapshot alongside it, per the model detail page's snapshot list). Max input 2,000 tokens per request.
- **Cartesia Sonic 3.6** — New Cartesia flagship, GA August 27, 2026, superseding sonic-3.5 as the current-recommended TTS model. Latest snapshot sonic-3.6-2026-08-27; continuously-updated alias sonic-3.6; beta alias sonic-preview. 44 languages / 61 locales, adding Odia (or) and Urdu (ur) with instant voice cloning support (per https://docs.cartesia.ai/build-with-cartesia/tts-models/latest). Ranked #1 on Artificial Analysis's Speech Arena leaderboards per https://www.cartesia.ai/launch (secondary source); built on state-space-model architecture. Sub-90ms TTFB per https://www.cartesia.ai/launch; carried the existing 90ms convention pending a model-specific published number. SSML speed/volume tags apply ("available on sonic-3 and later" per https://docs.cartesia.ai/build-with-cartesia/sonic-3/ssml-tags); model achieves natural expression without requiring SSML. Pricing unchanged from sibling Sonic models: ~1 credit/char on every TTS endpoint (https://docs.cartesia.ai/pricing); USD anchor Pro plan $5/mo for 100K credits => $50/1M chars; no per-model pricing differentiation published. output_formats/sample-rate options are platform-level (shared with other Sonic snapshots), not documented separately per model. Added 2026-09-02 refresh; two-source verified via primary (cartesia.ai/pricing) plus docs.cartesia.ai/build-with-cartesia/tts-models/latest and /older-models.
- **Cartesia Sonic 3.5** — Cartesia now documents TTS billing as ~1 credit per character on every TTS endpoint (https://docs.cartesia.ai/pricing); Pro Voice Clone output is ~1.5 credits/char. USD anchor: Pro plan $5/mo for 100K credits => $50/1M chars; higher tiers are cheaper per credit (Startup ~$39.2, Scale ~$37.4 per 1M chars). Effective rate matches the prior $1-per-25-min + ~150 WPM anchoring, so no structured price change; per-char billing removes the speech-rate assumption. Plan-page minutes math (750 credits/min) is consistent with 1 credit/char at ~150 WPM. IVC voice cloning included (no clone fee). 90ms TTFB (sub-90ms per docs). 42 languages; latest snapshot sonic-3.5-2026-05-04. SSML speed/volume tags confirmed available on sonic-3.5 per https://docs.cartesia.ai/build-with-cartesia/sonic-3/ssml-tags (re-checked 2026-07-02). Emotion tags remain beta. Re-verified 2026-07-19: still Cartesia's current-recommended model. 2026-07-27 refresh: pricing and model status unchanged via primary (cartesia.ai/pricing) and secondary (docs.cartesia.ai/build-with-cartesia/tts-models/older-models) sources. 2026-08-11 refresh: pricing (Pro $5/mo => 133 TTS minutes, consistent with 100K credits at 750 credits/min), snapshot sonic-3.5-2026-05-04, 42 languages, and sub-90ms TTFB confirmed unchanged via primary (cartesia.ai/pricing) and secondary (docs.cartesia.ai/build-with-cartesia/tts-models/latest) sources; no newer Sonic model announced. 2026-09-02 refresh: still stable and callable (no sunset date published), but superseded as Cartesia's top-recommended model by sonic-3.6 (added as new row this pass); pricing and specs unchanged. confidence set to medium to reflect the same per-tier credit-rate variance as sibling rows now that this is no longer the flagship.
- **Cartesia Sonic 3** — Predecessor to sonic-3.5, still stable and a valid model_id on the TTS API. Latest snapshot sonic-3-2026-01-12; older-models page lists 44 languages for this snapshot, updated from prior 42+ count (non-price change). Cartesia recommends sonic-3.5 for best results, most languages, and naturalness. SSML speed/volume tags confirmed available per https://docs.cartesia.ai/build-with-cartesia/sonic-3/ssml-tags. TTFB not published for this specific snapshot. Pricing equality with sonic-3.5 now documented: https://docs.cartesia.ai/pricing applies ~1 credit/char to every TTS endpoint; USD anchor Pro plan $5 per 100K credits => $50/1M chars. Confidence medium because the USD-per-credit rate varies by plan tier (Scale ~$37.4/1M chars). Model sunsets October 20, 2026 per https://docs.cartesia.ai/build-with-cartesia/tts-models/older-models. 2026-07-27 refresh: language count corrected to 44 (non-price), pricing and sunset date confirmed via secondary source. 2026-08-11 refresh: snapshot, 44 languages, pricing, and October 20 2026 sunset date confirmed unchanged via secondary source. 2026-09-02 refresh: snapshot, languages, pricing, and October 20 2026 sunset date confirmed unchanged via secondary source; replaced_by_model_id updated to sonic-3.6, Cartesia's new top-recommended model added this pass.
- **Cartesia Sonic 2** — Predecessor to sonic-3.5; still stable and callable, but Cartesia recommends sonic-3.5 for new builds. Latest snapshot sonic-2-2025-06-11, supporting 8 core languages (en, fr, de, es, pt, zh, ja, ko). The 7 additional languages previously carrying a 2026-06-01 EOL are no longer listed as supported on this snapshot as of 2026-07-02; languages field reduced from 15+ accordingly. 90ms model latency. Higher-fidelity voice cloning capability. Pricing equality with sonic-3.5 now documented (~1 credit/char on every TTS endpoint per https://docs.cartesia.ai/pricing); USD anchor Pro plan $5 per 100K credits => $50/1M chars. Confidence medium because the USD-per-credit rate varies by plan tier. Model sunsets October 20, 2026 per https://docs.cartesia.ai/build-with-cartesia/tts-models/older-models. 2026-07-27 refresh: pricing, snapshot, and language count confirmed unchanged; sunset date noted. 2026-08-11 refresh: pricing, snapshot, and language count confirmed unchanged via secondary source. 2026-09-02 refresh: pricing, snapshot, language count, and October 20 2026 sunset date confirmed unchanged via secondary source; replaced_by_model_id updated to sonic-3.6, Cartesia's new top-recommended model added this pass.
- **Cartesia Sonic Turbo** — Lowest-latency Sonic variant (~40ms TTFB). Still stable and callable, but Cartesia recommends sonic-3.5 for new builds. Latest snapshot sonic-turbo-2025-06-04, supporting 9 languages (en, fr, de, es, pt, zh, ja, hi, ko). The 6 additional languages previously carrying a 2026-06-01 EOL are no longer listed as supported on this snapshot as of 2026-07-02; languages field reduced from 15+ accordingly. Pricing equality with sonic-3.5 now documented (~1 credit/char on every TTS endpoint per https://docs.cartesia.ai/pricing); USD anchor Pro plan $5 per 100K credits => $50/1M chars. Confidence medium because the USD-per-credit rate varies by plan tier. Model sunsets October 20, 2026 per https://docs.cartesia.ai/build-with-cartesia/tts-models/older-models. 2026-07-27 refresh: pricing, snapshot, and language count confirmed unchanged; sunset date noted. 2026-08-11 refresh: pricing, snapshot, and language count confirmed unchanged via secondary source. 2026-09-02 refresh: pricing, snapshot, language count, and October 20 2026 sunset date confirmed unchanged via secondary source; replaced_by_model_id updated to sonic-3.6, Cartesia's new top-recommended model added this pass.
- **Groq Canopy Labs Orpheus English (Groq)** — Hosted on Groq. Output speed ~100 characters/second. Vendor API model_id is canopylabs/orpheus-v1-english per console.groq.com/docs/models; this row keeps the pre-existing Hail slug for stability. Re-verified 2026-07-19 against groq.com/pricing and console.groq.com/docs/models: price and model_id unchanged; still listed as a preview model. 2026-08-11 refresh: groq.com/pricing now 308-redirects to the marketing homepage (no pricing content) and groq.com/groqcloud-models 404s; verified instead via console.groq.com/docs/models (price $22.00/1M chars, model_id, preview status unchanged) and console.groq.com/docs/deprecations (no deprecation scheduled for this model; playai-tts, its pre-Orpheus predecessor, was already shut down 2025-12-31). 2026-09-02 refresh: verified via console.groq.com/docs/models (price $22.00/1M chars, model_id, preview status unchanged) and console.groq.com/docs/deprecations (no new deprecation scheduled). No new Orpheus language variants found — only English and Arabic Saudi listed.
- **Groq Canopy Labs Orpheus Arabic Saudi (Groq)** — Hosted on Groq. Saudi Arabic variant. Output speed ~100 characters/second. Vendor API model_id is canopylabs/orpheus-arabic-saudi per console.groq.com/docs/models; this row keeps the pre-existing Hail slug for stability. Re-verified 2026-07-19 against groq.com/pricing and console.groq.com/docs/models: price and model_id unchanged; still listed as a preview model. 2026-08-11 refresh: groq.com/pricing now 308-redirects to the marketing homepage (no pricing content) and groq.com/groqcloud-models 404s; verified instead via console.groq.com/docs/models (price $40.00/1M chars, model_id, preview status unchanged) and console.groq.com/docs/deprecations (no deprecation scheduled for this model; playai-tts-arabic, its pre-Orpheus predecessor, was already shut down 2025-12-31). 2026-09-02 refresh: verified via console.groq.com/docs/models (price $40.00/1M chars, model_id, preview status unchanged) and console.groq.com/docs/deprecations (no new deprecation scheduled). No new Orpheus language variants found.
- **Google Google Cloud TTS — Studio** — Google's premium TTS tier for professional media production (long-form narration, advertising). Vendor price $0.00016/char = $160/1M chars (sku 84AB-48C0-F9C3); the single-speaker Studio class is GA and the multispeaker class is experimental per https://docs.cloud.google.com/text-to-speech/docs/voices. SSML supported except <mark>, <emphasis>, <prosody pitch>, and <lang>. Model_id is a Hail-coined tier slug (Google bills per-voice-tier rather than per API model name). Free tier: first 1M chars/month included per current pricing table (earlier rows recorded 100K; free tier is not representable in schema free_tier fields, and free-tier size is not a structured price field, so last_changed_at is not bumped). 2026-07-19 refresh: clean primary-source read achieved — raw curl of the pricing page HTML yielded the full pricing table (WebFetch still truncates). Studio is now grouped under the page's 'Legacy TTS models' section (no deprecation notice; still GA per secondary docs/voices). Price unchanged; confidence raised to high (primary table + secondary docs). 2026-08-11 refresh: price unchanged ($160/1M, sku 84AB-48C0-F9C3, still under 'Legacy TTS models'); no deprecation banner; secondary docs/voices corroborates GA status. 2026-09-02 refresh: price unchanged ($160/1M, sku 84AB-48C0-F9C3, still under 'Legacy TTS models'); no deprecation banner; secondary docs/voices (single-speaker Studio GA, multispeaker experimental) corroborates.
- **Google Google Cloud TTS — Neural2** — Google's recommended general-purpose neural tier (newer architecture than WaveNet; no longer the same rate — WaveNet dropped to the $4/1M Standard rate). Vendor price $0.000016/char = $16/1M chars (sku FEBD-04B6-769B, shared with the Polyglot preview tier). SSML fully supported. Model_id is a Hail-coined tier slug (Google bills per-voice-tier rather than per API model name). Free tier: first 1M chars/month included (not representable in schema free_tier fields). 2026-07-19 refresh: clean primary-source read achieved — raw curl of the pricing page HTML yielded the full pricing table (WebFetch still truncates). Neural2 now sits under the page's 'Legacy TTS models' section (no deprecation notice; still GA per secondary docs/voices). Price unchanged; confidence raised to high (primary table + secondary docs). 2026-08-11 refresh: price unchanged ($16/1M, sku FEBD-04B6-769B, still under 'Legacy TTS models'); no deprecation banner; secondary docs/voices corroborates GA status. 2026-09-02 refresh: price unchanged ($16/1M, sku FEBD-04B6-769B, still under 'Legacy TTS models'); no deprecation banner; secondary docs/voices corroborates GA status.
- **Google Google Cloud TTS — WaveNet** — Original neural-net voice family from DeepMind. PRICE CHANGE 2026-07-19: $16/1M -> $4/1M chars. A clean primary-source read (raw curl of the pricing page HTML; WebFetch still truncates) shows WaveNet now listed under 'Legacy TTS models' and billed on the same SKU as Standard voices (sku 9D01-5995-B545) at $0.000004/char = $4/1M chars, with the free tier enlarged to the first 4M chars/month (was 1M). Cross-confirmed by third-party aggregators (costbench.com, diyai.io, xpay.sh) all listing WaveNet at $4/1M as of 2026-07; the WaveNet doc page (docs.cloud.google.com/text-to-speech/docs/wavenet) no longer states a price and defers to the pricing page. This retroactively validates the 2026-07-02 aggregator claim of $4/1M that the 07-13 pass dismissed as a tier mismatch. Not deprecated — no deprecation banner; still GA per secondary docs/voices — but Google recommends Neural2 ($16/1M) or Chirp 3: HD for new projects. SSML fully supported. Model_id is a Hail-coined tier slug. 2026-08-11 refresh: price unchanged ($4/1M, sku 9D01-5995-B545, 4M-char free tier, still under 'Legacy TTS models'); no deprecation banner; secondary docs/voices corroborates GA status. 2026-09-02 refresh: price unchanged ($4/1M, sku 9D01-5995-B545, 4M-char free tier, still under 'Legacy TTS models'); no deprecation banner; secondary docs/voices corroborates GA status.
- **Google Google Cloud TTS — Chirp 3: HD** — Google's newest-generation TTS family with 30 voice styles in 30+ languages. Vendor price $0.00003/char = $30/1M chars (sku F977-2280-6F1B), listed under the pricing page's 'Latest TTS models' section. Per https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd, Chirp 3: HD explicitly does NOT support SSML, speaking-rate adjustments, or pitch parameters; streaming synthesis IS supported. Model_id is a Hail-coined tier slug. Free tier: first 1M chars/month included. 2026-07-19 refresh: clean primary-source read achieved — raw curl of the pricing page HTML yielded the full pricing table (WebFetch still truncates); tier still GA with unchanged SSML/streaming behavior per secondary docs/voices. Price unchanged; confidence raised to high (primary table + secondary docs). 2026-08-11 refresh: price unchanged ($30/1M, sku F977-2280-6F1B, still under 'Latest TTS models'); no deprecation banner; secondary docs/voices corroborates GA status and no-SSML behavior. 2026-09-02 refresh: price unchanged ($30/1M, sku F977-2280-6F1B, still under 'Latest TTS models'); no deprecation banner; secondary docs/voices corroborates GA status and no-SSML behavior.
- **Google Google Cloud TTS — Chirp 3: Instant custom voice** — Row added 2026-07-19. Chirp 3: Instant custom voice — voice cloning from ~10s of reference audio plus a recorded consent statement, synthesized via a reusable voice cloning key. Vendor price $0.00006/char = $60/1M chars (sku A247-37D7-C094), listed under the pricing page's 'Latest TTS models' section; no free tier ('Not available' on the pricing table). Sources: primary pricing page (full table via raw curl of the page HTML) + https://docs.cloud.google.com/text-to-speech/docs/chirp3-instant-custom-voice. 34 languages; streaming and long-form synthesis supported; supports Chirp 3 pace controls and pause tags (SSML support not documented, so ssml_supported omitted). Model_id is a Hail-coined tier slug. confidence medium: the feature is in preview and access is restricted to allow-listed users, though the price itself is published on the primary pricing table. 2026-08-11 refresh: price unchanged ($60/1M, sku A247-37D7-C094, no free tier, still under 'Latest TTS models'); still preview/allow-listed, confidence remains medium; secondary docs/voices does not list this row (stale relative to the pricing page, not a contradiction). 2026-09-02 refresh: price unchanged ($60/1M, sku A247-37D7-C094, no free tier, still under 'Latest TTS models'); still preview/allow-listed, confidence remains medium; secondary docs/voices still does not list this row.
- **Microsoft Azure Azure AI Speech — Neural** — Azure's standard neural TTS tier (called 'Neural' on the pricing page; 'Standard voice' in docs). 500+ prebuilt voices across 100+ locales per https://learn.microsoft.com/en-us/azure/ai-services/speech-service/text-to-speech (reconfirmed 2026-07-13: 100+ languages/locales, full SSML support, unchanged). S0 pay-as-you-go price for both real-time and batch synthesis (HD, AOAI, Custom Neural Voice, and Personal Voice priced separately). Chinese characters counted as 2 chars for billing. Free tier (F0): 500K chars/month. 2026-07-13 refresh: primary marketing pricing page again unreachable via WebFetch (timeout across 4 attempts, multiple URL variants — same persistent issue noted in the 2026-07-02 refresh). Queried Microsoft's own public Azure Retail Prices API directly (https://prices.azure.com/api/retail/prices, meterName 'S1 Neural Text To Speech Characters', productName 'Azure Speech', meterId 0f98e708-a16c-407b-8089-a0ed9e14ab49): retailPrice is $15.00/1M chars for the Global meter and all commercial regions (US Gov regions bill $18.75/1M), with effectiveStartDate 2024-02-01 — i.e. this rate has been continuously in effect since Feb 2024, not a fresh vendor price change this week. This is Microsoft's own first-party billing-meter API, more authoritative than the marketing page or any third-party recap. price_per_1m_chars_usd corrected from '16.0' to '15.0' to match the live meter — this appears to fix a stale figure carried forward from earlier refreshes. Note: multiple third-party aggregators (costbench.com, texttolab.com) and general web search still report '$16/1M' as of mid-2026, apparently echoing each other or a stale marketing-page snapshot rather than the live billing meter; none could be reconciled against prices.azure.com. confidence held at medium pending a clean fetch of the live marketing pricing page to confirm its displayed sticker price matches the billing meter. 2026-07-19 refresh: marketing pricing page still unreachable via WebFetch (ETIMEDOUT). $15.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId 0f98e708-a16c-407b-8089-a0ed9e14ab49, retailPrice 15.0 per 1M for the Global meter and all commercial regions, effectiveStartDate 2024-02-01 — queried 2026-07-19). 100+ languages/locales, full SSML support, and Chinese-characters-count-as-2 billing reconfirmed via learn.microsoft.com/en-us/azure/ai-services/speech-service/text-to-speech (doc now brands the service 'Azure Speech in Foundry Tools'; tier still called 'Neural' on the pricing page / 'Standard voice' in docs). No structured field changed. 2026-08-11 refresh: primary marketing pricing page still rendered client-side '$-' placeholders (WebFetch reached the page but prices are JS-populated, not a fetch failure). $15.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId 0f98e708-a16c-407b-8089-a0ed9e14ab49, Global meter, effectiveStartDate still 2024-02-01). 100+ languages/locales and full SSML support reconfirmed via learn.microsoft.com/en-us/azure/ai-services/speech-service/text-to-speech (page unchanged since the 2026-07-13 pass, still branded 'Azure Speech in Foundry Tools'). No structured field changed. 2026-09-02 refresh: primary marketing pricing page again rendered client-side '$-' placeholders (WebFetch reached the page, prices are JS-populated). $15.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId 0f98e708-a16c-407b-8089-a0ed9e14ab49, Global meter, effectiveStartDate still 2024-02-01, queried directly with curl since the field is only visible via the raw API, not the WebFetch-rendered page). 100+ languages/locales and SSML support reconfirmed unchanged via learn.microsoft.com/en-us/azure/ai-services/speech-service/text-to-speech (still branded 'Azure Speech in Foundry Tools'). No structured field changed.
- **Microsoft Azure Azure AI Speech — Neural HD (DragonHD)** — Azure's premium HD neural tier (DragonHD architecture, 30+ GA voices). Per https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/azure-speech-%E2%80%93-neural-hd-text-to-speech-recent-voice-updates/4505380, Azure reduced Neural HD pricing to $22/1M chars effective March 2026 (down from $30/1M). Latency <300ms, real-time only. SSML support is partial; per the current supported/unsupported SSML table at learn.microsoft.com/en-us/azure/ai-services/speech-service/high-definition-voices, DragonHD (non-Omni) supports <voice>, <lang>, <phoneme>, <lexicon> (alias only), <say-as>, <sub>, <break>, <p>, <s>, but NOT <mstts:express-as>, <prosody>, <emphasis>, <audio>, <mstts:audioduration>, <mstts:backgroundaudio>, <math>, <bookmark>, <mstts:silence>, <mstts:viseme>; we mark ssml_supported=false because <prosody>/<emphasis>, the elements most callers want, are unsupported (2026-07-13: corrected note — <break> IS actually supported for DragonHD per the current doc table, unlike what a prior refresh's note implied). Automatic emotion/sentiment detection drives delivery (emotion_control_supported=true). DragonHDOmni (700+ voices, mixed GA/preview) and DragonHDFlash (en-US/zh-CN only) are distinct models tracked separately if added later. Confidence medium because the pricing-page value was originally sourced via third-party recap (techcommunity blog) — verify on the live Azure pricing page before high-volume use. 2026-07-13 refresh: primary marketing pricing page again unreachable via WebFetch (timeout across 4 attempts, multiple URL variants — same persistent issue as 2026-07-02). $22/1M reconfirmed unchanged via Microsoft's own Azure Retail Prices API (https://prices.azure.com/api/retail/prices, meterName 'Neural HD Text to Speech Characters', productName 'Azure Speech', meterId ad55d150-6182-5565-acf9-364d6ed979fb): retailPrice $22.00/1M chars (Global meter), effectiveStartDate 2026-03-01 — matches the March 2026 cut recorded here exactly, confirming no further change since. 30+ voice count, the en-US/de-DE/es-ES/fr-FR/ja-JP/zh-CN locale set, and the SSML support table reconfirmed via a direct fetch of learn.microsoft.com/en-us/azure/ai-services/speech-service/high-definition-voices (28 named DragonHD voices listed, 3 of them Preview status; six locales exactly matching this row). DragonHD Omni (700+ voices) and DragonHD Flash (zh-CN/en-US optimized variants, confirmed via the same doc to list 17 named Flash voices) still have no distinct billing meter in the Azure Retail Prices API (checked directly, no 'Omni' or 'Flash' Speech meters found) — deferred as new rows pending a priced, in-file-verifiable model_id. 2026-07-19 refresh: marketing pricing page still unreachable via WebFetch (ETIMEDOUT). $22.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId ad55d150-6182-5565-acf9-364d6ed979fb, Global + regional meters all 22.0 per 1M, effectiveStartDate 2026-03-01 — queried 2026-07-19). high-definition-voices doc re-checked: 28 named DragonHD voices (3 Preview) across the same six locales, comparison table still states 30 voices, real-time only, <300ms latency, styles/paralinguistics supported (emotion_control_supported=true stands), <prosody>/<emphasis> still unsupported (ssml_supported=false stands). Full Azure Speech Global consumption-meter list re-pulled 2026-07-19: still no distinct meter for DragonHD Omni (now documented at 700+ voices), DragonHD Flash (17 voices, zh-CN/en-US), or the new MAI-Voice-1/MAI-Voice-2 base models — all remain deferred. Personal Voice does have its own meter and was added as row azure-tts-personal-voice this pass. No structured field changed. 2026-08-11 refresh: marketing pricing page rendered client-side '$-' placeholders again (not unreachable, just JS-populated). $22.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId ad55d150-6182-5565-acf9-364d6ed979fb, Global meter, effectiveStartDate still 2026-03-01). high-definition-voices doc re-checked: same 28 named DragonHD voices (3 Preview: Andrew3, Ava3, MultiTalker-Ava-Andrew) across the same six locales, comparison table still states 30, real-time only, <300ms latency, <prosody>/<emphasis> still unsupported, <break> still supported. DragonHD Omni (now 700+ voices), DragonHD Flash (17 voices), and MAI-Voice-1/MAI-Voice-2 still have no distinct billing meter in the Azure Retail Prices API — remain deferred. No structured field changed. 2026-09-02 refresh: marketing pricing page still rendered client-side '$-' placeholders. $22.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId ad55d150-6182-5565-acf9-364d6ed979fb, Global meter, effectiveStartDate still 2026-03-01). high-definition-voices doc re-checked: same 28 named DragonHD voices (3 Preview: Andrew3, Ava3, MultiTalker-Ava-Andrew) across the same six locales, comparison table still states 30 voices, <300ms latency, real-time only, <prosody>/<emphasis> still unsupported, <break> still supported — no field changes. Doc's top summary table now describes DragonHDOmni as '500+ voices (all released voices)' while the body and comparison table elsewhere on the same page still say '700+ voices' — an internal inconsistency in Microsoft's own doc, not a product change; treating 700+ as authoritative since it's stated in two of three places and matches prior passes. Queried the Azure Retail Prices API directly for serviceName='Foundry Tools' with contains(meterName,'Omni'|'Flash'|'MAI'|'Dragon'): zero results for all four — DragonHD Omni, DragonHD Flash, and MAI-Voice-1/MAI-Voice-2 still have no distinct billing meter and remain deferred. No structured field changed.
- **Microsoft Azure Azure AI Speech — Personal Voice** — Azure's zero-shot voice-cloning tier ('Personal voice'): creates a voice from ~1 minute of speech plus a recorded verbal consent statement, then synthesizes in 91 languages / 100+ locales with automatic sentence-level language detection (https://learn.microsoft.com/en-us/azure/ai-services/speech-service/personal-voice-overview). Limited-access feature: the API is restricted to approved customers and use cases (Microsoft intake form). Row added 2026-07-19; the tier has been priced since 2024-02 but was previously left out ('priced separately' in the azure-tts-neural notes). Synthesis billed per character at $24/1M, verified via Microsoft's first-party Azure Retail Prices API (meterName 'Text to Speech - Personal Voice Characters', productName 'Azure Speech', skuName 'Text to Speech - Personal Voice', retailPrice 24.0 per 1M, Global meter, effectiveStartDate 2024-02-01). Separate voice-profile storage charge is NOT represented in structured fields: meter 'Text to Speech - Personal Voice Voice Storage' bills $600 per 1K units (docs: each stored profile is billed per voice per day until deleted; partial days count as full days). Synthesis uses a base model selected in SSML via <voice name> + <mstts:ttsembedding speakerProfileId>: PhoenixLatestNeural (~200ms latency), DragonLatestNeural (~500ms), DragonHDOmniLatestNeural (~300ms, styles/paralinguistics), MAI-Voice-1/MAI-Voice-2 — per https://learn.microsoft.com/en-us/azure/ai-services/speech-service/personal-voice-how-to-use. ssml_supported=false by the same rule as azure-tts-hd: <prosody> pitch/contour/range/volume and <emphasis> are unsupported in both Phoenix and Dragon (prosody rate IS supported); emotion_control_supported=false because <mstts:express-as> is unsupported for personal voice. Real-time synthesis via the standard Speech SDK/REST endpoints (streaming_supported=true). Chinese characters counted as 2 chars for billing (same policy as other Azure TTS tiers). confidence medium: the $24/1M figure rests on the first-party billing-meter API plus learn.microsoft.com docs confirming the per-character billing model; the marketing pricing page (primary source) remains unfetchable (ETIMEDOUT 2026-07-19, persistent issue since 2026-07-02) and personal voice pricing is only displayed there for regions where the feature is available. 2026-08-11 refresh: marketing pricing page rendered client-side '$-' placeholders (not a fetch failure). $24.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId 2172330f-8d48-59d4-a35a-24c604f8cae1, Global meter, effectiveStartDate still 2024-02-01; voice-storage meter 4ba2e583-8439-5f6b-887f-1a9aca8a48a9 also unchanged at $600/1K). personal-voice-overview doc re-checked: still limited-access (intake-form approval required), still 91 languages / 100+ locales, automatic sentence-level language detection unchanged. No structured field changed. 2026-09-02 refresh: marketing pricing page still rendered client-side '$-' placeholders. $24.00/1M reconfirmed unchanged via Azure Retail Prices API (same meterId 2172330f-8d48-59d4-a35a-24c604f8cae1, Global meter, effectiveStartDate still 2024-02-01). personal-voice-overview doc re-checked: still limited-access (intake-form approval required), still '90+ languages' in the intro / '91 languages across 100+ locales' in the how-to-create section — same wording as prior passes, no field changes. No structured field changed.
- **Inworld Inworld Realtime TTS-2** — Inworld's flagship Realtime TTS-2 model with natural-language steering. Pricing tiers restructured since last check: On-Demand (PAYG) rate now $25/1M chars (was $35), Creator $20, Builder $17.50, Developer $15, Growth $12.50, Enterprise as low as $5/1M. Instant voice cloning, custom pronunciation, and enhanced timestamp/viseme alignment included. The older inworld-tts-1 and inworld-tts-1-max were discontinued June 15, 2026 per docs.inworld.ai/tts/tts-models (auto-rerouted to inworld-tts-1.5-mini and inworld-tts-1.5-max respectively, neither of which was previously tracked — added as separate rows in a prior pass). 2026-07-19 refresh: pricing page and docs re-verified via WebFetch, all tiers and rates unchanged; no new models launched; no further deprecation activity. 2026-08-11 refresh: pricing page reconfirmed all tiers/rates unchanged. docs.inworld.ai now lists a new model 'Realtime TTS-2 Flash' (inworld-tts-2-flash, ~20ms TTFB) — deferred, not added: not mentioned on the pricing page and no price found on either source, so the two-source rule can't be met. Same docs page also now states inworld-tts-1.5-max and inworld-tts-1.5-mini are 'deprecated and not recommended for new projects' (see those rows for detail); TTS-2 remains the recommended flagship and unaffected by that note. 2026-09-02 refresh: On-Demand rate reconfirmed unchanged at $25/1M via pricing page. Languages field corrected from '100+' to '200+': the pricing page's Features comparison table now explicitly states '200+ languages' (matching docs.inworld.ai, which previously was the lone source for that figure — no longer a source conflict). 'Realtime TTS-2 Flash' now has verified On-Demand pricing ($15/1M chars) on the pricing page plus a confirmed model_id (inworld-tts-2-flash) on docs.inworld.ai — two-source rule met, added as a new row. No deprecation changes; TTS-2 remains the recommended flagship.
- **Inworld Inworld Realtime TTS-2 Flash** — New row, added 2026-09-02. Low-latency sibling of inworld-tts-2: 'our lowest latency — 20ms TTFB, 5x faster than inworld-tts-2', same 200+ languages/locales, high-quality instant voice cloning (per docs.inworld.ai/tts/tts-models). On-Demand (PAYG) rate $15/1M chars; Creator $10, Builder $9, Developer $8, Growth $7, Enterprise sub-$5/1M (per inworld.ai/pricing). model_id inworld-tts-2-flash confirmed on docs.inworld.ai/tts/tts-models; streaming_supported and deployment_options inferred from the 'Realtime' product family naming and TTFB metric, consistent with the sibling inworld-tts-2 row — not separately itemized on either source. Previously observed on docs.inworld.ai during the 2026-08-11 pass but deferred for lack of pricing; now meets the two-source rule.
- **Inworld Inworld Realtime TTS 1.5 Max** — Alternative to TTS-2, tuned for maximum stability across 15 languages. Vendor-reported <200ms median latency. On-Demand (PAYG) rate $35/1M chars; Creator $25, Builder $22.50, Developer $20, Growth $17.50, Enterprise custom. Instant voice cloning and enhanced timestamps included. Successor to discontinued inworld-tts-1-max per https://docs.inworld.ai/tts/tts-models. 2026-07-19 refresh: rate and language list reconfirmed unchanged via pricing page and docs. 2026-08-11 refresh: pricing page still lists this row with full active pricing (no deprecation banner, rate unchanged at $35/1M), but docs.inworld.ai/tts/tts-models now states 'previous-generation models (inworld-tts-1, inworld-tts-1-max, inworld-tts-1.5-max, inworld-tts-1.5-mini) are deprecated and not recommended for new projects,' with no vendor-published sunset date. Primary/secondary sources conflict on deprecation status; deprecated_at set to this verification date per the same convention used for eleven_turbo_v2_5 (still callable and priced, docs recommend migrating). replaced_by_model_id points to inworld-tts-2, the only in-file successor docs recommend; a faster inworld-tts-2-flash model was also observed on docs.inworld.ai but has no verifiable pricing yet (deferred, see inworld-tts-2 notes). 2026-09-02 refresh: pricing page no longer lists this row at all (only TTS-2 and TTS-2 Flash appear); docs.inworld.ai still lists it under 'deprecated and not recommended for new projects' with no sunset date. $35/1M chars is the last-confirmed price, kept since no live successor price applies to prior usage and the row still requires a required price field; treat as fully retired from new sign-up. replaced_by_model_id kept at inworld-tts-2 (the quality/steerability-tier successor); no vendor-published sunset date yet, so deprecated_at is unchanged from 2026-08-11 (the date this pass first identified deprecation).
- **Inworld Inworld Realtime TTS 1.5 Mini** — Low-latency tier across the same 15 languages as 1.5 Max. Vendor-reported ~120ms median latency (~160ms P90 first-chunk). On-Demand (PAYG) rate $15/1M chars; Creator $10, Builder $9, Developer $8, Growth $7, Enterprise custom. Instant voice cloning and enhanced timestamps included; also available for on-premises deployment. Successor to discontinued inworld-tts-1 per https://docs.inworld.ai/tts/tts-models. 2026-07-19 refresh: rate and language list reconfirmed unchanged via pricing page and docs. 2026-08-11 refresh: pricing page still lists this row with full active pricing (no deprecation banner, rate unchanged at $15/1M), but docs.inworld.ai/tts/tts-models now states 'previous-generation models (inworld-tts-1, inworld-tts-1-max, inworld-tts-1.5-max, inworld-tts-1.5-mini) are deprecated and not recommended for new projects,' with no vendor-published sunset date. Primary/secondary sources conflict on deprecation status; deprecated_at set to this verification date per the same convention used for eleven_turbo_v2_5 (still callable and priced, docs recommend migrating). replaced_by_model_id points to inworld-tts-2, the only in-file successor docs recommend; a low-latency inworld-tts-2-flash model was also observed on docs.inworld.ai (positioned as this row's likely direct successor) but has no verifiable pricing yet (deferred, see inworld-tts-2 notes) — re-check next pass. 2026-09-02 refresh: pricing page no longer lists this row at all (only TTS-2 and TTS-2 Flash appear); docs.inworld.ai still lists it under 'deprecated and not recommended for new projects.' $15/1M chars kept as the last-confirmed price for prior usage; treat as fully retired from new sign-up. inworld-tts-2-flash now has verified pricing and a confirmed model_id, and is the closer low-latency successor by design intent, so replaced_by_model_id is updated from inworld-tts-2 to inworld-tts-2-flash. deprecated_at unchanged (2026-08-11, the date this pass first identified deprecation); no vendor-published sunset date.
- **Smallest.ai Lightning v3.1** — Smallest.ai's current standard TTS model (44.1 kHz native, ~200ms TTFB at 40 concurrent requests). 2026-07-02 refresh: canonical API model_id corrected from 'lightning-v3.1' (hyphen, Hail-coined) to 'lightning_v3.1' (underscore) per exact string confirmed in code samples on https://docs.smallest.ai/waves/model-cards/text-to-speech/lightning-v-3-1 ('"model": "lightning_v3.1"'); old id kept as alias. Price dropped from $25/1M to $17.5/1M chars (~$0.175/10k chars), confirmed via rendered pricing-page text (both AI-summarized and raw-HTML grep of the live page); note the page's stale JSON-LD schema.org markup still advertises the old ~$0.25/10k figure for its 'Pro' *subscription* tier description, so confidence downgraded to medium pending a secondary source that states the exact number (docs.smallest.ai defers pricing to the pricing page and does not restate figures). Output formats corrected to match vendor's exact list (added alaw; mulaw renamed to ulaw per vendor terminology) — sample rates and voice/language counts unchanged. New sibling model added this refresh: lightning_v3.1_pro. Lightning v2 and lightning-large remain deprecated per https://docs.smallest.ai/waves/documentation/getting-started/models — new integrations should use lightning_v3.1 or lightning_v3.1_pro. WebSocket streaming for real-time/conversational use. Instant + professional voice cloning supported. On-prem available on Enterprise plan. 2026-07-13 refresh: re-verified via smallest.ai/pricing ($17.5/1M chars unchanged, no deprecation banner) and docs.smallest.ai model card (12-language list, voice cloning support, sample rates/output formats, 200ms TTFB all unchanged). Vendor's getting-started/models catalog still lists only lightning_v3.1 and lightning_v3.1_pro as active — no new TTS model launches found. 2026-07-19 refresh: vendor restructured pricing — smallest.ai/pricing is now an agent-pricing page (per-minute layer rates, no 'lightning' or per-character figures) with model pricing moved to smallest.ai/pricing/models; source_url updated accordingly, not a price change. Raw-HTML grep of the new page confirms Lightning V3.1 at ~$0.175/10K chars = $17.5/1M, unchanged. Docs catalog (getting-started/models) still lists only lightning_v3.1 and lightning_v3.1_pro as active TTS (new Electron LLM, Pulse/Pulse Pro STT, Hydra S2S are out of TTS scope); docs still defer exact pricing to the pricing page, so confidence stays medium. 2026-08-11 refresh: re-verified via smallest.ai/pricing/models ($17.5/1M chars, i.e. $0.175/10K, unchanged) and docs.smallest.ai model card (12 trained-voice languages, 217 voices, voice cloning, sample rates/output formats, 200ms TTFB all unchanged). docs.smallest.ai/waves/documentation/getting-started/models still lists only lightning_v3.1 and lightning_v3.1_pro as active TTS; Hydra (speech-to-speech, beta) is out of TTS scope. No new TTS launches found; no field changes on this row. 2026-09-02 refresh: re-verified via smallest.ai/pricing/models ($17.5/1M chars, i.e. $0.175/10K, unchanged) and docs.smallest.ai/waves/model-cards/text-to-speech/lightning-v-3-1 model card, which confirms the 12-trained-voice-language set (en, hi, ta, es, kn, mr, te, or, pa, ml, gu, bn) recorded in this row's `languages` field is correct — the 20-code figure quoted on some docs pages includes 8 additional codes (fr, de, it, nl, sv, pt, pl, ru) that are routed through English/Hindi voices rather than having dedicated trained voices, so this row correctly lists only the 12 trained languages per existing convention. 217 voices, voice cloning, output formats/sample rates, and 200ms TTFB ("at 40 concurrent requests" per model card) all reconfirmed unchanged. docs.smallest.ai/waves/documentation/text-to-speech-lightning/overview reconfirms Lightning v2 deprecated in favor of v3.1/v3.1 Pro; no new TTS model launches found. No field changes on this row.
- **Smallest.ai Lightning v3.1 Pro** — New row (2026-07-02): premium 44.1 kHz-native pool with a curated voice catalog (American, British, and Indian accents), same ~200ms TTFB latency profile as standard Lightning v3.1. Voice cloning explicitly NOT available on the Pro pool per https://docs.smallest.ai/waves/model-cards/text-to-speech/lightning-v-3-1-pro ('No voice cloning. Voice cloning is not available on the Pro pool.'). Model_id 'lightning_v3.1_pro' confirmed verbatim in both the pricing page audio-sample caption and the docs model card. Vendor rate ~$0.195/10k chars = $19.5/1M chars, confirmed via rendered pricing-page text (AI-summarized + raw-HTML grep); confidence medium because the secondary docs source does not restate the exact price (defers to pricing page). voices_count omitted — vendor describes a 'curated catalog' without publishing an exact number for this pool. 2026-07-13 refresh: languages field corrected 2->29 — the docs model card (docs.smallest.ai/waves/model-cards/text-to-speech/lightning-v-3-1-pro) now states 'English + Hindi, plus 27 more (9 Indian, 8 Asian & Middle Eastern, 10 European)' vs. the en/hi-only description recorded at row creation; not a price change so last_changed_at not bumped. Price ($19.5/1M), voice-cloning-unavailable status, sample rates/formats, and 200ms TTFB reconfirmed unchanged via smallest.ai/pricing and the docs model card. 2026-07-19 refresh: vendor restructured pricing — per-character model pricing moved from smallest.ai/pricing (now agent-focused, per-minute rates only) to smallest.ai/pricing/models; source_url updated accordingly, not a price change. Raw-HTML grep of the new page confirms Lightning V3.1 Pro at ~$0.195/10K chars = $19.5/1M, unchanged. Docs TTS overview reconfirms lightning_v3.1_pro active with 29-language support; docs still defer exact pricing to the pricing page, so confidence stays medium. 2026-08-11 refresh: TWO field changes found and applied. (1) voice_cloning flipped false->true: docs.smallest.ai/waves/model-cards/text-to-speech/lightning-v-3-1-pro now states 'Yes. Pass model: lightning-v3.1-pro to the [voice-cloning API] to clone onto the Pro pool, then use the resulting voice_id with TTS model: lightning_v3.1_pro' — a direct reversal of the 'No voice cloning' text recorded at row creation; corroborated by smallest.ai/pricing/models, which now states both Lightning models support 'voice cloning capabilities from 5 seconds of audio.' (2) languages expanded 29->31: docs model card now enumerates Dutch (nl) and Swedish (sv) among the European set (12 European + 8 Asian/Middle Eastern + 9 Indic + en/hi = 31), not present in the prior 29-language list; docs.smallest.ai/waves/documentation/text-to-speech-lightning/overview independently states 'Lightning v3.1 Pro: All 31 languages with dedicated voices', corroborating the count. Neither change is a price change, so last_changed_at not bumped. Price ($19.5/1M = $0.195/10K), sample rates/formats, and 200ms TTFB reconfirmed unchanged via smallest.ai/pricing/models and the docs model card. No new TTS launches found (Hydra speech-to-speech, beta, is out of TTS scope). 2026-09-02 refresh: re-verified via smallest.ai/pricing/models ($19.5/1M chars, i.e. $0.195/10K, unchanged) and docs.smallest.ai/waves/documentation/text-to-speech-lightning/overview, which reconfirms 'Lightning v3.1 Pro: All 31 languages with dedicated voices' and voice cloning support. Sample rates/output formats and 200ms TTFB reconfirmed unchanged. No new TTS launches found; no field changes on this row.
- **Rime Rime Mist v3** — Rime's current Mist-family flagship, optimized for lowest latency. Model_id, capabilities, and lack of custom pronunciation on mistv3 reconfirmed unchanged via https://docs.rime.ai/api-reference/models, which now quotes '~37ms P50' TTFA/TTFB for Mist v3 (tightened from the prior 'sub-100ms' figure; time_to_first_byte_ms updated 100->37, not a price field so last_changed_at not bumped). 2026-07-02 refresh: rime.ai/pricing was restructured (page timestamp Jul 1, 2026) from the old per-model $/char table to a flat 'Starter: $0.05/1K characters, Enterprise: custom' layout with no per-model breakdown visible for Mist/Coda/Arcana; three separate fetches of the live page and a web search found no text attributing $0.05/1K specifically to mistv3. A Mar 2025 Rime blog post (rime.ai/resources/new-pricing) shows an unrelated subscription-tier ladder ($5/100k, $19/500k, $99/3M, $249/10M chars/mo) that also doesn't differentiate by model, while an older post (rime.ai/resources/introducing-new-pricing) restates the legacy per-model figures matching this row's existing $30/1M for Mist alongside Arcana $40/1M and Coda $50/1M -- consistent with the value on file but not a fresh primary-source confirmation. Given the primary source no longer states a model-specific rate, kept the existing $30/1M value rather than guess; confidence downgraded to medium pending a vendor page that re-publishes per-model pricing or direct sales confirmation. Coda (184 voices, 6 languages incl. ES/FR/PT/DE/JA) and Arcana (now at v3, 94 voices, multilingual; Rime recommends migrating Arcana deployments to Coda) remain distinct models tracked separately if added later -- deferred again this refresh for the same pricing-opacity reason. Voice cloning not documented for Mist; available via Enterprise plan for custom voices. 2026-07-13 refresh: rime.ai/pricing (fetched, page timestamp Jul 10, 2026 6:27pm UTC) still shows only the flat Starter $0.05/1K-chars / Enterprise-custom layout with zero per-model breakdown -- same opacity as the 2026-07-02 pass, so the existing $30/1M value is kept unverified-by-primary and confidence stays medium. docs.rime.ai/api-reference/models reconfirms mistv3 as 'Current (Mar 2026)' with model_id, 94 voices, EN-only, streaming, no voice cloning, and ~37ms P50 latency all unchanged, and additionally documents two more models not in this dataset: mistv2 (model_id 'mistv2', 94 voices, EN+ES, ~70ms on-prem, supports custom pronunciation) and legacy 'mist' (model_id 'mist', deprecated in favor of v2/v3) -- neither is added this pass because, like Coda/Arcana, no source states a per-model price and the two-source price rule can't be met. No deprecation banner on mistv3 itself. 2026-07-19 refresh: rime.ai/pricing unchanged -- still flat Starter $0.05/1K chars (3,000 free minutes, 20 concurrent generations) / Enterprise custom, no per-model breakdown, so $30/1M is kept per the standing 2026-07-02 decision and confidence stays medium. docs.rime.ai/api-reference/models reconfirms mistv3 as current (released Mar 2026), 94 voices, EN-only, streaming, no custom pronunciation, TTFB 'well below 100ms', no deprecation notice. Docs now present Coda (released May 2026, 184 voices, EN/ES/FR/PT/DE/JA) as the flagship and mark Arcana as deprecated in favor of Coda; Coda, mistv2, and Arcana remain deferred -- still no source attributing a per-model price, so the two-source rule can't be met. 2026-08-11 refresh: rime.ai/pricing was restructured again (two separate fetches) and now shows an explicit per-model comparison table -- 'Mist v3: $0.03/1K characters (~$0.03/minute)' and 'Coda: $0.05/1K characters (~$0.05/minute)' -- resolving the opacity from the 2026-07-02..07-19 passes. $0.03/1K chars = $30/1M chars, an exact match for the existing price, so last_changed_at is not bumped; confidence raised back to (default) high since the primary source now states the per-model rate directly, satisfying the two-source rule together with docs.rime.ai. docs.rime.ai/api-reference/models (two fetches) shows two field changes for mistv3 vs the prior refresh: voices_count 94->78 and languages en-only->en/fr/de/es (Rime's docs describe Mist v3 as now shipping English, French, German, and Spanish voices); status remains 'Current', released March 2026, no deprecation banner, still no inline pronunciation control (supports custom pauses instead). Not in scope this pass but noted for a future add: docs.rime.ai now lists Coda at 184 voices / 8 languages (adds AR, HI to the prior EN/ES/FR/PT/DE/JA set) and flags 'Arcana: Sunset Aug 15, 2026' -- Arcana is being sunset in favor of Coda per Rime's own docs; Arcana was never added to this dataset (deferred each prior refresh for pricing opacity) so no deprecated_at/replaced_by_model_id action is needed here, but the sunset date is worth flagging for the next refresh pass. mistv2 and legacy 'mist' remain deferred -- still no source attributing a per-model price for either. 2026-09-02 refresh: rime.ai/pricing reconfirms the per-model comparison table -- 'Mist v3: $0.03/1K characters' -- unchanged ($30/1M), so last_changed_at not bumped. docs.rime.ai/api-reference/models reconfirms model_id, 78 voices, en/fr/de/es, streaming, no custom pronunciation, ~37ms P50 TTFB, no deprecation banner, all unchanged. Arcana no longer appears anywhere on either source (past its Aug 15, 2026 sunset date flagged last refresh); it was never a row in this dataset so no deprecated_at action needed. mistv2 (model_id 'mistv2', now listed 'Current', 138 voices, en/fr/de/es, ~175ms on-prem, custom pronunciation supported) still has no per-model price on the pricing page, so it remains deferred -- two-source price rule not met. Coda is added as a new row this refresh: the pricing page now states 'Coda: $0.05/1K characters' explicitly (the per-model breakdown that resolved the 2026-07-02..07-19 opacity for mistv3 also covers Coda), and docs.rime.ai/api-reference/models corroborates model_id 'coda', 253 voices across 9 languages (en/ar/fr/de/hi/it/ja/pt/es), HTTP+WebSocket streaming, and sub-100ms GPU-engine latency (imprecise range, not set as time_to_first_byte_ms) -- satisfying the two-source rule. ssml_supported, emotion_control_supported, and voice_cloning are not documented for Coda on either source, so left unset rather than guessed.
- **Rime Rime Coda** — New row (2026-09-02): Rime's flagship model for natural expression, released May 2026. Price $0.05/1K characters = $50/1M chars per rime.ai/pricing's per-model comparison table, corroborated by docs.rime.ai/api-reference/models which confirms model_id 'coda', 253 total voices across 9 languages (English, Arabic, French, German, Hindi, Italian, Japanese, Portuguese, Spanish), and HTTP+WebSocket streaming. Latency documented only as 'sub-100ms model latency on the GPU engine' plus 25-50ms cloud-API network RTT -- not a single figure, so time_to_first_byte_ms left unset rather than guessed. SSML support, emotion control, and voice cloning are not documented on either source for this model, so those fields are left unset. No deprecation banner. Rime's docs no longer mention Arcana (which prior refreshes flagged as sunsetting Aug 15, 2026) -- Coda appears to be its de facto successor, but Arcana was never a row in this dataset so no replaced_by_model_id linkage applies.
- **LMNT LMNT Blizzard** — LMNT's flagship Blizzard 2.0 model (canonical model_id 'blizzard' per https://docs.lmnt.com/models/overview). 31 languages with accent control, word timestamps, streaming, voice cloning, and speech sessions. Confidence medium because LMNT publishes plan-bundled pricing (Indie $10/mo for 200K chars + $0.05/1K overage; Pro $49/mo + $0.045/1K overage; Premium $199/mo + $0.035/1K overage) rather than a standalone PAYG per-char rate — $50/1M shown here is the Indie-tier overage rate. Free tier includes 15K characters/month with no overage rate (not representable as tokens_per_day in schema). Premium tier overage is $0.035/1K = $35/1M — large customers should benchmark on their own plan. 2026-07-02 refresh: re-verified via www.lmnt.com/pricing and docs.lmnt.com/models/overview — Indie/Pro/Premium plan pricing, overage rates, and Blizzard 2.0 as sole canonical model_id all unchanged; no new model_ids or deprecation notices found. 2026-07-13 refresh: re-verified via both sources again — Free $0/15K chars, Indie $10/mo+$0.05/1K, Pro $49/mo+$0.045/1K, Premium $199/mo+$0.035/1K, Blizzard 2.0 (model_id 'blizzard') as sole canonical model_id, 31-language/streaming/voice-cloning feature set, all unchanged; no new model_ids or deprecation notices found. 2026-07-19 refresh: re-verified via both sources — all plan pricing and overage rates unchanged, Blizzard 2.0 still sole canonical model_id; docs mention enterprise-only preview versions with no public model_id or pricing, so nothing addable under the canonical-model_id rule. 2026-08-11 refresh: re-verified via both sources — Free $0/15K chars, Indie $10/mo+$0.05/1K, Pro $49/mo (1.25M chars included)+$0.045/1K, Premium $199/mo (5.7M chars included)+$0.035/1K, all unchanged; Blizzard 2.0 (model_id 'blizzard') still sole canonical model_id, 31-language/streaming/voice-cloning feature set unchanged; docs.lmnt.com repeats the enterprise-preview-versions line with no public model_id/pricing, so nothing addable; no deprecation notices found. 2026-09-02 refresh: LMNT has shut down entirely. www.lmnt.com/pricing and docs.lmnt.com/models/overview both now display only "Our speech generation journey has come to an end. Thank you for being part of it."; app.lmnt.com's page title reads "LMNT has shut down" (confirmed via web search). No successor product or model_id exists, so replaced_by_model_id is omitted. Marking deprecated_at 2026-09-02 (date of first observation) and freezing price_per_1m_chars_usd at the last-verified $50/1M rate as a historical record; no further re-verification against a live pricing page is possible.
- **Deepgram Deepgram Aura 2** — Deepgram's current TTS family, addressed as 'aura-2-<voice>-<lang>' (e.g., aura-2-thalia-en) per https://developers.deepgram.com/docs/tts-models. Vendor rate $0.030/1K chars = $30/1M chars on Pay-As-You-Go; Growth tier $0.027/1K = $27/1M. Voice counts by language (91 total as of 2026-09-02, was 90): en 40, es 18 (was 17), nl 9, fr 2, de 7, it 10, ja 5. Free tier ships $200 of signup credit applicable to all products (not representable as tokens_per_day in schema). Aura 1 (legacy, 'aura-<voice>-<lang>', English-only, $0.0150/1K PAYG) remains callable but Deepgram recommends Aura 2 for new integrations; not tracked as a separate row (not a new launch). 2026-07-19 refresh: re-verified via deepgram.com/pricing and developers.deepgram.com/docs/tts-models — PAYG $0.030/1K and Growth $0.027/1K rates, 7-language/90-voice count, and streaming support all unchanged; no new Deepgram TTS model_ids found and no deprecation notice for Aura 1 or Aura 2. 2026-08-11 refresh: re-verified via deepgram.com/pricing (PAYG $0.030/1K, Growth $0.027/1K unchanged) and developers.deepgram.com/docs/tts-models (7-language voice family, naming convention, English-accent and Spanish codeswitching notes unchanged); no new Deepgram TTS model_ids found and no deprecation notice for Aura 1 or Aura 2. 2026-09-02 refresh: re-verified via deepgram.com/pricing (PAYG $0.030/1K, Growth $0.027/1K unchanged) and developers.deepgram.com/docs/tts-models (Spanish voice count moved 17→18, total 90→91; 7-language family, naming convention, streaming support otherwise unchanged); no new Aura TTS model_ids found and no deprecation notice for Aura 1 or Aura 2. Deepgram also launched Flux (flux-general-en, flux-general-multi) during this period, confirmed via developers.deepgram.com/docs/flux to be an STT/conversational product, not TTS — out of scope for this file.
- **Resemble AI Resemble Chatterbox Turbo** — Resemble bills TTS per second of generated audio at a flat $0.0005/sec on their Flex (pay-as-you-go) plan; the rate is not split per model. Chatterbox Turbo is Resemble's flagship English TTS per https://www.resemble.ai (also open-sourced at https://github.com/resemble-ai/chatterbox, 'SoTA open-source TTS'). Confidence medium because the pricing page lists the rate against the service category 'Text-to-speech' rather than naming Chatterbox Turbo specifically, and the API's exact 'model' parameter slug was not verified against live docs. Multilingual ('Chatterbox Multilingual') and dramatic-read ('DramaBox') variants are marketed separately; modeled here as the English flagship only. Re-verified 2026-07-02: rate unchanged at $0.0005/sec, model still listed as active on both resemble.ai/pricing and github.com/resemble-ai/chatterbox; no per-model pricing split found, so the model-id ambiguity noted above persists. Re-verified 2026-07-13: rate still $0.0005/sec flat on resemble.ai/pricing; Chatterbox Turbo (350M params, English, paralinguistic tags) confirmed still current per github.com/resemble-ai/chatterbox. Also checked docs.resemble.ai (TTS overview + API reference) for a documented 'model' request parameter that would resolve the model-id ambiguity — none found; no per-model pricing split exists. Pricing/marketing copy also names 'Chatterbox Nano' and 'Chatterbox Flash' alongside Turbo/Multilingual/DramaBox, but none have a discoverable canonical API model_id or distinct price, so not added as rows per the two-source + canonical-model_id rule. Re-verified 2026-07-19: resemble.ai/pricing has been restructured — its USAGE RATES section now lists only deepfake detection ($0.04/sec audio) and verification (watermark encode $0.0005/sec) rates; the TTS 'Text-to-speech' $0.0005/sec line is gone, and no TTS rate is published on the pricing page, docs.resemble.ai, or resemble.ai/products/text-to-speech (raw-HTML grep confirmed absence, not a fetch artifact). Chatterbox Turbo itself remains an active current product ('NEW', 350M params, English, paralinguistic tags) per resemble.ai/products/text-to-speech, resemble.ai/learn/models/chatterbox-turbo, and github.com/resemble-ai/chatterbox. Retaining the last published rate of $0.0005/sec (last confirmed live 2026-07-13) and lowering confidence to low, since the current rate is no longer publicly verifiable; no evidence of an actual price change, so last_changed_at not bumped. GitHub README now lists Chatterbox Multilingual V3 (500M, 23+ languages; V2 checkpoint legacy via t3_model='v2') plus six Single Language Pack finetunes, and marketing adds 'Chatterbox Pro' — none have a published hosted-API price or canonical API model_id, so still deferred per the two-source rule. Re-verified 2026-08-11: still no TTS rate published on resemble.ai/pricing (USAGE RATES section still lists only deepfake detection and watermark encode, no TTS line) or resemble.ai/products/text-to-speech; Chatterbox Turbo confirmed still active/current (350M params, English, paralinguistic tags, 'NEW' badge) per resemble.ai/products/text-to-speech and github.com/resemble-ai/chatterbox. No new pricing surfaced for Nano, Flash, or Pro variants either. Retaining last published rate of $0.0005/sec unchanged; last_changed_at not bumped (no evidence of an actual price move). Re-verified 2026-09-02: no change. resemble.ai/pricing still shows only deepfake-detection and watermark-encode usage rates, no TTS line. Chatterbox Turbo confirmed still active ('NEW' badge, 350M params, English, paralinguistic tags) per resemble.ai/products/text-to-speech and github.com/resemble-ai/chatterbox. GitHub README still lists Chatterbox-Nano (110M, English, CPU-optimized), Chatterbox-Multilingual V3 (500M, 23+ languages), six Single Language Pack finetunes, and legacy original Chatterbox (500M); marketing site adds Chatterbox Pro (enterprise tier) and DramaBox — none have a published hosted-API price or canonical API model_id, so all still deferred per the two-source rule. Retaining last published rate of $0.0005/sec unchanged; last_changed_at not bumped (no evidence of an actual price move).


---

_This page is generated from the JSON datasets at 2026-09-02T17:21:39.271Z. Verify any price against the provider's official pricing page before billing decisions; dataset is updated via [GitHub PRs](https://github.com/hail-hq/hail/tree/main/costs)._
