> This page is for version v2 API (default).
> For other versions, use one of these documentation indexes:
> - v2 API (default): https://docs.cohere.com/v2/llms.txt
> - v1 API: https://docs.cohere.com/v1/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.cohere.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server.

# Standard Vault Pricing

> Standard Vault pricing models (Fixed and Flex) and per-model performance tiers and rates.

Standard Vault is billed as a Cohere-managed service. Pricing depends on the models you select and
each model's [performance tier](#performance-tiers). Cohere manages the underlying infrastructure and
scaling, and customers can choose between two pricing models:

| Feature      | Fixed                                                                                            | Flex                                                              |
| ------------ | ------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------- |
| Commitment   | Monthly or annual                                                                                | Monthly or annual                                                 |
| Capacity     | Fixed number of instances (no autoscaling)                                                       | Minimum baseline instances, plus autoscaling                      |
| Sizing       | Determined through a sizing exercise or a production trial (for example, based on expected load) | --                                                                |
| Autoscaling  | --                                                                                               | Scales up/down based on request rate and agreed latency SLOs      |
| Pause/resume | You can pause a model or restart a paused model to save on costs.                                | You can pause a model or restart a paused model to save on costs. |
| Overages     | --                                                                                               | Additional capacity billed per instance-hour                      |
| Max capacity | --                                                                                               | Maximum instance cap per model                                    |

The following table summarizes the available models and their rates. All rates are per instance.

| Model             | Performance Tier | Hourly rate | Monthly rate | Annual rate |
| ----------------- | ---------------- | ----------- | ------------ | ----------- |
| Embed 3           | Small            | \$4.00      | \$2,500      | \$25,000    |
| Embed 4           | Small            | \$4.00      | \$2,500      | \$25,000    |
| Embed 4           | Medium           | \$5.00      | \$3,250      | \$32,500    |
| Embed 5 Fast      | Small            | \$3.00      | \$2,000      | \$20,000    |
| Embed 5 Fast      | Medium           | \$5.00      | \$3,250      | \$32,500    |
| Embed 5 Pro       | Small            | \$3.00      | \$2,000      | \$20,000    |
| Embed 5 Pro       | Medium           | \$5.00      | \$3,250      | \$32,500    |
| Rerank 3.5        | Medium           | \$5.00      | \$3,250      | \$32,500    |
| Rerank 4 Fast     | Medium           | \$5.00      | \$3,250      | \$32,500    |
| Rerank 4 Pro      | Medium           | \$5.00      | \$3,250      | \$32,500    |
| Rerank 4 Pro      | Large            | \$10.00     | \$6,500      | \$65,000    |
| Parse 5           | Medium           | \$4.00      | \$2,500      | \$25,000    |
| Parse 5           | Large            | \$8.00      | \$5,000      | \$50,000    |
| Parse 5           | XL               | \$7.00      | \$4,300      | \$43,000    |
| Cohere Transcribe | Medium           | \$3.75      | \$2,500      | \$25,000    |
| Cohere Transcribe | Large            | \$7.50      | \$4,750      | \$47,500    |

You may also want to compare Standard Vault pricing against the operational and capacity costs of
running inference directly in your cloud provider account (for example, AWS), where cloud-provider
credits may apply.

## Generative models

The rates above cover **Embed**, **Rerank**, **Parse**, and **Transcribe**. Embed and Rerank are
available self-serve. Generative models (the Command A family and North) are also available on
Standard Vault, typically through a waitlist. All rates are per instance, per hour.

| Model               | L Hourly rate | XL hourly rate |
| ------------------- | ------------- | -------------- |
| Command A           | \$40.00       | \$48.00        |
| Command A Vision    | \$40.00       | \$48.00        |
| Command A Translate | \$40.00       | \$48.00        |
| Command A Reasoning | \$48.00       | \$57.50        |
| Command A+          | \$17.50       | \$32.50        |
| North Mini Code     | \$7.50        | \$10.50        |

Monthly and annual commitment pricing follows the same Fixed and Flex plans described above;
[contact Cohere](https://cohere.com/contact-sales) for those rates and for access. Generative access
typically requires a waitlist, so check the model dropdown when creating a vault, or contact Cohere.
Bundles and customized models are also available; see
[Supported Models](../../../docs/model-vault/standard/supported-models) for the full list.

For the pricing of the encrypted, confidential-computing product, see
[Model Vault Encrypted Pricing](../../../docs/model-vault/encrypted/pricing).

## Performance Tiers

Each model has a performance tier based on latency requirements and throughput service level
objectives (SLOs) per instance. You can see the tiers listed in the model dropdown selection as a size
letter (e.g., `S`, `M`, `L`). The tiers follow an **instance/hour pricing**, which is then
incorporated into your payment plan. We recommend selecting the model-performance tier combination
that matches the nature of your workflow and the required measures of performance.