> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.cohere.com/v2/docs/model-vault/standard/pricing/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server. # Standard Vault Pricing > Standard Vault pricing models (Fixed and Flex) and per-model performance tiers and rates. Standard Vault is billed as a Cohere-managed service. Pricing depends on the models you select and each model's [performance tier](#performance-tiers). Cohere manages the underlying infrastructure and scaling, and customers can choose between two pricing models: | Feature | Fixed | Flex | | ------------ | ------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------- | | Commitment | Monthly or annual | Monthly or annual | | Capacity | Fixed number of instances (no autoscaling) | Minimum baseline instances, plus autoscaling | | Sizing | Determined through a sizing exercise or a production trial (for example, based on expected load) | -- | | Autoscaling | -- | Scales up/down based on request rate and agreed latency SLOs | | Pause/resume | You can pause a model or restart a paused model to save on costs. | You can pause a model or restart a paused model to save on costs. | | Overages | -- | Additional capacity billed per instance-hour | | Max capacity | -- | Maximum instance cap per model | The following table summarizes the available models and their rates. All rates are per instance. | Model | Performance Tier | Hourly rate | Monthly rate | Annual rate | | ----------------- | ---------------- | ----------- | ------------ | ----------- | | Embed 3 | Small | \$4.00 | \$2,500 | \$25,000 | | Embed 4 | Small | \$4.00 | \$2,500 | \$25,000 | | Embed 4 | Medium | \$5.00 | \$3,250 | \$32,500 | | Embed 5 Fast | Small | \$3.00 | \$2,000 | \$20,000 | | Embed 5 Fast | Medium | \$5.00 | \$3,250 | \$32,500 | | Embed 5 Pro | Small | \$3.00 | \$2,000 | \$20,000 | | Embed 5 Pro | Medium | \$5.00 | \$3,250 | \$32,500 | | Rerank 3.5 | Medium | \$5.00 | \$3,250 | \$32,500 | | Rerank 4 Fast | Medium | \$5.00 | \$3,250 | \$32,500 | | Rerank 4 Pro | Medium | \$5.00 | \$3,250 | \$32,500 | | Rerank 4 Pro | Large | \$10.00 | \$6,500 | \$65,000 | | Parse 5 | Medium | \$4.00 | \$2,500 | \$25,000 | | Parse 5 | Large | \$8.00 | \$5,000 | \$50,000 | | Parse 5 | XL | \$7.00 | \$4,300 | \$43,000 | | Cohere Transcribe | Medium | \$3.75 | \$2,500 | \$25,000 | | Cohere Transcribe | Large | \$7.50 | \$4,750 | \$47,500 | You may also want to compare Standard Vault pricing against the operational and capacity costs of running inference directly in your cloud provider account (for example, AWS), where cloud-provider credits may apply. ## Generative models The rates above cover **Embed**, **Rerank**, **Parse**, and **Transcribe**. Embed and Rerank are available self-serve. Generative models (the Command A family and North) are also available on Standard Vault, typically through a waitlist. All rates are per instance, per hour. | Model | L Hourly rate | XL hourly rate | | ------------------- | ------------- | -------------- | | Command A | \$40.00 | \$48.00 | | Command A Vision | \$40.00 | \$48.00 | | Command A Translate | \$40.00 | \$48.00 | | Command A Reasoning | \$48.00 | \$57.50 | | Command A+ | \$17.50 | \$32.50 | | North Mini Code | \$7.50 | \$10.50 | Monthly and annual commitment pricing follows the same Fixed and Flex plans described above; [contact Cohere](https://cohere.com/contact-sales) for those rates and for access. Generative access typically requires a waitlist, so check the model dropdown when creating a vault, or contact Cohere. Bundles and customized models are also available; see [Supported Models](../../../docs/model-vault/standard/supported-models) for the full list. For the pricing of the encrypted, confidential-computing product, see [Model Vault Encrypted Pricing](../../../docs/model-vault/encrypted/pricing). ## Performance Tiers Each model has a performance tier based on latency requirements and throughput service level objectives (SLOs) per instance. You can see the tiers listed in the model dropdown selection as a size letter (e.g., `S`, `M`, `L`). The tiers follow an **instance/hour pricing**, which is then incorporated into your payment plan. We recommend selecting the model-performance tier combination that matches the nature of your workflow and the required measures of performance. > Cohere's API documentation helps developers easily integrate natural language processing and generation into their products.