Standard Vault Pricing
Standard Vault is billed as a Cohere-managed service. Pricing depends on the models you select and each model’s performance tier. Cohere manages the underlying infrastructure and scaling, and customers can choose between two pricing models:
The following table summarizes the available models and their rates. All rates are per instance.
You may also want to compare Standard Vault pricing against the operational and capacity costs of running inference directly in your cloud provider account (for example, AWS), where cloud-provider credits may apply.
Generative models
The rates above cover Embed and Rerank, which are available self-serve. Generative models (the Command A family and North) are also available on Standard Vault, typically through a waitlist. All rates are per instance, per hour.
Monthly and annual commitment pricing follows the same Fixed and Flex plans described above; contact Cohere for those rates and for access. Generative access typically requires a waitlist, so check the model dropdown when creating a vault, or contact Cohere. Bundles and customized models are also available; see Supported Models for the full list.
For the pricing of the encrypted, confidential-computing product, see Model Vault Encrypted Pricing.
Performance Tiers
Each model has a performance tier based on latency requirements and throughput service level
objectives (SLOs) per instance. You can see the tiers listed in the model dropdown selection as a size
letter (e.g., S, M, L). The tiers follow an instance/hour pricing, which is then
incorporated into your payment plan. We recommend selecting the model-performance tier combination
that matches the nature of your workflow and the required measures of performance.