> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.cohere.com/v1/docs/model-vault/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server. # Model Vault > Cohere's API documentation helps developers easily integrate natural language processing and generation into their products. ## Docs - [Quickstart](https://docs.cohere.com/docs/model-vault/quickstart.md): Create your first vault from the Model Vault app and make an inference request in a few minutes. - [Model Vault Home Page](https://docs.cohere.com/docs/model-vault/vault-home.md): Find and manage all of your vaults (Standard and Encrypted) from one place on the Model Vault home page. - [Creating a Vault](https://docs.cohere.com/docs/model-vault/creating-a-vault.md): Create a new vault, choose Standard or Encrypted, and select a model, performance tier, and replicas. - [Managing Vaults](https://docs.cohere.com/docs/model-vault/managing-vaults.md): View vault details and edit, pause, resume, or delete models from the Model Vault app. - [Monitoring](https://docs.cohere.com/docs/model-vault/monitoring.md): Monitor latency, queuing, and GPU utilization for any vault with the Grafana dashboard. - [Standard Vault Overview](https://docs.cohere.com/docs/model-vault/standard.md): Standard Vault is Cohere's managed, single-tenant inference environment with dedicated infrastructure and no confidential-computing layer. - [Supported Models](https://docs.cohere.com/docs/model-vault/standard/supported-models.md): Cohere models and GPUs available in a Standard Vault. - [Calling a Standard Vault over the API](https://docs.cohere.com/docs/model-vault/standard/api-access.md): Call a Standard Vault with the Cohere SDK, raw HTTP, or an OpenAI-compatible client by pointing requests at your vault endpoint URL. - [Standard Vault Pricing](https://docs.cohere.com/docs/model-vault/standard/pricing.md): Standard Vault pricing models (Fixed and Flex) and per-model performance tiers and rates. - [Encrypted Vault Overview](https://docs.cohere.com/docs/model-vault/encrypted.md): Encrypted Vaults add confidential computing to Model Vault, so prompts and responses stay protected end to end with verifiable attestation. - [Supported Models](https://docs.cohere.com/docs/model-vault/encrypted/supported-models.md): Which Cohere models are available in Model Vault Encrypted, the supported confidential-computing GPUs, and the isolating architecture. - [Calling an Encrypted Vault over the API](https://docs.cohere.com/docs/model-vault/encrypted/api-usage.md): Call an Encrypted Vault through the Cohere OHTTP proxy that verifies the TEE and encrypts end to end before any data is sent. - [Confidential Computing Primer](https://docs.cohere.com/docs/model-vault/encrypted/confidential-computing.md): A primer on the trusted execution environments and GPU confidential computing that power Model Vault Encrypted. - [Security Model](https://docs.cohere.com/docs/model-vault/encrypted/security-model.md): The trust boundary and threat model for Model Vault Encrypted: who can and cannot access your data. - [Remote Attestation](https://docs.cohere.com/docs/model-vault/encrypted/attestation.md): How remote attestation and the Passport model with Intel Trust Authority prove which code is running inside a Model Vault Encrypted deployment. - [Verifying Your Deployment](https://docs.cohere.com/docs/model-vault/encrypted/verifying-deployment.md): How to verify a Model Vault Encrypted deployment: automatic client-side checks and the attestation details you can inspect in the Model Vault app. - [Encryption & Key Management](https://docs.cohere.com/docs/model-vault/encrypted/encryption-key-management.md): How Model Vault Encrypted protects data in transit, at rest, and in use, and how encryption keys and Zero Data Retention are handled. - [Compliance](https://docs.cohere.com/docs/model-vault/encrypted/compliance.md): How Model Vault Encrypted supports compliance requirements such as GDPR, HIPAA, and SOC 2 through hardware-enforced confidentiality and verifiable attestation. - [Model Vault Encrypted Pricing](https://docs.cohere.com/docs/model-vault/encrypted/pricing.md): Pricing for Model Vault Encrypted, Cohere's confidential-computing inference environment. - [Frequently Asked Questions About Model Vault Encrypted](https://docs.cohere.com/docs/model-vault/encrypted/faq.md): Answers to common questions about Model Vault Encrypted: data privacy, attestation and verification, the trust boundary, keys, compliance, and performance. - [Model Vault with North](https://docs.cohere.com/docs/model-vault/model-vault-with-north.md): Run the North application in your environment and route model inference to your vault endpoints. > **Note:** This page contains both a page directory (above) and the landing page content (below). The page directory is generated for agent use and does not appear on the landing page. > For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.cohere.com/v1/docs/model-vault/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server. # Model Vault Overview > Model Vault is a Cohere-managed, single-tenant environment for deploying and serving Cohere models. Every vault is either Standard or Encrypted. Model Vault is a Cohere-managed inference environment for deploying and serving Cohere models in an isolated, single-tenant setup. It provides dedicated infrastructure with full control over model selection, scaling, and performance monitoring, without you operating the underlying serving stack. Because your infrastructure isn't shared with other tenants, you get the security and isolation of private hosting with the convenience of an API: no noisy neighbors, no rate limits, and predictable performance at scale. You manage all of your vaults in one place from the Model Vault app, where you can track spend, usage hours, and activity across them. Model Vault comes in two types, **[Standard](../docs/model-vault/standard)** and **[Encrypted](../docs/model-vault/encrypted)**. The deployment and management experience is the same; the only difference is the level of data protection (see [how they compare](#how-they-compare) below). vault.cohere.com ![The Model Vault home page showing token spend, vault counts, and usage hours, a usage-over-time chart, and separate rows of Encrypted vault and Standard vault cards.](/_fern-img/2270750af1c4685100974be6ac19b50ffe17462f01834346e440bf0b9b8f22ac.webp) ## Why Model Vault ![Dedicated and single-tenant](/_fern-img/8f270872aaca8ea38dfdfd62371824964443da25b067da6799606342087b55aa.webp) ### Dedicated and single-tenant Your load balancer, serving middleware, inference servers, and GPU accelerators are dedicated to you, so there are no noisy neighbors competing for capacity. ![Fully managed](/_fern-img/62369ae95a92fd9b478910246cd8e85332b95cd16a830cbcdc84ec1d02bbaefd.webp) ### Fully managed Create a vault from the Model Vault app and let Cohere handle maintenance, deployments, updates, and scaling. There's no serving stack for you to operate. ![Models and performance tiers](/_fern-img/01b5fc559f51ff2a8b44b8f15270a773435d8d30c8faeca9f6536de2fdf4393e.webp) ### Models and performance tiers Pick a model and a size tier (S, M, L, XL) to match your latency and throughput needs. ![Elastic capacity](/_fern-img/1002cdc48938ac3d0395867b7dc66580d294bea58200353fe048d2f83e102f51.webp) ### Elastic, unthrottled capacity Set a minimum and maximum replica range (1–25 per model), with autoscaling available. You get dedicated throughput with no rate limits, and only pay for what you use. ![Built for production](/_fern-img/149f62bd848085abcff8424925da190b74602be0fbd177975f2a4bc7ad0e74e2.webp) ### Built for production Real-time monitoring of request rates, latency, token throughput, and GPU utilization helps you tune capacity for production workloads. ![Standalone or with North](/_fern-img/343565220cb0d80b1ce36ff7d88b6b85061ef374eb4d24851ee167af46866f34.webp) ### Standalone or with North Call a vault directly over the API, or use it as the inference backend for [North](../docs/model-vault/model-vault-with-north). ## How it works ### Create a vault In the [Model Vault app](../docs/model-vault/vault-home), [create a vault](../docs/model-vault/creating-a-vault): name it, choose a model and performance tier, and set its replica range. ### Get your endpoint Once the vault is `Ready`, copy its **endpoint URL** and **model name** from the vault's details page (see [Managing Vaults](../docs/model-vault/managing-vaults)). ### Call it with the Cohere SDK Point the SDK's `base_url` at your vault endpoint and send chat, embed, or rerank requests (see [Calling a Vault over the API](../docs/model-vault/standard/api-access)). ## Two types of vault [![Standard Vault](/_fern-img/b168a7cda99fa4df5c5a8125bda839775bb39ab41934978e797e8c2f6c57bcca.webp)](../docs/model-vault/standard) ### Standard Vault A Cohere-managed, single-tenant deployment with data protected in transit and at rest. Best when you want dedicated inference without managing the serving stack. EXPLORE STANDARD VAULT → [![Encrypted Vault](/_fern-img/3ff52944a019b7e0e7e8466c4130a99b988bff3d2bb9f152b0dcec5d6dd63afa.webp)](../docs/model-vault/encrypted) ### Encrypted Vault Everything in a Standard Vault, plus confidential computing: prompts, responses, and everything in between stay protected end to end inside hardware-backed trusted execution environments, with verifiable remote attestation. Best for regulated or highly sensitive workloads. EXPLORE ENCRYPTED VAULT → ## How they compare | | Standard Vault | Encrypted Vault | | -------------------------------------------------- | ---------------------------------------------------------------------- | ------------------------------------------------------------------------ | | Managed, single-tenant deployment | Yes | Yes | | Home page, monitoring, usage & billing | Yes | Yes | | Use standalone or with North | Yes | Yes | | Data protected in transit and at rest | Yes | Yes | | Data protected **in use** (confidential computing) | No | Yes | | Verifiable **remote attestation** | No | Yes | | Compliance support (GDPR, HIPAA, SOC 2) | Supported | Supported, plus verifiable attestation evidence | | Supported models | [Standard vault models](../docs/model-vault/standard/supported-models) | [Encrypted vault models](../docs/model-vault/encrypted/supported-models) | | Pricing | [Standard vault pricing](../docs/model-vault/standard/pricing) | [Encrypted vault pricing](../docs/model-vault/encrypted/pricing) | ## How this documentation is organized Start with the shared sections that cover the day-to-day flow for **every vault**, regardless of type: * **Deploy & manage**: [Home Page](../docs/model-vault/vault-home), [Creating a Vault](../docs/model-vault/creating-a-vault), and [Managing Vaults](../docs/model-vault/managing-vaults). * **Operate & observe**: [Monitoring](../docs/model-vault/monitoring). Then dive into the section for your vault type for what's specific to it, including how to call it over the API: * **[Standard Vault](../docs/model-vault/standard)**: supported models, [calling the API](../docs/model-vault/standard/api-access) (Cohere SDK, raw HTTP, or OpenAI-compatible), and pricing. * **[Encrypted Vault](../docs/model-vault/encrypted)**: supported models, [calling the API](../docs/model-vault/encrypted/api-usage) through the attestation-verifying Cohere OHTTP proxy, the confidential-computing deep dive (confidential computing, security model, remote attestation, key management, compliance), and pricing. And to connect a vault to North: * **[Model Vault with North](../docs/model-vault/model-vault-with-north)**: use a vault as the inference backend for North. ## Get started * New to Model Vault? Start with the [Quickstart](../docs/model-vault/quickstart). * Need confidential computing? See [Encrypted Vaults](../docs/model-vault/encrypted).