> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.cohere.com/v1/docs/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server. # Guides and concepts > Cohere's API documentation helps developers easily integrate natural language processing and generation into their products. ## Docs - [Welcome to Cohere](https://docs.cohere.com/docs/welcome.md): Find the right Cohere product: the Platform for direct API access, Model Vault for dedicated deployments, or North for a ready-made agentic AI solution. - [An Overview of The Cohere Platform](https://docs.cohere.com/docs/the-cohere-platform.md): Cohere offers world-class Large Language Models (LLMs) like Command, Rerank, and Embed. These help developers and enterprises build LLM-powered applications. - [Installation](https://docs.cohere.com/docs/get-started-installation.md): A guide for installing the Cohere SDK, supported in 4 different languages – Python, TypeScript, Java, and Go. - [Creating a client](https://docs.cohere.com/docs/create-client.md): A guide for creating Cohere API client using Cohere SDK, supported in 4 different languages – Python, TypeScript, Java, and Go. - [Text Generation](https://docs.cohere.com/docs/text-gen-quickstart.md): A quickstart guide for performing text generation with Cohere's Command models (v1 API). - [Retrieval Augmented Generation (RAG)](https://docs.cohere.com/docs/rag-quickstart.md): A quickstart guide for performing retrieval augmented generation (RAG) with Cohere's Command models (v1 API). - [Tool Use & Agents](https://docs.cohere.com/docs/tool-use-quickstart.md): A quickstart guide for using tool use and building agents with Cohere's Command models (v1 API). - [Semantic Search](https://docs.cohere.com/docs/sem-search-quickstart.md): A quickstart guide for performing text semantic search with Cohere's Embed models (v1 API). - [Reranking](https://docs.cohere.com/docs/reranking-quickstart.md): A quickstart guide for performing reranking with Cohere's Reranking models (v1 API). - [Document Parsing](https://docs.cohere.com/docs/parse-quickstart.md): A quickstart guide for parsing documents with Cohere's Parse model. - [Document Parsing - best practices](https://docs.cohere.com/docs/parse-best-practices.md): Best practices for image format, resolution, and throughput when using the Cohere Parse API. - [An Overview of the Developer Playground](https://docs.cohere.com/docs/playground-overview.md): The Cohere Playground is a powerful visual interface for testing Cohere's generation and embedding language models without coding. - [Frequently Asked Questions About Cohere](https://docs.cohere.com/docs/cohere-faqs.md): Cohere is a powerful platform for using Large Language Models (LLMs). This page covers FAQs related to functionality, pricing, troubleshooting, and more. - [Model Vault Overview](https://docs.cohere.com/docs/model-vault.md): Model Vault is a Cohere-managed, single-tenant environment for deploying and serving Cohere models. Every vault is either Standard or Encrypted. - [Quickstart](https://docs.cohere.com/docs/model-vault/quickstart.md): Create your first vault from the Model Vault app and make an inference request in a few minutes. - [Model Vault Home Page](https://docs.cohere.com/docs/model-vault/vault-home.md): Find and manage all of your vaults (Standard and Encrypted) from one place on the Model Vault home page. - [Creating a Vault](https://docs.cohere.com/docs/model-vault/creating-a-vault.md): Create a new vault, choose Standard or Encrypted, and select a model, performance tier, and replicas. - [Managing Vaults](https://docs.cohere.com/docs/model-vault/managing-vaults.md): View vault details and edit, pause, resume, or delete models from the Model Vault app. - [Monitoring](https://docs.cohere.com/docs/model-vault/monitoring.md): Monitor latency, queuing, and GPU utilization for any vault with the Grafana dashboard. - [Standard Vault Overview](https://docs.cohere.com/docs/model-vault/standard.md): Standard Vault is Cohere's managed, single-tenant inference environment with dedicated infrastructure and no confidential-computing layer. - [Supported Models](https://docs.cohere.com/docs/model-vault/standard/supported-models.md): Cohere models and GPUs available in a Standard Vault. - [Calling a Standard Vault over the API](https://docs.cohere.com/docs/model-vault/standard/api-access.md): Call a Standard Vault with the Cohere SDK, raw HTTP, or an OpenAI-compatible client by pointing requests at your vault endpoint URL. - [Standard Vault Pricing](https://docs.cohere.com/docs/model-vault/standard/pricing.md): Standard Vault pricing models (Fixed and Flex) and per-model performance tiers and rates. - [Encrypted Vault Overview](https://docs.cohere.com/docs/model-vault/encrypted.md): Encrypted Vaults add confidential computing to Model Vault, so prompts and responses stay protected end to end with verifiable attestation. - [Supported Models](https://docs.cohere.com/docs/model-vault/encrypted/supported-models.md): Which Cohere models are available in Model Vault Encrypted, the supported confidential-computing GPUs, and the isolating architecture. - [Calling an Encrypted Vault over the API](https://docs.cohere.com/docs/model-vault/encrypted/api-usage.md): Call an Encrypted Vault through the Cohere OHTTP proxy that verifies the TEE and encrypts end to end before any data is sent. - [Confidential Computing Primer](https://docs.cohere.com/docs/model-vault/encrypted/confidential-computing.md): A primer on the trusted execution environments and GPU confidential computing that power Model Vault Encrypted. - [Security Model](https://docs.cohere.com/docs/model-vault/encrypted/security-model.md): The trust boundary and threat model for Model Vault Encrypted: who can and cannot access your data. - [Remote Attestation](https://docs.cohere.com/docs/model-vault/encrypted/attestation.md): How remote attestation and the Passport model with Intel Trust Authority prove which code is running inside a Model Vault Encrypted deployment. - [Verifying Your Deployment](https://docs.cohere.com/docs/model-vault/encrypted/verifying-deployment.md): How to verify a Model Vault Encrypted deployment: automatic client-side checks and the attestation details you can inspect in the Model Vault app. - [Encryption & Key Management](https://docs.cohere.com/docs/model-vault/encrypted/encryption-key-management.md): How Model Vault Encrypted protects data in transit, at rest, and in use, and how encryption keys and Zero Data Retention are handled. - [Compliance](https://docs.cohere.com/docs/model-vault/encrypted/compliance.md): How Model Vault Encrypted supports compliance requirements such as GDPR, HIPAA, and SOC 2 through hardware-enforced confidentiality and verifiable attestation. - [Model Vault Encrypted Pricing](https://docs.cohere.com/docs/model-vault/encrypted/pricing.md): Pricing for Model Vault Encrypted, Cohere's confidential-computing inference environment. - [Frequently Asked Questions About Model Vault Encrypted](https://docs.cohere.com/docs/model-vault/encrypted/faq.md): Answers to common questions about Model Vault Encrypted: data privacy, attestation and verification, the trust boundary, keys, compliance, and performance. - [Model Vault with North](https://docs.cohere.com/docs/model-vault/model-vault-with-north.md): Run the North application in your environment and route model inference to your vault endpoints. - [An Overview of Cohere's Models](https://docs.cohere.com/docs/models.md): Cohere has a variety of models that cover many different use cases. If you need more customization, you can train a model to tune it to your specific use case. - [Cohere's Command A+ Model](https://docs.cohere.com/docs/command-a-plus.md): Command A+ is a Mixture of Experts (MoE) model with 25B active and 218B total parameters, excelling in agentic, reasoning, vision, and multilingual tasks. - [Command A](https://docs.cohere.com/docs/command-a.md): Command A is a performant mode good at tool use, RAG, agents, and multilingual use cases. It has 111 billion parameters and a 256k context length. - [Cohere's Command A Reasoning Model](https://docs.cohere.com/docs/command-a-reasoning.md): Command A Reasoning excels in tool use, agentic workflows, and complex problem-solving. It has 111 billion parameters and a 256k context length. - [Cohere's Command A Translate Model](https://docs.cohere.com/docs/command-a-translate.md): Command A Translate is a state of the art model performant in 23 languages. It has a context length of 16K tokens and 111B parameters. - [Cohere's Command A Vision Model](https://docs.cohere.com/docs/command-a-vision.md): Command A Vision is a powerful visual language model capable of interacting with image inputs. This document contains information about its capabilities. - [Cohere's Command R7B Model](https://docs.cohere.com/docs/command-r7b.md): Command R7B is the smallest, fastest, and final model in our R family of enterprise-focused large language models. It excels at RAG, tool use, and agents. - [Cohere's Command R+ Model](https://docs.cohere.com/docs/command-r-plus.md): Command R+ is Cohere's optimized for conversational interaction and long-context tasks, best suited for complex RAG workflows and multi-step tool use. - [Cohere's Command R Model](https://docs.cohere.com/docs/command-r.md): Command R is a conversational model that excels in language tasks and supports multiple languages, making it ideal for coding use cases. - [Cohere's Embed Models (Details and Application)](https://docs.cohere.com/docs/cohere-embed.md): Explore Embed models for text classification and embedding generation in English and multiple languages, with details on dimensions and endpoints. - [North Small Translate](https://docs.cohere.com/docs/north-small-translate-1.0.md): North Small Translate is a 218B total / 25B active parameter MoE model purpose-built for machine translation across more than 50 languages. - [North Mini Code](https://docs.cohere.com/docs/north-mini-code-1.0.md): North Mini Code is a 30B total / 3B active parameter MoE model trained for agentic coding, released under Apache 2.0 and suitable for local deployment. - [Cohere's Rerank Model (Details and Application)](https://docs.cohere.com/docs/rerank.md): This page describes how Cohere's Rerank models work and how to use them. - [Aya Family of Models](https://docs.cohere.com/docs/aya.md): Understand Cohere Labs groundbreaking multilingual Aya models, which aim to bring many more languages into generative AI. - [Aya Vision](https://docs.cohere.com/docs/aya-vision.md): Understand Cohere Labs groundbreaking multilingual model Aya Vision, a state-of-the-art multimodal language model excelling at multiple tasks. - [Aya Expanse](https://docs.cohere.com/docs/aya-expanse.md): Understand Cohere Labs highly performant multilingual Aya models, which aim to bring many more languages into generative AI. - [Tiny Aya](https://docs.cohere.com/docs/tiny-aya.md): Tiny Aya is a compact yet powerful 3.35B-parameter multilingual model supporting 70 languages, designed for efficient and practical multilingual AI deployment. - [Introduction to Text Generation at Cohere](https://docs.cohere.com/docs/introduction-to-text-generation-at-cohere.md): This page describes how a large language model generates textual output. - [Using the Cohere Chat API for Text Generation](https://docs.cohere.com/docs/chat-api.md): How to use the Chat API endpoint with Cohere LLMs to generate text responses in a conversational interface. - [A Guide to Streaming Responses](https://docs.cohere.com/docs/streaming.md): The document explains how the Chat API can stream events like text generation in real-time. - [How do Structured Outputs Work?](https://docs.cohere.com/v1/docs/structured-outputs.md): This page describes how to get Cohere models to create outputs in a certain format, such as JSON, using parameters such as `response_format`. - [How to Get Predictable Outputs with Cohere Models](https://docs.cohere.com/docs/predictable-outputs.md): Strategies for decoding text, and the parameters that impact the randomness and predictability of a language model's output. - [Advanced Generation Parameters](https://docs.cohere.com/docs/advanced-generation-hyperparameters.md): This page describes advanced parameters for controlling generation. - [Retrieval Augmented Generation (RAG)](https://docs.cohere.com/v1/docs/retrieval-augmented-generation-rag.md): Generate text with external data and inline citations using Retrieval Augmented Generation and Cohere's Chat API. - [An Overview of Tool Use with Cohere](https://docs.cohere.com/docs/tools.md): Understand single-step and multi-step tool use, and learn when to use each in your workflows. - [Multi-step Tool Use (Agents)](https://docs.cohere.com/v1/docs/multi-step-tool-use.md): "Cohere's tool use feature enhances AI capabilities by connecting external tools for dynamic, adaptable, and sequential actions." - [Implementing a Multi-Step Agent with Langchain](https://docs.cohere.com/v1/docs/implementing-a-multi-step-agent-with-langchain.md): This page describes how to building a powerful, flexible AI agent with Cohere and LangChain. (V1) - [How Does Single-Step Tool Use Work?](https://docs.cohere.com/v1/docs/tool-use.md): Enable your large language models to connect with external tools for more advanced and dynamic interactions (V1). - [What Parameter Types are Available in Tool Use?](https://docs.cohere.com/v1/docs/parameter-types-in-tool-use.md): This page describes Cohere's tool use parameters and how to work with them. - [A Guide to Tokens and Tokenizers](https://docs.cohere.com/docs/tokens-and-tokenizers.md): This document describes how to use the tokenize and detokenize API endpoints. - [Migrating from the Generate API to the Chat API](https://docs.cohere.com/v1/docs/migrating-from-cogenerate-to-cochat.md): Learn about the transition from Generate to Chat for improved generative capabilities with Cohere. - [Summarizing Text with the Chat Endpoint](https://docs.cohere.com/docs/summarizing-text.md): Learn how to perform text summarization using Cohere's Chat endpoint with features like length control and RAG. - [Introduction to Embeddings at Cohere](https://docs.cohere.com/docs/embeddings.md): Embeddings transform text into numerical data, enabling language-agnostic similarity searches and efficient storage with compression. - [Semantic Search with Embeddings](https://docs.cohere.com/docs/semantic-search-embed.md): Examples on how to use the Embed endpoint to perform semantic search (API v1). - [Unlocking the Power of Multimodal Embeddings](https://docs.cohere.com/docs/multimodal-embeddings.md): Multimodal embeddings convert text and images into embeddings for search and classification. - [Batch Embedding Jobs with the Embed API](https://docs.cohere.com/docs/embed-jobs-api.md): Learn how to use the Embed Jobs API to handle large text data efficiently with a focus on creating datasets and running embed jobs. - [An Overview of Cohere's Rerank Model](https://docs.cohere.com/docs/rerank-overview.md): This page describes how Cohere's Rerank models work. - [Best Practices for using Rerank](https://docs.cohere.com/docs/reranking-best-practices.md): Tips for optimal endpoint performance, including constraints on the number of documents, tokens per document, and tokens per query. - [Different Types of API Keys and Rate Limits](https://docs.cohere.com/docs/rate-limits.md): This page describes Cohere API rate limits for production and evaluation keys. - [Going Live with a Cohere Model](https://docs.cohere.com/docs/going-live.md): Learn to upgrade from a Trial to a Production key; understand the limitations and benefits of each and go live with Cohere. - [Deprecations](https://docs.cohere.com/docs/deprecations.md): Learn about Cohere's deprecation policies and recommended replacements - [How Does Cohere's Pricing Work?](https://docs.cohere.com/docs/how-does-cohere-pricing-work.md): This page details Cohere's pricing model. Our models can be accessed directly through our API, allowing for the creation of scalable production workloads. - [Integrating Embedding Models with Other Tools](https://docs.cohere.com/docs/integrations.md): Learn how to integrate Cohere embeddings with open-source vector search engines for enhanced applications. - [Elasticsearch and Cohere (Integration Guide)](https://docs.cohere.com/docs/elasticsearch-and-cohere.md): Learn how to create a semantic search pipeline with Elasticsearch and Cohere's generative AI capabilities. - [MongoDB and Cohere (Integration Guide)](https://docs.cohere.com/docs/mongodb-and-cohere.md): Build semantic search and RAG systems using Cohere and MongoDB Atlas Vector Search. - [Redis and Cohere (Integration Guide)](https://docs.cohere.com/docs/redis-and-cohere.md): Learn how to integrate Cohere with Redis for similarity searches on text data with this step-by-step guide. - [Haystack and Cohere (Integration Guide)](https://docs.cohere.com/docs/haystack-and-cohere.md): Build custom LLM applications with Haystack, now integrated with Cohere for embedding, generation, chat, and retrieval. - [Pinecone and Cohere (Integration Guide)](https://docs.cohere.com/docs/pinecone-and-cohere.md): This page describes how to integrate Cohere with the Pinecone vector database. - [Weaviate and Cohere (Integration Guide)](https://docs.cohere.com/docs/weaviate-and-cohere.md): This page describes how to integrate Cohere with the Weaviate database. - [Open Search and Cohere (Integration Guide)](https://docs.cohere.com/docs/opensearch-and-cohere.md): Unlock the power of search and analytics with OpenSearch, enhanced by ML connectors like Cohere and Amazon Bedrock. - [Vespa and Cohere (Integration Guide)](https://docs.cohere.com/docs/vespa-and-cohere.md): This page describes how to integrate Cohere with the Vespa database. - [Qdrant and Cohere (Integration Guide)](https://docs.cohere.com/docs/qdrant-and-cohere.md): This page describes how to integrate Cohere with the Qdrant vector database. - [Milvus and Cohere (Integration Guide)](https://docs.cohere.com/docs/milvus-and-cohere.md): This page describes integrating Cohere with the Milvus vector database. - [Zilliz and Cohere (Integration Guide)](https://docs.cohere.com/docs/zilliz-and-cohere.md): This page describes how to integrate Cohere with the Zilliz database. - [Chroma and Cohere (Integration Guide)](https://docs.cohere.com/docs/chroma-and-cohere.md): This page describes how to integrate Cohere and Chroma. - [Cohere and LangChain (Integration Guide)](https://docs.cohere.com/docs/cohere-and-langchain.md): Integrate Cohere with LangChain for advanced chat features, RAG, embeddings, and reranking; this guide includes code examples for each feature. - [Cohere Chat on LangChain (Integration Guide)](https://docs.cohere.com/docs/chat-on-langchain.md): Integrate Cohere with LangChain to build applications using Cohere's models and LangChain tools. - [Cohere Embed on LangChain (Integration Guide)](https://docs.cohere.com/docs/embed-on-langchain.md): This page describes how to work with Cohere's embeddings models and LangChain. - [Cohere Rerank on LangChain (Integration Guide)](https://docs.cohere.com/docs/rerank-on-langchain.md): This page describes how to integrate Cohere's ReRank models with LangChain. - [Cohere Tools on LangChain (Integration Guide)](https://docs.cohere.com/docs/tools-on-langchain.md): Explore code examples for multi-step and single-step tool usage in chatbots, harnessing internet search and vector storage. - [LlamaIndex and Cohere's Models](https://docs.cohere.com/docs/llamaindex.md): Learn how to use Cohere and LlamaIndex together to generate responses based on data. - [Deployment Options - Overview](https://docs.cohere.com/docs/deployment-options-overview.md): This page provides an overview of the available options for deploying Cohere's models. - [Cohere SDK Cloud Platform Compatibility](https://docs.cohere.com/docs/cohere-works-everywhere.md): This page describes various places you can use Cohere's SDK. - [Private Deployment Overview](https://docs.cohere.com/docs/private-deployment-overview.md): This page provides an overview of private deployments of Cohere's models. - [Private Deployment – Setting Up](https://docs.cohere.com/docs/private-deployment-setup.md): This page describes the setup required for private deployments of Cohere's models. - [Single Container on Private Clouds](https://docs.cohere.com/docs/single-container-on-private-clouds.md): Learn how to pull and test Cohere's container images using a license with Docker and Kubernetes. - [AWS Private Deployment Guide (EC2 and EKS)](https://docs.cohere.com/docs/aws-private-deployment.md): Deploying Cohere models in AWS via EC2 or EKS for enhanced security, compliance, and control. - [Private Deployment Usage](https://docs.cohere.com/docs/private-deployment-usage.md): This page describes how to use Cohere's SDK to access privately deployed Cohere models. - [Cohere on Amazon Web Services (AWS)](https://docs.cohere.com/docs/cohere-on-aws.md): Access Cohere's language models on AWS with customization options through Amazon SageMaker and Amazon Bedrock. - [Cohere Models on Amazon Bedrock](https://docs.cohere.com/docs/amazon-bedrock.md): This document provides a guide for using Cohere's models on Amazon Bedrock. - [An Amazon SageMaker Setup Guide](https://docs.cohere.com/docs/amazon-sagemaker-setup-guide.md): This document will guide you through enabling development teams to access Cohere’s offerings on Amazon SageMaker. - [Deploy Finetuned Command Models from AWS Marketplace](https://docs.cohere.com/docs/bring-your-finetuned-models-to-sagemaker.md): This document provides a guide for bringing your own finetuned models to Amazon SageMaker. - [Cohere on the Microsoft Azure Platform](https://docs.cohere.com/docs/cohere-on-microsoft-azure.md): This page describes how to work with Cohere models on Microsoft Azure. - [Cohere on Oracle Cloud Infrastructure (OCI)](https://docs.cohere.com/docs/oracle-cloud-infrastructure-oci) - [Cohere Cookbooks: AI Agents, RAG, Search, and More](https://docs.cohere.com/docs/cookbooks.md): Get started with Cohere's cookbooks to build agents, QA bots, perform searches, and more, all organized by category. - [Welcome to LLM University!](https://docs.cohere.com/docs/llmu-2.md): LLM University (LLMU) offers in-depth, practical NLP and LLM training. Ideal for all skill levels. Learn, build, and deploy Language AI with Cohere. - [Build an Onboarding Assistant with Cohere!](https://docs.cohere.com/docs/build-things-with-cohere.md): This page describes how to build an onboarding assistant with Cohere's large language models. - [Cohere Text Generation Tutorial](https://docs.cohere.com/docs/text-generation-tutorial.md): This page walks through how Cohere's generation models work and how to use them. - [Building a Chatbot with Cohere](https://docs.cohere.com/docs/building-a-chatbot-with-cohere.md): This page describes building a generative-AI powered chatbot with Cohere. - [Semantic Search with Cohere Models](https://docs.cohere.com/docs/semantic-search-with-cohere.md): This is a tutorial describing how to leverage Cohere's models for semantic search. - [Master Reranking with Cohere Models](https://docs.cohere.com/docs/reranking-with-cohere.md): This page contains a tutorial on using Cohere's ReRank models. - [Building RAG models with Cohere](https://docs.cohere.com/docs/rag-with-cohere.md): This page walks through building a retrieval-augmented generation model with Cohere. - [Building a Generative AI Agent with Cohere](https://docs.cohere.com/docs/building-an-agent-with-cohere.md): This page describes building a generative-AI powered agent with Cohere. - [Usage Policy](https://docs.cohere.com/docs/usage-policy.md): Developers must outline and get approval for their use case to access the Cohere API, understanding the models and limitations. They should refer to model cards for detailed information and document potential harms of their application. Certain use cases, such as violence, hate speech, fraud, and privacy violations, are strictly prohibited. - [Command R and Command R+ Model Card](https://docs.cohere.com/docs/responsible-use.md): This doc provides guidelines for using Cohere generation models ethically and constructively. - [Cohere Web Crawlers](https://docs.cohere.com/docs/cohere-web-crawlers.md): Cohere's policy on web crawlers and robots.txt for generative AI training. - [Cohere Labs Acceptable Use Policy](https://docs.cohere.com/docs/cohere-labs-acceptable-use-policy.md): "Promoting safe and ethical use of generative AI with guidelines to prevent misuse and abuse." - [How to Start with the Cohere Toolkit](https://docs.cohere.com/docs/cohere-toolkit.md): Build and deploy RAG applications quickly with the Cohere Toolkit, which offers pre-built front-end and back-end components. - [The Cohere Datasets API (and How to Use It)](https://docs.cohere.com/docs/datasets.md): Learn about the Dataset API, including its file size limits, data retention, creation, validation, metadata, and more, with provided code snippets. - [Help Us Improve The Cohere Docs](https://docs.cohere.com/docs/contribute.md): Contribute to our docs content, stored in the cohere-developer-experience repo; we welcome your pull requests!