> This page is for version v1 API.
> For other versions, use one of these documentation indexes:
> - v2 API (default): https://docs.cohere.com/v2/llms.txt
> - v1 API: https://docs.cohere.com/v1/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.cohere.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server.

# Release Notes

## Announcing Cohere's Embed 5 Models

> Release announcement for Embed 5 Pro and Embed 5 Fast, Cohere's most powerful family of embeddings models for enterprise search and retrieval.

We're pleased to announce the release of [Embed 5](/docs/cohere-embed), Cohere's most powerful embeddings family yet.

Embed 5 delivers frontier retrieval quality on complex enterprise data, with major gains over Embed 4 on visually rich documents, financial filings, parsed PDFs, code, and multilingual retrieval.

## Key features

* **Two model variants available:**
  * `embed-v5.0-pro`: Optimized for the highest retrieval quality, particularly for offline indexing and quality-critical retrieval
  * `embed-v5.0-fast`: Optimized for low latency and high throughput, particularly for interactive search, agent loops, and high-volume query traffic
* **Shared embedding space**: Pro and Fast share an embedding space, so a corpus indexed with one model can be queried with the other. We recommend indexing with Pro and querying with Fast.
* **Multimodal inputs**: Embed text, images, and mixed text-and-image inputs (e.g. PDF pages) in a single vector
* **Multilingual support**: Supports over 100 languages
* **Extended context length**: 128k token context window
* **Flexible storage**: Matryoshka embeddings in the following dimensions: `[256, 512, 768, 1024, 1536, 2048]`, with `float`, `int8`, and `binary` output types

## Availability

Embed 5 is available through the [Embed API](/reference/embed), as well as Microsoft Foundry
([Pro](https://ai.azure.com/catalog/models/Cohere-Embed-V5-Pro),
[Fast](https://ai.azure.com/catalog/models/Cohere-Embed-V5-Fast)) and Amazon SageMaker
([Pro](https://aws.amazon.com/marketplace/pp/prodview-pobc74bcuaohi),
[Fast](https://aws.amazon.com/marketplace/pp/prodview-qffayhd7sgkja)).
For single-tenant deployment, Embed 5 is also available in Cohere's [Model Vault](/docs/model-vault).

For more details, see the [model documentation](/docs/cohere-embed).

## Announcing Cohere's North Small Translate

> This announcement covers the release of North Small Translate, Cohere's open-weights machine translation model.

We're pleased to announce the release of
[North Small Translate](/docs/north-small-translate-1.0), an open-weights mixture-of-experts model purpose-built
for machine translation across more than 50 languages.

North Small Translate is designed to give researchers, developers, and enterprises flexible ways to evaluate and
deploy machine translation while retaining control over their data and infrastructure.

## Key features

* **Purpose-built translation**: Optimized for machine translation across more than 50 languages and locale
  variants.
* **Efficient MoE architecture**: 218 billion total parameters with 25 billion active parameters.
* **Flexible deployment**: Available through the free-tier Chat V2 API and as open weights in W4A16, FP8, and BF16
  for non-commercial use.
* **Private deployment**: Suggested deployment hardware by quantization format:
  * **W4A16**: Two H100s or one B200
  * **FP8**: Four H100s or two B200s
  * **BF16**: Eight H100s or four B200s

## Technical details

* **Model name**: `north-small-translate-1-0`
* **Context length**: 16K
* **License**:
  [Creative Commons Attribution-NonCommercial 4.0](https://creativecommons.org/licenses/by-nc/4.0/)
* **Open-weights formats**: W4A16, FP8, and BF16

## Availability

North Small Translate is available on the free tier through the [Chat V2 API](/reference/chat). Open weights are
available in W4A16, FP8, and BF16 on [Hugging Face](https://huggingface.co/CohereLabs/North-Small-Translate-1.0)
for non-commercial use under the CC BY-NC 4.0 license.

For supported languages, use cases, and an API example, see the
[model documentation](/docs/north-small-translate-1.0).

## Meet Cohere Parse

> This announcement covers the release of Cohere Parse, Cohere's document parsing model for visual understanding and document intelligence workflows.

Today we are releasing Cohere [Parse](/docs/parse).

Parse (model ID: `parse-v5.0`) turns complex documents into clean, structured Markdown ready for downstream AI workflows. The 2.3B-parameter multimodal model extracts text in reading order, tables, lists, forms, images and captions, page boundaries, and visual element locations.

Outputs include Markdown/HTML content, HTML-formatted tables, bounding boxes, and image descriptions — preserving both document structure and layout for easier rendering and processing.

Key specs: 8K context window · \~4.6GB model size · Markdown output

## Availability

Cohere Parse is available through the Parse API, as well as [Microsoft Foundry](https://ai.azure.com/catalog/models/Cohere-parse-v5) and [AWS SageMaker](https://aws.amazon.com/marketplace/pp/prodview-25vdn5x53zqgo).
For single-tenant deployment, Parse is also available in [Model Vault](/docs/model-vault).

For more details, see the [model documentation](/docs/parse).

## Meet Cohere Transcribe Arabic

> This announcement covers the release of Cohere Transcribe Arabic, Cohere's specialist speech recognition model for transcribing Arabic audio.

Today we are releasing [Cohere Transcribe Arabic](/docs/transcribe-arabic).

This open-source speech-to-text model is a fine-tune of Cohere Transcribe using Arabic speech data. It
lets Arabic speakers transcribe their voice with unmatched accuracy and support for regional dialects or
speech patterns.

It is currently the most accurate open-source Arabic ASR model available today and is optimized for
production inference and throughput.

## Technical Details

* **Model Name**: `cohere-transcribe-arabic-07-2026`
* **Size**: 2B
* **Architecture**: conformer-based encoder-decoder
* **Languages supported**: Arabic (all major dialects), English (including English spoken with an
  Arabic accent)
* **License**: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)

## Availability

Cohere Transcribe Arabic is available through the V2 Audio Transcriptions API and as open weights on
[Hugging Face](https://huggingface.co/CohereLabs/cohere-transcribe-arabic-06-2026). For
production use, [Model Vault](/docs/model-vault) deployment is also supported.

For more details, see the [model documentation](/docs/transcribe-arabic).

## Announcing Cohere's North Mini Code

> This announcement covers the release of North Mini Code, Cohere's first open-source agentic coding model.

We're pleased to announce the release of [North Mini Code](/docs/north-mini-code-1.0), Cohere's first
agentic coding model. It is a 30 billion total / 3 billion active parameter Mixture of Experts model
trained specifically for agentic coding, with a small enough active footprint to run on local hardware.

## Technical Details

* **Model Name**: `north-mini-code-1-0`
* **Context Length**: 256K input, 64K output
* **License**: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)

## Availability

North Mini Code is available through the Chat V2 API and as open weights on Hugging Face. For
production use, [Model Vault](/docs/model-vault) deployment is also supported.

For more details, see the [model documentation](/docs/north-mini-code-1.0).

## Announcing Cohere's Command A+

> This announcement covers the release of Command A+, Cohere's last model in the Command A family.

We're pleased to announce the release of [Command A+](/docs/command-a-plus), the last model in the Command A
family of models, combining support for vision inputs, reasoning capabilities, translation capabilities, and
agentic tasks all within the same model. It is also notably our first Mixture of Experts (MoE) model with 25
billion active parameters ands 218 billion total parameters.

## Key Features

* **Agentic Applications**: With notable performance increases in tool use and agentic tasks, Command A+ is
  the strongest agentic model in the Command family.
* **Expanded Multilingual Support**: With 48 languages supported, including all official EU languages, this
  more than doubles the support of languages from our prior models.
* **Efficient & Fast**: With as few as 1 x B200 or 2 x H100s required to deploy the model, and up to 110%
  throughput increase and 30% decrease in latency over Command A Reasoning, the model is designed for
  production-grade deployments.

## Technical Details

* **Model Name**: command-a-plus-05-2026
* **Context Length**: 128K input, 64K output
* **Languages covered**: English, Arabic, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, Spanish, Estonian, Persian, Finnish, Filipino, French, Irish, Hebrew, Hindi, Croatian, Hungarian, Indonesian, Icelandic, Italian, Japanese, Korean, Lithuanian, Latvian, Malay, Maltese, Dutch, Norwegian, Punjabi, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Serbian, Swedish, Tamil, Telugu, Thai, Turkish, Ukrainian, Urdu, Vietnamese, Chinese.
* **License**: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)

## Availability

Command A+ (`command-a-plus-05-2026`) is now available for all Cohere users through our standard API
endpoints. For enterprise customers, [private deployment](/docs/private-deployment-overview) options are
available to ensure maximum security and control over your translation workflows.

For more detailed information about Command A+, including technical specifications and implementation
examples, visit our [model documentation](/docs/command-a-plus).

## Retirement of Embed v2.0 and Aya Expanse / Vision 8B

> Effective April 4, 2026, five models are no longer available on the Cohere API. Migrate to Embed v3/v4 and Command or Aya 32B alternatives.

## Retirement notice

Effective April 4, 2026, the following models are no longer available. Requests using these model IDs will fail.

Retired models:

* `embed-english-v2.0`
* `embed-english-light-v2.0`
* `embed-multilingual-v2.0`
* `c4ai-aya-expanse-8b`
* `c4ai-aya-vision-8b`

We recommend these replacements:

**Embedding tasks**

* `embed-english-v3.0`
* `embed-multilingual-v3.0`
* `embed-v4.0`

**Chat tasks**

* `command-r7b-12-2024`
* `command-a-03-2025`
* `command-a-reasoning-08-2025`

For the full announcement and lifecycle context, see the [Deprecations](/docs/deprecations) page. For questions or
assistance, contact [support@cohere.com](mailto:support@cohere.com).

## Announcing the Cohere Transcribe model

> This announcement covers the release of Cohere Transcribe, Cohere's first transcription model.

We're pleased to announce the release of [Cohere Transcribe](/docs/transcribe), our first transcription model.
Cohere Transcribe specializes in audio-in, text-out, automatic speech recognition (ASR).

## Technical details

* **Model name**: `cohere-transcribe-03-2026`
* **Input**: Audio waveform
* **Output**: Text
* **Languages covered**: English, German, French, Italian, Spanish, Portuguese, Greek, Dutch, Polish,
  Vietnamese, Chinese, Arabic, Japanese, Korean.
* **License**: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)
* **API endpoint**: [Audio Transcriptions API](/reference/create-audio-transcription)

## Getting started

The model is available immediately through Cohere's [Audio Transcriptions API endpoint](/reference/create-audio-transcription).
You can start transcribing audio using the following example query:

**`PYTHON`**

```python PYTHON
import cohere

co = cohere.ClientV2()

response = co.audio.transcriptions.create(
    model="cohere-transcribe-03-2026",
    language="en",
    file=open("./sample.wav", "rb"),
)

print(response)
```

## Availability

You can access Cohere Transcribe via our [API](http://dashboard.cohere.com) for free, low-setup experimentation
subject to rate limits. See the [Different Types of API Keys and Rate Limits](/docs/rate-limits) page for
usage details and integration guidance.

For production deployment without rate limits, provision a dedicated [Model Vault](/docs/model-vault).
This enables low-latency, private cloud inference without having to manage infrastructure. Pricing is
calculated per hour-instance, with discounted plans for longer-term commitments.
[Contact our team](https://cohere.com/contact-sales) to discuss your requirements.

## Cohere's Rerank v4.0 Model is Here!

> Release announcment for Rerank 4.0 - our new state of the art model for ranking.

We're pleased to announce the release of [Rerank 4.0](/docs/rerank) our newest and most performant foundational model for ranking.

## Technical Details

* **Two model variants available:**
  * `rerank-v4.0-pro`: Optimized for state-of-the-art quality and complex use-cases
  * `rerank-v4.0-fast`: Optimized for low latency and high throughput use-cases
* **Multilingual support**: Re-rank both English and non-English documents
* **Semi-structured data support**: Re-rank JSON documents
* **Extended context length**: 32k token context window

## Example Query

**`PYTHON`**

```python PYTHON
import cohere

co = cohere.ClientV2()

query = "What is the capital of the United States?"
docs = [
    "Carson City is the capital city of the American state of Nevada. At the 2010 United States Census, Carson City had a population of 55,274.",
    "The Commonwealth of the Northern Mariana Islands is a group of islands in the Pacific Ocean that are a political division controlled by the United States. Its capital is Saipan.",
    "Charlotte Amalie is the capital and largest city of the United States Virgin Islands. It has about 20,000 people. The city is on the island of Saint Thomas.",
    "Washington, D.C. (also known as simply Washington or D.C., and officially as the District of Columbia) is the capital of the United States. It is a federal district. The President of the USA and many major national government offices are in the territory. This makes it the political center of the United States of America.",
    "Capital punishment has existed in the United States since before the United States was a country. As of 2017, capital punishment is legal in 30 of the 50 states. The federal government (including the United States military) also uses capital punishment.",
]

results = co.rerank(
    model="rerank-v4.0-pro", query=query, documents=docs, top_n=5
)
```

## Announcing Major Command Deprecations

> This announcement covers a series of major deprecations, including of classic Command models, several parameters, and entire endpoints.

As part of our ongoing commitment to delivering advanced AI solutions, we are deprecating the following models, features, and API endpoints:

Deprecated Models:

* `command-r-03-2024`  (and the alias `command-r`)
* `command-r-plus-04-2024`  (and the alias `command-r-plus`)
* `command-light`
* `command`
* `summarize` (Refer to [the migration guide](https://docs.cohere.com/docs/summarizing-text#migration-from-summarize-to-chat-endpoint) for alternatives).

For command model replacements, we recommend you use `command-r-08-2024`, `command-r-plus-08-2024`, or `command-a-03-2025` (which is the strongest-performing model across domains) instead.

Retired Fine-Tuning Capabilities:
All fine-tuning options via dashboard and API for models including `command-light`, `command`, `command-r`, `classify`, and `rerank` are being retired. Previously fine-tuned models will no longer be accessible.

Deprecated Features and API Endpoints:

* `/v1/connectors` (Managed connectors for RAG)
* `/v1/chat` parameters: `connectors`, `search_queries_only`
* `/v1/generate` (Legacy generative endpoint)
* `/v1/summarize` (Legacy summarization endpoint)
* `/v1/classify`
* Slack App integration
* Coral Web UI (chat.cohere.com and coral.cohere.com)

For questions, reach out to [support@cohere.com](mailto:support@cohere.com)

## Announcing Cohere's Command A Translate Model

> This announcement covers the release of Command A Translate, Cohere's most powerful translation model.

We're excited to announce the release of **Command A Translate**, Cohere's first machine translation model. It achieves state-of-the-art performance at producing accurate, fluent translations across 23 languages.

## Key Features

* **23 supported languages**: English, French, Spanish, Italian, German, Portuguese, Japanese, Korean, Chinese, Arabic, Russian, Polish, Turkish, Vietnamese, Dutch, Czech, Indonesian, Ukrainian, Romanian, Greek, Hindi, Hebrew, and Persian
* **111 billion parameters** for superior translation quality
* **16K token context length** (8K input + 8K output) for handling longer texts
* **Optimized for deployment** on 1-2 GPUs (A100s/H100s)
* **Secure deployment options** for sensitive data translation

## Getting Started

The model is available immediately through Cohere's Chat API endpoint. You can start translating text with simple prompts or integrate it programmatically into your applications.

```python
from cohere import ClientV2

co = ClientV2(api_key="<YOUR API KEY>")

response = co.chat(
    model="command-a-translate-08-2025",
    messages=[
        {
            "role": "user",
            "content": "Translate this text to Spanish: Hello, how are you?",
        }
    ],
)
```

## Availability

Command A Translate (`command-a-translate-08-2025`) is now available for all Cohere users through our standard API endpoints. For enterprise customers, [private deployment](https://docs.cohere.com/docs/private-deployment-overview) options are available to ensure maximum security and control over your translation workflows.

For more detailed information about Command A Translate, including technical specifications and implementation examples, visit our [model documentation](/docs/command-a-translate).

## Announcing Cohere's Command A Reasoning Model

> This announcement covers the release of Command A Reasoning, Cohere's first model able to engage in thinking and reasoning.

We’re excited to announce the release of **Command A Reasoning**, a hybrid reasoning model designed to excel at complex agentic tasks, in English and 22 other languages. With 111 billion parameters and a 256K context length, this model brings advanced reasoning capabilities to your applications through the familiar Command API interface.

**Key Features**

* **Tool Use**: Provides the strongest tool use performance out of the Command family of models.
* **Agentic Applications**: Demonstrates proactive problem-solving, autonomously using tools and resources to complete highly complex tasks.
* **Multilingual**: With 23 languages supported, the model solves reasoning and agentic problems in the language your business operates in.

**Technical Specifications**

* **Model Name**: `command-a-reasoning-08-2025`
* **Context Length**: 256K tokens
* **Maximum Output**: 32K tokens
* **API Endpoint**: Chat API

## Getting Started

Integrating Command A Reasoning is straightforward using the Chat API. Here’s a non-streaming example:

**`PYTHON`**

```python PYTHON 
from cohere import ClientV2

co = ClientV2("<YOUR_API_KEY>")

prompt = """
Alice has 3 brothers and she also has 2 sisters. How many sisters does Alice's brother have?
"""

response = co.chat(
    model="command-a-reasoning-08-2025",
    messages=[
        {
            "role": "user",
            "content": prompt,
        }
    ],
)

for content in response.message.content:
    if content.type == "thinking":
        print("Thinking:", content.thinking)

    if content.type == "text":
        print("Response:", content.text)
```

**`PYTHON (Streaming)`**

```python PYTHON (Streaming) 
from cohere import ClientV2

co = ClientV2(api_key="<YOUR_API_KEY>")        

prompt = """
Alice has 3 brothers and she also has 2 sisters. How many sisters does Alice's brother have?
"""

response = co.chat_stream(
    model="command-a-reasoning-08-2025",
    messages=[
        {
            "role": "user",
            "content": prompt,
        }
    ],
)
for event in response:
    if event.type == "content-delta":
        if event.delta.message.content.thinking:
            print(event.delta.message.content.thinking, end="")
        if event.delta.message.content.text:
            print(event.delta.message.content.text, end="")
```

**Customization Options**

You can enable and disable thinking capabilities using the `thinking` parameter, and steer the model's output with a flexible user-controlled thinking budget; for more details on token budgets, advanced configurations, and best practices, refer to our dedicated [Reasoning documentation](/docs/reasoning).

## Announcing Cohere's Command A Vision Model

> This announcement covers the release of Command A Vision, Cohere's first model able to understand and interpret image inputs.

We're excited to announce the release of **Command A Vision**, Cohere's first commercial model capable of understanding and interpreting visual data alongside text. This addition to our Command family brings enterprise-grade vision capabilities to your applications with the same familiar Command API interface.

## Key Features

### Multimodal Capabilities

* **Text + Image Processing**: Combine text prompts with image inputs
* **Enterprise-Focused Use Cases**: Optimized for business applications like document analysis, chart interpretation, and OCR
* **Multiple Languages**: Officially supports English, Portuguese, Italian, French, German, and Spanish

### Technical Specifications

* **Model Name**: `command-a-vision-07-2025`
* **Context Length**: 128K tokens
* **Maximum Output**: 8K tokens
* **Image Support**: Up to 20 images per request (or 20MB total)
* **API Endpoint**: Chat API

## What You Can Do

Command A Vision excels in enterprise use cases including:

* **📊 Chart & Graph Analysis**: Extract insights from complex visualizations
* **📋 Table Understanding**: Parse and interpret data tables within images
* **📄 Document OCR**: Optical character recognition with natural language processing
* **🌐 Image Processing for Multiple Languages**: Handle text in images across multiple languages
* **🔍 Scene Analysis**: Identify and describe objects within images

## 💻 Getting Started

The API structure is identical to our existing Command models, making integration straightforward:

```python
import cohere

co = cohere.Client("your-api-key")

response = co.chat(
    model="command-a-vision-07-2025",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "text",
                    "text": "Analyze this chart and extract the key data points",
                },
                {
                    "type": "image_url",
                    "image_url": {"url": "your-image-url"},
                },
            ],
        }
    ],
)
```

There's much more to be said about working with images, various limitations, and best practices, which you can find in our dedicated [Command A Vision](https://docs.cohere.com/docs/command-a-vision) and [Image Inputs](https://docs.cohere.com/docs/image-inputs) documents.

## Announcing Cutting-Edge Cohere Models on OCI

> This announcement covers the release of Command A, Rerank v3.5, and Embed v3.0 multimodal on Oracle Cloud Infrastructure's platform.

We are thrilled to announce that the Oracle Cloud Infrastructure (OCI) Generative AI service now supports Cohere Command A, Rerank v3.5, Embed v3.0 multimodal. This marks a major advancement in providing OCI's customers with enterprise-ready AI solutions.

Command A 03-2025 is the most performant Command model to date, delivering 150% of the throughput of its predecessor on only two GPUs.

Embed v3.0 is a cutting-edge AI search model enhanced with multimodal capabilities, allowing it to generate embeddings from both text and images.

Rerank 3.5, Cohere's newest AI search foundation model, is engineered to improve the precision of enterprise search and retrieval-augmented generation (RAG) systems across a wide range of data formats (such as lengthy documents, emails, tables, JSON, and code) and in over 100 languages.

Check out [Oracle's announcement](https://blogs.oracle.com/ai-and-datascience/post/cohere-command-a-rerank-oci-gen-ai) and [documentation](https://docs.oracle.com/en-us/iaas/Content/generative-ai/pretrained-models.htm) for more details.

## Announcing Embed Multimodal v4

> Release of Embed Multimodal v4, a performant search model, with Matryoshka embeddings and a 128k context length.

We’re thrilled to announce the release of [Embed 4](https://docs.cohere.com/docs/cohere-embed), the most recent entrant into the Embed family of enterprise-focused [large language models](https://docs.cohere.com/docs/the-cohere-platform#large-language-models-llms) (LLMs).

Embed v4 is Cohere’s most performant search model to date, and supports the following new features:

1. Matryoshka Embeddings in the following dimensions: '\[256, 512, 1024, 1536]'
2. Unified Embeddings produced from mixed modality input (i.e. a single payload of image(s) and text(s))
3. Context length of 128k

Embed v4 achieves state of the art in the following areas:

1. Text-to-text retrieval
2. Text-to-image retrieval
3. Text-to-mixed modality retrieval (from e.g. PDFs)

Embed v4 is available today on the [Cohere Platform](https://docs.cohere.com/docs/the-cohere-platform), [AWS Sagemaker](https://docs.cohere.com/docs/amazon-sagemaker-setup-guide#embeddings), and [Azure AI Foundry](https://docs.cohere.com/docs/cohere-on-microsoft-azure#embeddings). For more information, check out our [dedicated blog post](https://cohere.com/blog/embed-4).

## Announcing Command A

> Release of Command A, a performant model suited for tool use, RAG, agents, and multilingual uses, with 111 billion parameters and a 256k context length.

We're thrilled to announce the release of Command A, the most recent entrant into the Command family of enterprise-focused [large language models](https://docs.cohere.com/docs/the-cohere-platform#large-language-models-llms) (LLMs).

[Command A](https://docs.cohere.com/docs/command-a) is Cohere's most performant model to date, excelling at real world enterprise tasks including tool use, retrieval augmented generation (RAG), agents, and multilingual use cases. With 111B parameters and a context length of 256K, Command A boasts a considerable increase in inference-time efficiency -- 150% higher throughput compared to its predecessor Command R+ 08-2024 -- and only requires two GPUs (A100s / H100s) to run.

Command A is available today on the [Cohere Platform](https://docs.cohere.com/docs/the-cohere-platform), [HuggingFace](https://huggingface.co/CohereForAI/c4ai-command-a-03-2025), or through the SDK with `command-a-03-2025`. For more information, check out our [dedicated blog post](https://cohere.com/blog/command-a/).

## Our Groundbreaking Multimodal Model, Aya Vision, is Here!

> Release announcement for the new multimodal Aya Vision model

Today, Cohere Labs, Cohere’s research arm, is proud to announce
[Aya Vision](https://cohere.com/blog/aya-vision), a state-of-the-art multimodal large language model
excelling across multiple languages and modalities. Aya Vision outperforms the leading open-weight models
in critical benchmarks for language, text, and image capabilities.

## Technical Details

Built as a foundation for multilingual and multimodal communication, this groundbreaking AI model supports
tasks such as image captioning, visual question answering, text generation, and translations from both
texts and images into coherent text.

Refer to [Aya Vision](/docs/aya-multimodal) for more information.

## Cohere Releases Arabic-Optimized Command Model!

> Release announcement for the Command R7B Arabic model

Cohere is thrilled to announce the release of Command R7B Arabic (`c4ai-command-r7b-12-2024`). This is an open weights release of an advanced, 8-billion parameter custom model optimized for the Arabic language (MSA dialect), in addition to English. As with Cohere's other command models, this one comes with context length of 128,000 tokens; it excels at a number of critical enterprise tasks -- instruction following, length control, [retrieval-augmented generation (RAG)](https://docs.cohere.com/docs/retrieval-augmented-generation-rag), minimizing code-switching -- and it demonstrates excellent general purpose knowledge and understanding of the Arabic language and culture.

## Try Command R7B Arabic

If you want to try Command R7B Arabic, it's very easy: you can use it through the [Cohere playground](https://dashboard.cohere.com/playground/chat) or in our dedicated [Hugging Face Space](https://huggingface.co/spaces/CohereForAI/c4ai-command-r-plus).

Alternatively, you can use the model in your own code. To do that, first install the `transformers` library from its source repository:

```bash
pip install 'git+https://github.com/huggingface/transformers.git'
```

Then, use this Python snippet to run a simple text-generation task with the model:

```python
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "CohereForAI/c4ai-command-r7b-12-2024"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

# Format message with the c4ai-command-r7b-12-2024 chat template
messages = [{"role": "user", "content": "مرحبا، كيف حالك؟"}]
input_ids = tokenizer.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
)

gen_tokens = model.generate(
    input_ids,
    max_new_tokens=100,
    do_sample=True,
    temperature=0.3,
)

gen_text = tokenizer.decode(gen_tokens[0])
print(gen_text)
```

## Chat Capabilities

Command R7B Arabic can be operated in two modes, "conversational" and "instruct" mode:

* *Conversational mode* conditions the model on interactive behaviour, meaning it is expected to reply in a conversational fashion, provide introductory statements and follow-up questions, and use Markdown as well as LaTeX where appropriate. This mode is optimized for interactive experiences, such as chatbots, where the model engages in dialogue.
* *Instruct mode* conditions the model to provide concise yet comprehensive responses, and to not use Markdown or LaTeX by default. This mode is designed for non-interactive, task-focused use cases such as extracting information, summarizing text, translation, and categorization.

Note: Command R7B Arabic is delivered without a system preamble by default, though we encourage you to experiment with the conversational and instruct mode preambles. More information can be [found in our docs](https://docs.cohere.com/docs/command-r7b-hf).

## Multilingual RAG Capabilities

Command R7B Arabic has been trained specifically for Arabic and English tasks, such as the *generation* step of Retrieval Augmented Generation (RAG).

Command R7B Arabic's RAG functionality is supported through chat templates in Transformers. Using our RAG chat template, the model takes a conversation (with an optional user-supplied system preamble) and a list of document snippets as input. The resulting output contains a response with in-line citations. Here's what that looks like:

```python
# Define conversation input
conversation = [
    {
        "role": "user",
        "content": "اقترح طبقًا يمزج نكهات من عدة دول عربية",
    }
]

# Define documents for retrieval-based generation
documents = [
    {
        "heading": "المطبخ العربي: أطباقنا التقليدية",
        "body": "يشتهر المطبخ العربي بأطباقه الغنية والنكهات الفريدة. في هذا المقال، سنستكشف ...",
    },
    {
        "heading": "وصفة اليوم: مقلوبة",
        "body": "المقلوبة هي طبق فلسطيني تقليدي، يُحضر من الأرز واللحم أو الدجاج والخضروات. في وصفتنا اليوم ...",
    },
]

# Get the RAG prompt
input_prompt = tokenizer.apply_chat_template(
    conversation=conversation,
    documents=documents,
    tokenize=False,
    add_generation_prompt=True,
    return_tensors="pt",
)
# Tokenize the prompt
input_ids = tokenizer.encode_plus(input_prompt, return_tensors="pt")
```

You can then generate text from this input as normal.

## Notes on Usage

We recommend document snippets be short chunks (around 100-400 words per chunk) instead of long documents. They should also be formatted as key-value pairs, where the keys are short descriptive strings and the values are either text or semi-structured.

You may find that simply including relevant documents directly in a user message works as well as or better than using the `documents` parameter to render the special RAG template (though the template is a strong default for those wanting [citations](https://docs.cohere.com/docs/retrieval-augmented-generation-rag#citation-modes)). We encourage users to experiment with both approaches, and to evaluate which mode works best for their specific use case.

## Cohere via OpenAI SDK Using Compatibility API

> With the Compatibility API, you can use Cohere models via the OpenAI SDK without major refactoring.

Today, we are releasing our Compatibility API, enabling developers to seamlessly use Cohere's models via OpenAI's SDK.

This API enables you to switch your existing OpenAI-based applications to use Cohere's models without major refactoring.

It includes comprehensive support for chat completions, such as function calling and structured outputs, as well as support for text embeddings generation.

Check out [our documentation](https://docs.cohere.com/docs/compatibility-api) on how to get started with the Compatibility API, with examples in Python, TypeScript, and cURL.

## Cohere's Rerank v3.5 Model is on Azure AI Foundry!

> Release announcement for the ability to work with Cohere Rerank v3.5 models in the Azure's AI Foundry.

In December 2024, Cohere released [Rerank v3.5 model](https://docs.cohere.com/changelog/rerank-v3.5). It demonstrates SOTA performance on multilingual retrieval, reasoning, and tasks in domains as varied as finance, eCommerce, hospitality, project management, and email/messaging retrieval.

This model has been available through the Cohere API, but today we’re pleased to announce that it can also be utilized through Microsoft Azure's AI Foundry!

You can find more information about using Cohere’s embedding models on AI Foundry in the [Cohere on the Microsoft Azure Platform](https://docs.cohere.com/docs/cohere-on-microsoft-azure) section.

_Showing the 20 most recent of 65 entries. Append `/llms.txt` to the changelog URL for the complete index._