> This page is for version v2 API (default).
> For other versions, use one of these documentation indexes:
> - v2 API (default): https://docs.cohere.com/v2/llms.txt
> - v1 API: https://docs.cohere.com/v1/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.cohere.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server.

# Cohere Transcribe Arabic

> This page describes how the Cohere Transcribe Arabic model works and how to use it.

## About this model

Cohere Transcribe Arabic is a finetuned version of Cohere Transcribe that is optimized for Arabic audio inputs. Like Cohere Transcribe, it is a 2B parameters dedicated audio-in, text-out,
automatic speech recognition (ASR) model available open source. The model is designed to accurately capture the diversity of dialects, accents, and acoustic conditions among Arabic speakers during transcription.

## Model details

* **Input**: Audio waveform
* **Output**: Text
* **Model name**: `cohere-transcribe-arabic-07-2026`
* **Languages covered**: Arabic, English
* **Multidialectal support**: Yes
* **Code-switching support**: Yes
* **Maximum file size**: 25MB
* **License**: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0)

## Availability

You can access Cohere Transcribe Arabic via our [API](https://dashboard.cohere.com) for free, low-setup
experimentation subject to [rate limits](rate-limits).

For production deployment without rate limits, provision a dedicated [Model Vault](model-vault).
This enables low-latency, private cloud inference without having to manage infrastructure.
Pricing is calculated per hour-instance, with discounted plans for longer-term commitments.
[Contact our team](https://cohere.com/contact-sales) to discuss your requirements.

## Strengths

Cohere Transcribe Arabic achieves state-of-the-art transcription accuracy for Arabic speech. These results are robust to diverse or variable audio inputs, including: bilingual dialogues (Arabic-English); regional dialects and phrasing, and enterprise-specific vocabulary.
Like its parent model, Cohere Transcribe Arabic has been optimized for high-throughput, production inference, and is at the frontier for far-field (where the speaker is not close to the receiver) transcription tasks.

## Key Limitations

* **Timestamps**: The model does not output timestamps alongside transcripts.
* **Speaker diarization**: The model does not automatically identify individual speakers in a multispeaker audio file.

## Model architecture

Cohere Transcribe is built on a speech-optimized Transformer variant: a Conformer.
Input audio waveforms are converted into a Mel spectrogram and then processed by a Conformer encoder
that holds the majority of the model’s parameters.
The encoder’s representations are then passed to a lightweight Transformer decoder that generates text tokens.
Cohere Transcribe is trained using standard supervised cross-entropy.

## Further Resources

* [Audio Transcriptions quickstart](audio-transcription-quickstart)
* [Audio Transcriptions API reference documentation](../reference/create-audio-transcription)
* [Jupyter notebook](https://github.com/cohere-ai/cohere-developer-experience/blob/main/notebooks/guides/audio/audio_transcription.ipynb)