> This page is for version v1 API.
> For other versions, use one of these documentation indexes:
> - v2 API (default): https://docs.cohere.com/v2/llms.txt
> - v1 API: https://docs.cohere.com/v1/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.cohere.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server.

# Document Parsing

> A quickstart guide for parsing documents with Cohere's Parse model.

Cohere's Parse model converts unstructured enterprise documents (PDFs, images, slides) into structured Markdown output. It extracts text, tables, lists, forms, images, captions, and bounding box coordinates.

This quickstart guide shows you how to parse a document image with the Parse endpoint.

### Setup

First, install the Cohere Python SDK with the following command.

```bash
pip install -U cohere
```

Next, import the library and create a client.

#### Cohere Platform

**`PYTHON`**

```python PYTHON
import cohere

co = cohere.ClientV2(
    "COHERE_API_KEY"
)  # Get your free API key here: https://dashboard.cohere.com/api-keys
```

#### Private Deployment

**`PYTHON`**

```python PYTHON
import cohere

co = cohere.ClientV2(
    api_key="",  # Leave this blank
    base_url="<YOUR_DEPLOYMENT_URL>",
)
```

#### SageMaker

**`PYTHON`**

```python PYTHON
import cohere

co = cohere.SagemakerClientV2(
    aws_region="AWS_REGION",
    aws_access_key="AWS_ACCESS_KEY_ID",
    aws_secret_key="AWS_SECRET_ACCESS_KEY",
    aws_session_token="AWS_SESSION_TOKEN",
)
```

## Prepare the Document

Parse accepts documents as base64-encoded data URIs. Convert your image to a data URI.

**`PYTHON`**

```python PYTHON
import base64

with open("document.png", "rb") as f:
    b64 = base64.b64encode(f.read()).decode("utf-8")

data_uri = f"data:image/png;base64,{b64}"
```

### Parse the Document

Pass the document to the Parse endpoint. By default, the response contains Markdown output.

#### Cohere Platform

**`PYTHON`**

```python PYTHON
response = co.parse(
    model="parse-v5.0",
    document={"type": "image_url", "image_url": data_uri},
)

for page in response.pages:
    print(page.markdown.content)
```

#### Private Deployment

**`PYTHON`**

```python PYTHON
response = co.parse(
    model="parse-v5.0",
    document={"type": "image_url", "image_url": data_uri},
)

for page in response.pages:
    print(page.markdown.content)
```

#### SageMaker

**`PYTHON`**

```python PYTHON
response = co.parse(
    model="YOUR_ENDPOINT_NAME",
    document={"type": "image_url", "image_url": data_uri},
)

for page in response.pages:
    print(page.markdown.content)
```

### Blocks Output

To get structured content blocks, set `output_format` to `"blocks"`. Each block has a `type` (e.g. `text`, `table`) with type-specific fields including bounding boxes for tables.

**`PYTHON`**

```python PYTHON
response = co.parse(
    model="parse-v5.0",
    document={"type": "image_url", "image_url": data_uri},
    output_format="blocks",
)

for page in response.pages:
    for block in page.blocks:
        if block.type == "text":
            print(block.text.content)
        elif block.type == "table":
            print(f"[Table] bbox={block.table.bounding_box}")
            print(block.table.html)
            print(block.table.description)
        print()
```

## Further Resources

* [Parse model documentation](/docs/parse)
* [SageMaker notebook](https://github.com/cohere-ai/cohere-developer-experience/blob/main/notebooks/sagemaker/Parse%20Models.ipynb)