> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.cohere.com/v1/docs/parse-quickstart/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server. # Document Parsing > A quickstart guide for parsing documents with Cohere's Parse model. Cohere's Parse model converts unstructured enterprise documents (PDFs, images, slides) into structured Markdown output. It extracts text, tables, lists, forms, images, captions, and bounding box coordinates. This quickstart guide shows you how to parse a document image with the Parse endpoint. ### Setup First, install the Cohere Python SDK with the following command. ```bash pip install -U cohere ``` Next, import the library and create a client. #### Cohere Platform **`PYTHON`** ```python PYTHON import cohere co = cohere.ClientV2( "COHERE_API_KEY" ) # Get your free API key here: https://dashboard.cohere.com/api-keys ``` #### Private Deployment **`PYTHON`** ```python PYTHON import cohere co = cohere.ClientV2( api_key="", # Leave this blank base_url="", ) ``` #### SageMaker **`PYTHON`** ```python PYTHON import cohere co = cohere.SagemakerClientV2( aws_region="AWS_REGION", aws_access_key="AWS_ACCESS_KEY_ID", aws_secret_key="AWS_SECRET_ACCESS_KEY", aws_session_token="AWS_SESSION_TOKEN", ) ``` ## Prepare the Document Parse accepts documents as base64-encoded data URIs. Convert your image to a data URI. **`PYTHON`** ```python PYTHON import base64 with open("document.png", "rb") as f: b64 = base64.b64encode(f.read()).decode("utf-8") data_uri = f"data:image/png;base64,{b64}" ``` ### Parse the Document Pass the document to the Parse endpoint. By default, the response contains Markdown output. #### Cohere Platform **`PYTHON`** ```python PYTHON response = co.parse( model="parse-v5.0", document={"type": "image_url", "image_url": data_uri}, ) for page in response.pages: print(page.markdown.content) ``` #### Private Deployment **`PYTHON`** ```python PYTHON response = co.parse( model="parse-v5.0", document={"type": "image_url", "image_url": data_uri}, ) for page in response.pages: print(page.markdown.content) ``` #### SageMaker **`PYTHON`** ```python PYTHON response = co.parse( model="YOUR_ENDPOINT_NAME", document={"type": "image_url", "image_url": data_uri}, ) for page in response.pages: print(page.markdown.content) ``` ### Blocks Output To get structured content blocks, set `output_format` to `"blocks"`. Each block has a `type` (e.g. `text`, `table`) with type-specific fields including bounding boxes for tables. **`PYTHON`** ```python PYTHON response = co.parse( model="parse-v5.0", document={"type": "image_url", "image_url": data_uri}, output_format="blocks", ) for page in response.pages: for block in page.blocks: if block.type == "text": print(block.text.content) elif block.type == "table": print(f"[Table] bbox={block.table.bounding_box}") print(block.table.html) print(block.table.description) print() ``` ## Further Resources * [Parse model documentation](/docs/parse) * [SageMaker notebook](https://github.com/cohere-ai/cohere-developer-experience/blob/main/notebooks/sagemaker/Parse%20Models.ipynb) > Cohere's API documentation helps developers easily integrate natural language processing and generation into their products. ## Docs - [Document Parsing - best practices](https://docs.cohere.com/docs/parse-best-practices.md): Best practices for image format, resolution, and throughput when using the Cohere Parse API.