> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.cohere.com/v2/docs/reasoning/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server. # Reasoning Capabilities > Reasoning models excel at tool use, agentic workflows, and complex problem-solving. This page provides a general overview of Cohere's reasoning capalities. Reasoning models represent an advanced approach to AI that enables more sophisticated problem-solving capabilities. Cohere's reasoning models are *hybrid*, meaning reasoning can be enabled (in which case they generate internal reasoning processes before delivering their final responses) or disabled (in which case they function the way any other LLM would). ## How Reasoning Models Work When a reasoning model processes a request, it first works internally to break the problem down step-by-step. This reasoning process happens in dedicated "thinking" content blocks where the model works through its analysis, planning, and logical steps. Only after completing this internal reasoning does the model produce its final text response, and this allows them to tackle complex tasks with deeper analysis. The key benefit is that reasoning models can handle complex problems—such as leveraging tools and agentic problem solving in the 23 supported languages—by first working through the problem internally before presenting a well-reasoned solution. This approach leads to more accurate and thorough responses, while pushing the boundary for the complexity of problems the model is able to solve. ## Getting Started Models with Reasoning capabilities are accessible via the Chat API. Here's an example: **`PYTHON`** ```python PYTHON from cohere import ClientV2 co = ClientV2(api_key="") prompt = """ Alice has 3 brothers and she also has 2 sisters. How many sisters does Alice's brother have? """ response = co.chat( model="command-a-reasoning-08-2025", messages=[ { "role": "user", "content": prompt, } ], ) for content in response.message.content: if content.type == "thinking": print("Thinking:", content.thinking) if content.type == "text": print("Response:", content.text) ``` **`PYTHON (Streaming)`** ```python PYTHON (Streaming) from cohere import ClientV2 co = ClientV2(api_key="") prompt = """ Alice has 3 brothers and she also has 2 sisters. How many sisters does Alice's brother have? """ response = co.chat_stream( model="command-a-reasoning-08-2025", messages=[ { "role": "user", "content": prompt, } ], ) for event in response: if event.type == "content-delta": if event.delta.message.content.thinking: print(event.delta.message.content.thinking, end="") if event.delta.message.content.text: print(event.delta.message.content.text, end="") ``` **`Curl`** ```bash Curl curl --request POST \ --url https://api.cohere.ai/v2/chat \ --header 'accept: application/json' \ --header 'content-type: application/json' \ --header "Authorization: bearer $CO_API_KEY" \ --data '{ "model": "command-a-reasoning-08-2025", "messages": [ { "role": "user", "content": "Alice has 3 brothers and she also has 2 sisters. How many sisters does Alice'\''s brother have?" } ] }' ``` ### Enabling / Disabling Reasoning Capabilities For reasoning models, `thinking` is enabled by default. To disable it, send the following value to the `"thinking"` parameter: **`PYTHON`** ```python PYTHON thinking={ "type": "disabled" # turns off thinking. It is set to "enabled" by default. } ``` ### Thinking Budgets A thinking token budget can also be specified, to set an upper limit on how many thinking tokens the model can produce. Our recommendation is to use unlimited thinking (i.e. `reasoning = on`). However, if you plan to use thinking budgets, please make sure to leave at least 1K tokens for the response. For example, if you want the model to reason until the maximum limit, we recommend 31K as the token budget. When the budget is exceeded, the model will immediately proceed with the final response. **`PYTHON`** ```python PYTHON thinking = { "token_budget": 500 # limits the model's thinking output to at most 500 tokens } ``` ## Use Cases and Applications Reasoning models excel at tasks that benefit from step-by-step analysis, including: * **Agentic Use Cases**: Taking autonomous actions and interacting with the environment to solve problems. * **Tool Use**: Able to leverage a variety of tools, such as search engines and APIs. * **Multilingual**: Able to reason over multilingual inputs, providing support to user queries in 23 different languages. ## Technical Implementation The reasoning process is controlled through specific parameters that allow developers to: * Enable or disable reasoning capabilities * Set token budgets to control the depth of reasoning * Stream responses to see reasoning and final answers in real-time This architecture makes reasoning models particularly valuable for applications requiring high accuracy, transparency in reasoning, and the ability to handle complex, multi-faceted problems that benefit from systematic analysis. > Cohere's API documentation helps developers easily integrate natural language processing and generation into their products.