> This page is for version v2 API (default).
> For other versions, use one of these documentation indexes:
> - v2 API (default): https://docs.cohere.com/v2/llms.txt
> - v1 API: https://docs.cohere.com/v1/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.cohere.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server.

# Generate with Streaming

POST https://api.cohere.com/v1/generate
Content-Type: application/json

This API is marked as "Legacy" and is no longer maintained. Follow the [migration guide](https://docs.cohere.com/docs/migrating-from-cogenerate-to-cochat) to start using the Chat with Streaming API.

Generates realistic text conditioned on a given input.

Reference: https://docs.cohere.com/reference/generate-stream-v1

## Authentication

- `Authorization` header (bearer token, required) — Bearer authentication of the form `Bearer <token>`, where token is your auth token.

## Request

### Headers

- `X-Client-Name` (string, optional) — The name of the project that is making the request.

### Body (application/json)

This endpoint expects an object.

- `prompt` (string, required) — The input text that serves as the starting point for generating the response. Note: The prompt will be pre-processed and modified before reaching the model.
- `stream` (true, required) — When `true`, the response will be a JSON stream of events. Streaming is beneficial for user interfaces that render the contents of the response piece by piece, as it gets generated. The final event will contain the complete response, and will contain an `is_finished` field set to `true`. The event will also contain a `finish_reason`, which can be one of the following: - `COMPLETE` - the model sent back a finished reply - `MAX_TOKENS` - the reply was cut off because the model reached the maximum number of tokens for its context length - `ERROR` - something went wrong when generating the reply - `ERROR_TOXIC` - the model generated a reply that was deemed toxic
- `model` (string, optional) — The identifier of the model to generate with. Currently available models are `command` (default), `command-nightly` (experimental), `command-light`, and `command-light-nightly` (experimental). Smaller, "light" models are faster, while larger models will perform better. [Custom models](https://docs.cohere.com/docs/training-custom-models) can also be supplied with their full ID.
- `num_generations` (integer, optional) — The maximum number of generations that will be returned. Defaults to `1`, min value of `1`, max value of `5`.
- `max_tokens` (integer, optional) — The maximum number of tokens the model will generate as part of the response. Note: Setting a low value may result in incomplete generations. This parameter is off by default, and if it's not specified, the model will continue generating until it emits an EOS completion token. See [BPE Tokens](/bpe-tokens-wiki) for more details. Can only be set to `0` if `return_likelihoods` is set to `ALL` to get the likelihood of the prompt.
- `truncate` (enum, optional, default: END) — One of `NONE|START|END` to specify how the API will handle inputs longer than the maximum token length. Passing `START` will discard the start of the input. `END` will discard the end of the input. In both cases, input is discarded until the remaining input is exactly the maximum input token length for the model. If `NONE` is selected, when the input exceeds the maximum input token length an error will be returned.
  - Allowed values: `NONE`, `START`, `END`
- `temperature` (double, optional) — A non-negative float that tunes the degree of randomness in generation. Lower temperatures mean less random generations. See [Temperature](/temperature-wiki) for more details. Defaults to `0.75`, min value of `0.0`, max value of `5.0`.
- `seed` (integer, optional) — If specified, the backend will make a best effort to sample tokens deterministically, such that repeated requests with the same seed and parameters should return the same result. However, determinism cannot be totally guaranteed. Compatible Deployments: Cohere Platform, Azure, AWS Sagemaker/Bedrock, Private Deployments
- `preset` (string, optional) — Identifier of a custom preset. A preset is a combination of parameters, such as prompt, temperature etc. You can create presets in the [playground](https://dashboard.cohere.com/playground/generate). When a preset is specified, the `prompt` parameter becomes optional, and any included parameters will override the preset's parameters.
- `end_sequences` (list of string, optional) — The generated text will be cut at the beginning of the earliest occurrence of an end sequence. The sequence will be excluded from the text.
- `stop_sequences` (list of string, optional) — The generated text will be cut at the end of the earliest occurrence of a stop sequence. The sequence will be included the text.
- `k` (integer, optional) — Ensures only the top `k` most likely tokens are considered for generation at each step. Defaults to `0`, min value of `0`, max value of `500`.
- `p` (double, optional) — Ensures that only the most likely tokens, with total probability mass of `p`, are considered for generation at each step. If both `k` and `p` are enabled, `p` acts after `k`. Defaults to `0.75`. min value of `0.01`, max value of `0.99`.
- `frequency_penalty` (double, optional) — Used to reduce repetitiveness of generated tokens. The higher the value, the stronger a penalty is applied to previously present tokens, proportional to how many times they have already appeared in the prompt or prior generation. Using `frequency_penalty` in combination with `presence_penalty` is not supported on newer models.
- `presence_penalty` (double, optional) — Defaults to `0.0`, min value of `0.0`, max value of `1.0`. Can be used to reduce repetitiveness of generated tokens. Similar to `frequency_penalty`, except that this penalty is applied equally to all tokens that have already appeared, regardless of their exact frequencies. Using `frequency_penalty` in combination with `presence_penalty` is not supported on newer models.
- `return_likelihoods` (enum, optional, default: NONE) — One of `GENERATION|NONE` to specify how and if the token likelihoods are returned with the response. Defaults to `NONE`. If `GENERATION` is selected, the token likelihoods will only be provided for generated text. WARNING: `ALL` is deprecated, and will be removed in a future release.
  - Allowed values: `GENERATION`, `ALL`, `NONE`
- `raw_prompting` (boolean, optional) — When enabled, the user's prompt will be sent to the model without any pre-processing.

## Response

### 200

This API is marked as "Legacy" and is no longer maintained. Follow the [migration guide](https://docs.cohere.com/docs/migrating-from-cogenerate-to-cochat) to start using the Chat with Streaming API. Generates realistic text conditioned on a given input.

- Streaming response of `generate_Response_stream_streaming`.
- `event_type`: `text-generation` (text-generation)
  - `is_finished` (boolean, required)
  - `text` (string, required) — A segment of text of the generation.
  - `index` (integer, optional) — Refers to the nth generation. Only present when `num_generations` is greater than zero, and only when text responses are being streamed.
- `event_type`: `stream-end` (stream-end)
  - `is_finished` (boolean, required)
  - `response` (GenerateStreamEndResponse, required)
  - `finish_reason` (enum, optional)
    - Allowed values: `COMPLETE`, `STOP_SEQUENCE`, `ERROR`, `ERROR_TOXIC`, `ERROR_LIMIT`, `USER_CANCEL`, `MAX_TOKENS`, `TIMEOUT`
- `event_type`: `stream-error` (stream-error)
  - `err` (string, required) — Error message
  - `finish_reason` (enum, required)
    - Allowed values: `COMPLETE`, `STOP_SEQUENCE`, `ERROR`, `ERROR_TOXIC`, `ERROR_LIMIT`, `USER_CANCEL`, `MAX_TOKENS`, `TIMEOUT`
  - `is_finished` (boolean, required)
  - `index` (integer, optional) — Refers to the nth generation. Only present when `num_generations` is greater than zero.

## Errors

### 400 Bad Request Error

This error is returned when the request is not well formed. This could be because: - JSON is invalid - The request is missing required fields - The request contains an invalid combination of fields

- `message` (string, optional)
- `id` (string, optional)

### 401 Unauthorized Error

This error indicates that the operation attempted to be performed is not allowed. This could be because: - The api token is invalid - The user does not have the necessary permissions

- `message` (string, optional)
- `id` (string, optional)

### 403 Forbidden Error

This error indicates that the operation attempted to be performed is not allowed. This could be because: - The api token is invalid - The user does not have the necessary permissions

- `message` (string, optional)
- `id` (string, optional)

### 404 Not Found Error

This error is returned when a resource is not found. This could be because: - The endpoint does not exist - The resource does not exist eg model id, dataset id

- `message` (string, optional)
- `id` (string, optional)

### 422 Unprocessable Entity Error

This error is returned when the request is not well formed. This could be because: - JSON is invalid - The request is missing required fields - The request contains an invalid combination of fields

- `message` (string, optional)
- `id` (string, optional)

### 429 Too Many Requests Error

Too many requests

- `message` (string, optional)
- `id` (string, optional)

### 498 Invalid Token Error

This error is returned when a request or response contains a deny-listed token.

- `message` (string, optional)
- `id` (string, optional)

### 499 Client Closed Request Error

This error is returned when a request is cancelled by the user.

- `message` (string, optional)
- `id` (string, optional)

### 500 Internal Server Error

This error is returned when an uncategorised internal server error occurs.

- `message` (string, optional)
- `id` (string, optional)

### 501 Not Implemented Error

This error is returned when the requested feature is not implemented.

- `message` (string, optional)
- `id` (string, optional)

### 503 Service Unavailable Error

This error is returned when the service is unavailable. This could be due to: - Too many users trying to access the service at the same time

- `message` (string, optional)
- `id` (string, optional)

### 504 Gateway Timeout Error

This error is returned when a request to the server times out. This could be due to: - An internal services taking too long to respond

- `message` (string, optional)
- `id` (string, optional)

## Types

### GenerateStreamEndResponse

- `id` (string, required)
- `prompt` (string, optional)
- `generations` (list of SingleGenerationInStream, optional)

### SingleGenerationInStream

- `id` (string, required)
- `text` (string, required) — Full text of the generation.
- `finish_reason` (enum, required)
  - Allowed values: `COMPLETE`, `STOP_SEQUENCE`, `ERROR`, `ERROR_TOXIC`, `ERROR_LIMIT`, `USER_CANCEL`, `MAX_TOKENS`, `TIMEOUT`
- `index` (integer, optional) — Refers to the nth generation. Only present when `num_generations` is greater than zero.

## Examples

**Request**

```json
{
  "prompt": "Please explain to me how LLMs work",
  "stream": true
}
```

**Response**

```json
[
  {
    "event_type": "text-generation",
    "is_finished": true,
    "text": "string"
  }
]
```

**SDK Code**

```python
import requests

url = "https://api.cohere.com/v1/generate"

payload = {
    "prompt": "Please explain to me how LLMs work",
    "stream": True
}
headers = {
    "X-Client-Name": "my-cool-project",
    "Authorization": "Bearer <token>",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
```

```javascript
const url = 'https://api.cohere.com/v1/generate';
const options = {
  method: 'POST',
  headers: {
    'X-Client-Name': 'my-cool-project',
    Authorization: 'Bearer <token>',
    'Content-Type': 'application/json'
  },
  body: '{"prompt":"Please explain to me how LLMs work","stream":true}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go
package main

import (
	"fmt"
	"strings"
	"net/http"
	"io"
)

func main() {

	url := "https://api.cohere.com/v1/generate"

	payload := strings.NewReader("{\n  \"prompt\": \"Please explain to me how LLMs work\",\n  \"stream\": true\n}")

	req, _ := http.NewRequest("POST", url, payload)

	req.Header.Add("X-Client-Name", "my-cool-project")
	req.Header.Add("Authorization", "Bearer <token>")
	req.Header.Add("Content-Type", "application/json")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby
require 'uri'
require 'net/http'

url = URI("https://api.cohere.com/v1/generate")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Post.new(url)
request["X-Client-Name"] = 'my-cool-project'
request["Authorization"] = 'Bearer <token>'
request["Content-Type"] = 'application/json'
request.body = "{\n  \"prompt\": \"Please explain to me how LLMs work\",\n  \"stream\": true\n}"

response = http.request(request)
puts response.read_body
```

```java
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.post("https://api.cohere.com/v1/generate")
  .header("X-Client-Name", "my-cool-project")
  .header("Authorization", "Bearer <token>")
  .header("Content-Type", "application/json")
  .body("{\n  \"prompt\": \"Please explain to me how LLMs work\",\n  \"stream\": true\n}")
  .asString();
```

```php
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('POST', 'https://api.cohere.com/v1/generate', [
  'body' => '{
  "prompt": "Please explain to me how LLMs work",
  "stream": true
}',
  'headers' => [
    'Authorization' => 'Bearer <token>',
    'Content-Type' => 'application/json',
    'X-Client-Name' => 'my-cool-project',
  ],
]);

echo $response->getBody();
```

```csharp
using RestSharp;

var client = new RestClient("https://api.cohere.com/v1/generate");
var request = new RestRequest(Method.POST);
request.AddHeader("X-Client-Name", "my-cool-project");
request.AddHeader("Authorization", "Bearer <token>");
request.AddHeader("Content-Type", "application/json");
request.AddParameter("application/json", "{\n  \"prompt\": \"Please explain to me how LLMs work\",\n  \"stream\": true\n}", ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
```

```swift
import Foundation

let headers = [
  "X-Client-Name": "my-cool-project",
  "Authorization": "Bearer <token>",
  "Content-Type": "application/json"
]
let parameters = [
  "prompt": "Please explain to me how LLMs work",
  "stream": true
] as [String : Any]

let postData = JSONSerialization.data(withJSONObject: parameters, options: [])

let request = NSMutableURLRequest(url: NSURL(string: "https://api.cohere.com/v1/generate")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "POST"
request.allHTTPHeaderFields = headers
request.httpBody = postData as Data

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```