> This page is for version v1 API.
> For other versions, use one of these documentation indexes:
> - v2 API (default): https://docs.cohere.com/v2/llms.txt
> - v1 API: https://docs.cohere.com/v1/llms.txt

> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.cohere.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server.

# Monitoring

> Monitor latency, queuing, and GPU utilization for any vault with the Grafana dashboard.

If you click into a Vault, you will see a `Monitoring` button in the top-right corner. Clicking it
opens a Grafana dashboard which offers various analytics into the performance of this particular
Vault, such as:

* First Token Latency
* Queuing Latency
* Average GPU Duty Cycle

![](/_fern-img/d396b458af124efd7cf0aa98282de5303e8728e77f9292ccda2a9888828cd7f9.webp)

This lets you gather analytics related to specific models, modify the time range over which your
analytics are gathered, inspect various on-page graphs, or export and share your data.

You can change the model with the `Model` dropdown in the top-left corner, use the **Search** bar at
the top of the screen to find particular pieces of information quickly and easily, and refresh your
data by clicking **Refresh** at the top of the screen.

> **Note**
>
> Performance monitoring is available for all vaults. For encrypted vaults, these operational metrics
> are derived from infrastructure telemetry and do not expose the contents of your prompts or responses,
> which remain protected inside the trusted execution environment.

## Next steps

* [Managing Vaults](../../docs/model-vault/managing-vaults)
* [Calling a Vault over the API](../../docs/model-vault/standard/api-access)