> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.cohere.com/v2/docs/model-vault/monitoring/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.cohere.com/_mcp/server. # Monitoring > Monitor latency, queuing, and GPU utilization for any vault with the Grafana dashboard. If you click into a Vault, you will see a `Monitoring` button in the top-right corner. Clicking it opens a Grafana dashboard which offers various analytics into the performance of this particular Vault, such as: * First Token Latency * Queuing Latency * Average GPU Duty Cycle ![](/_fern-img/d396b458af124efd7cf0aa98282de5303e8728e77f9292ccda2a9888828cd7f9.webp) This lets you gather analytics related to specific models, modify the time range over which your analytics are gathered, inspect various on-page graphs, or export and share your data. You can change the model with the `Model` dropdown in the top-left corner, use the **Search** bar at the top of the screen to find particular pieces of information quickly and easily, and refresh your data by clicking **Refresh** at the top of the screen. > **Note** > > Performance monitoring is available for all vaults. For encrypted vaults, these operational metrics > are derived from infrastructure telemetry and do not expose the contents of your prompts or responses, > which remain protected inside the trusted execution environment. ## Next steps * [Managing Vaults](../../docs/model-vault/managing-vaults) * [Calling a Vault over the API](../../docs/model-vault/standard/api-access) > Cohere's API documentation helps developers easily integrate natural language processing and generation into their products.