> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hyperbolic.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# GPU Monitoring

> Real-time GPU metrics on supported instances, no setup required

Most On-Demand and Reserved instances ship with GPU monitoring enabled. There is nothing to install and no extra charge — metrics start appearing within a few minutes of the instance becoming ready. A few machine types do not support it yet; see [Instances without monitoring](#instances-without-monitoring) below.

## Open monitoring for an instance

From your instance's details page, click **Monitoring**. The view is scoped to that instance and lists each GPU attached to it.

## What you can see

| Metric | What it tells you |
| - | - |
| **GPU utilization** | Percentage of time each GPU was executing kernels. Sustained low utilization during training usually means a data-loading or CPU bottleneck, not a GPU problem. |
| **GPU memory** | Used vs. total per GPU. Running near the ceiling risks out-of-memory errors; see the [CUDA out of memory](/docs/on-demand/quickstart#gpu-issues) fix. |
| **Temperature** | Per-GPU. Sustained readings well above the pack are worth flagging to support — thermal throttling shows up as lower clocks and slower steps before it shows up as an error. |

Metrics refresh automatically about every 30 seconds.

## Instances without monitoring

Monitoring depends on the platform agent being present on the instance. If an instance's details page has no **Monitoring** button, that machine type doesn't support it yet; `nvidia-smi` on the instance gives you the same numbers on demand — see [Managing Instances](/docs/on-demand/managing-instances#real-time-gpu-monitoring).

## What monitoring is not

Monitoring is a read-only view. It does not alert you, and it does not restart or replace a GPU automatically. If you see a GPU that has dropped out or is throttling, run the checks in [Verifying Instance Performance](/docs/on-demand/verifying-performance) and contact support with the instance ID and the `nvidia-smi -q` output.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.