> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hyperbolic.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Managed Kubernetes

> A Kubernetes cluster on your GPU instances, with the control plane run by Hyperbolic

<Note>
  Managed Kubernetes is in **beta**. It works, teams are running on it, and the surface is still changing. Expect the feature list below to grow, and tell us what's missing.
</Note>

Managed Kubernetes gives you a ready-to-use GPU Kubernetes cluster without running a control plane yourself. You choose the GPU nodes; Hyperbolic provisions the cluster, installs the NVIDIA stack, and gives you a single sign-on kubeconfig that authenticates `kubectl` as **you**.

## What you get

* **A hosted control plane**, run and kept available by Hyperbolic, located in a cloud region close to your GPU nodes.
* **GPU worker nodes** that are your rented instances — same hardware, same networking, same billing as renting them directly.
* **NVIDIA GPU Operator and Network Operator** preinstalled, so GPUs, drivers, and RDMA networking are exposed to pods on day one.
* **Single sign-on access.** You download a per-cluster kubeconfig; the first `kubectl` command signs you in through Hyperbolic in your browser and authenticates as **you** — short-lived, with every action attributed to your identity instead of a shared admin credential. Org admins get `cluster-admin`; other members get `edit`. RBAC inside the cluster is yours to extend. See [Connect with kubectl](/docs/managed-kubernetes/kubectl-access).
* **Multi-node performance without overhead.** NCCL all-reduce inside a pod matches the same test run on the bare host.

## Create a cluster

1. In the console, open **Kubernetes** and choose **Create cluster**.
2. Pick the GPU type, node count, and region. Nodes in one cluster are in the same region and on the same interconnect.
3. Wait for the cluster to show **Ready**. Provisioning bootstraps the control plane, joins the nodes, and installs the add-ons.
4. Install the `kubectl oidc-login` plugin (one-time) and click **Download kubeconfig** — see [Connect with kubectl](/docs/managed-kubernetes/kubectl-access) for the install commands. Then point `kubectl` at the file:

```bash theme={null}
export KUBECONFIG=~/Downloads/hyperbolic-cluster.yaml
kubectl get nodes -o wide
```

The first command opens your browser to sign in with Hyperbolic; after that, `kubectl` runs as you. You should see one node per GPU instance, each advertising `nvidia.com/gpu` capacity.

<Tip>
  You can still rent interconnected multi-node clusters **without** Kubernetes — see [On-Demand](/docs/on-demand/overview). Managed Kubernetes is one way to run on those nodes, not the only way.
</Tip>

## Node pools

Nodes are grouped into pools. From the cluster's **Pools** tab you can see each pool's GPU type, size, and status, and add or remove nodes. CPU-only worker pools for non-GPU workloads (proxies, monitoring, controllers) are on the roadmap.

## Running GPU workloads

Request GPUs like any other resource:

```yaml theme={null}
resources:
  limits:
    nvidia.com/gpu: 8
```

For multi-node training, the Network Operator exposes the RDMA interfaces to pods; use your framework's usual NCCL launcher. The NCCL and interconnect checks in [Verifying Instance Performance](/docs/on-demand/verifying-performance) apply unchanged inside Kubernetes.

## Exposing services

Nodes have public IPs. A `NodePort` or `hostNetwork` service is reachable from the internet only once the port is open on the node — see [Opening inbound ports](/docs/on-demand/storage-and-ports#opening-inbound-ports). Prefer an in-cluster ingress plus one open port over exposing many NodePorts.

## Storage

Node-local NVMe is available to pods as `hostPath` or a local volume and does **not** survive node replacement. Shared network storage for clusters is in development; until it ships, use your own object storage for checkpoints.

## Monitoring

The cluster page in the console shows the pods and events on each node and live GPU telemetry. You are free to install Prometheus, Grafana, or any observability agent in the cluster. A fuller **Monitoring** tab with history is in development.

## What Hyperbolic runs, and what you run

| Hyperbolic | You |
| - | - |
| Control plane availability and upgrades | Everything scheduled on the cluster |
| Node provisioning, GPU drivers, add-on installation at bootstrap | RBAC, namespaces, network policy |
| Single sign-on — authenticating users as themselves, short-lived | RBAC bindings for your own teams and service accounts |
| Node replacement on hardware failure | Add-on configuration after bootstrap — if you modify or remove an add-on, that's the new state |
| Platform telemetry | Application monitoring, logging, alerting |

Managed Slurm is on the roadmap; Kubernetes is the supported orchestration path today.

## Billing

GPU nodes are billed exactly like the same instances rented directly. The hosted control plane is billed separately and shown on the cluster's billing tab.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.