> ## Documentation Index
> Fetch the complete documentation index at: https://docs.hyperbolic.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Cluster Handover

> What we validate before delivering a Private Cloud cluster, and what your handover document contains

Every Private Cloud cluster is delivered with a handover document specific to your deployment. This page describes what we check before handover and what the document contains, so you know what to expect and what to verify on first login.

## Before handover

A cluster is not delivered until it has passed:

1. **Acceptance certification.** A benchmark run across the delivered nodes — GPU health, NCCL collectives (all-reduce, reduce-scatter, all-gather), per-rail and pairwise fabric bandwidth, and correctness — scored against pass thresholds. The certification run ID and score are recorded in your document.
2. **Hardening validation.** The effective SSH daemon configuration is audited (`sshd -T`) for public-key-only authentication; a password-only login attempt must fail with `Permission denied (publickey)`; default accounts are confirmed locked; the firewall or security-group policy is reviewed against the intended exposure.
3. **Account cleanup.** Provisioning, installer, and partner accounts are removed. Your user and public key are installed; the sudo decision (passwordless or not) is recorded.
4. **Data sanitization.** Drives that have been used before are block-erased at the hardware level and fresh filesystems created. Nothing from a previous tenant is copied forward.
5. **Runtime checks.** Docker and the NVIDIA Container Toolkit are verified on every node with a disposable GPU container that must see all GPUs; the test containers and images are removed afterwards.
6. **Cleanup.** Benchmark containers, temporary files, and benchmarking artifacts are removed. No benchmark jobs are left running at handover unless coordinated with you.
7. **Baseline capture.** OS, kernel, driver, container runtime, and RDMA port state are recorded per node.

<Note>
  When nodes are **added** to an existing cluster, the additions get access, runtime, and sanitization checks and a fresh baseline. The original certification is not automatically extended to the combined cluster — ask for a re-certification if you want a benchmark that covers the new topology.
</Note>

## What the handover document contains

| Section | What's in it |
| - | - |
| **Access and node inventory** | Hostname, public IP, and internal IP for every node; the SSH command and user; the fingerprint of the key we installed (never the key itself); sudo state |
| **System baseline** | Per node: OS release, kernel, NVIDIA driver, Docker version, active RDMA ports — recorded on the handover date |
| **GPU topology** | GPU model, GPUs per node, total GPUs, enumeration check, and the command to view per-node topology (`nvidia-smi topo -m`) |
| **Fabric** | Interconnect type — InfiniBand, or RoCE over Ethernet — active RDMA ports per node, observed link speed, and the verification commands (`ibstat`, `ibv_devinfo`). GPUDirect RDMA is verified with a GPU-aware NCCL run |
| **Storage** | Local NVMe layout per node (for example a striped LVM/ext4 volume mounted at `/data`), root filesystem size and usage, sanitization status, and any shared or network mounts. Anything marked not mounted or not recorded is yours to verify before the first workload |
| **Container runtime** | Docker, NVIDIA Container Toolkit, containerd, Buildx, and Compose versions; default runtime; GPU container smoke-test result |
| **Recommended NCCL configuration** | The environment variables used during certification — HCA list, GID index, `NCCL_CROSS_NIC`, `NCCL_SOCKET_IFNAME`, and the rest. These were tuned for multi-node jobs on this fabric; don't apply the forced fabric settings to single-node jobs, and confirm library compatibility before you rely on them |
| **Benchmark results** | Certification run, scorecard, and headline numbers — large-message all-reduce bandwidth, MFU, scaling efficiency — with the exact set of nodes they cover |
| **Verification commands** | A spot-check set for first login |
| **Handover status** | Each requirement above with its status and evidence |
| **Notes, support, and escalation** | Deployment-specific caveats, how to reach us, and your reference IDs |

## First login

Run these on each node and compare with your document:

```bash theme={null}
whoami && sudo -n whoami      # your user, passwordless sudo if recorded
hostname && uptime
nvidia-smi                    # all GPUs visible
nvidia-smi topo -m            # GPU <-> NIC topology
ibstat && ibv_devinfo         # RDMA ports active, link layer, rate
docker ps                     # nothing of ours left running
df -h / && lsblk -o NAME,SIZE,TYPE,FSTYPE,MOUNTPOINTS
```

Then confirm GPU access from a container:

```bash theme={null}
docker run --rm --gpus all --ipc=host --ulimit memlock=-1 --net=host \
  nvidia/cuda:12.8.1-base-ubuntu24.04 nvidia-smi
```

Anything that doesn't match the document — a missing mount, a port not active, a GPU not enumerated — goes to support with the node's short name and the command output.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.