NodeRelay: keyless node access across a multi-cloud GPU fleet
A single, audited path to a host shell through the Kubernetes control plane—without distributing provider-specific SSH keys across the on-call team.
Access was slower than diagnosis.
The fleet ran production GPU inference across many clouds and GPU providers. Each provider came with a different SSH key, username, bastion, console flow, and network path. During an incident, the engineer first had to identify the provider and reconstruct its access procedure before debugging could begin.
Static keys spread across systems also made access difficult to audit and revoke. I took the problem on during an internal hackathon, built the first working path, and then carried it through security review, operational hardening, and adoption for daily production use.
Faster access without weaker controls.
Use the control plane every cluster already has.
NodeRelay combines a CLI with a node-local agent. Sessions follow the existing Kubernetes control path, so authentication, authorization, and audit come from the platform rather than a parallel credential system.
Documented contingency procedures cover cases where the normal Kubernetes control path is unavailable, without exposing those internal procedures publicly.
The platform records authenticated identity, target, and timestamp so access can be reviewed without maintaining a separate credential trail.
Five steps from identity to diagnostic access.
Choose a cluster
The CLI reads existing kubeconfig contexts and reuses each cluster's cloud IAM or OIDC authentication.
Choose a node
NodeRelay discovers the access agent associated with the selected host.
Authorize the session
The request follows the Kubernetes control path, where the API server verifies identity and role-based access.
Open a diagnostic shell
The node agent establishes a controlled host-level troubleshooting session without introducing provider-specific credentials.
Leave an audit trail
The platform records the authenticated engineer, target, and timestamp for the access request.
Identity, policy, and controlled change.
Access is granted to an identity-provider group rather than individual users and is limited to the capabilities required for a session. Removing group membership revokes access consistently across the fleet.
The more important boundary is who can change NodeRelay itself. The deployment and its authorization policy are managed from a single GitOps source, with reviewed changes and tightly controlled write access.
A production tool should make its risks explicit.
The control path must be healthy
The primary workflow depends on parts of the cluster control path. Documented break-glass procedures cover failures outside that path.
Audit depth has limits
The initial design focused on identity and session-level accountability. Deeper session recording remains an area for improvement.
Host access requires elevated capability
The agent's elevated capability is constrained through image controls, isolated policy, group-based authorization, and reviewed GitOps changes.
Standing group access
SRE group members can open sessions at any time. Just-in-time, time-limited grants with approval would further reduce exposure.
From hackathon demo to the on-call team's default.
NodeRelay made node access one command across the fleet, removed provider-specific SSH keys from the normal workflow, and put every session behind the same identity, RBAC, GitOps, and audit controls. The core lesson was simple: the best way to fix access sprawl was not another access system—it was routing access through the control plane already present in every cluster.