open to work: staff / lead SRE

Reliability for AI systems
at GPU scale.

profile signal

I'm Sheshank Dudaboina, a Staff / Lead Site Reliability Engineer with 9+ years building production-grade infrastructure for model serving, Kubernetes, multi-cloud platforms, and the observability loops that keep them calm under load.

9+
years SRE
70+
Kubernetes clusters
~20
cloud providers
SLO
driven ops
inference-fleet/us-central1
healthy
gpu_utilization86%
p99_latency118ms
error_budget96.4%
sheshank@gpu-control-plane ~
❯ kubectl get nodes -l accelerator=nvidia -o wide
NAME GPU STATUS REGION VERSION
a100-pool-1 8 Ready us-central1 v1.29.2
h100-pool-2 8 Ready us-east1 v1.29.2
l4-burst-3 4 Ready us-west1 v1.29.2
❯ cat reliability.yaml # production priorities
role: SRE - AI Infrastructure
focus: ai platforms, observability, automation
currently: GPU infrastructure @ Baseten
status: open to staff / lead SRE roles
❯

01.Skills & Tools

KubernetesGPU infrastructureAWSGCPGrafanaVictoriaMetricsLokiFlux CDPrometheusincident.ioLinuxTerraformKServe / vLLM / TritonAI platforms

02.Latest Writing

All posts

03.Projects

All projects

Grafana Dashboard Generator

Generates Grafana dashboards from a simple YAML SLO definition. Supports burn-rate alerts and error budget visualization.

GoGrafonnetKubernetesHelm

k8s-runbook-operator

A Kubernetes operator that attaches runbooks to PodDisruptionBudgets and automatically links them in PagerDuty alerts.

Gocontroller-runtimePagerDuty API

NodeRelay

A hackathon project that became the production, keyless node-access path across a large multi-cloud GPU fleet.

KubernetesRBACGitOpsMulti-cloud