Sheshank Dudaboina
Staff / Lead Site Reliability Engineer
I've spent 9+ years making distributed systems more reliable, from global consumer services to multi-cloud Kubernetes and production GPU inference fleets. I currently operate AI infrastructure at Baseten and am exploring Staff or Lead SRE roles where I can improve reliability, platform engineering, observability, and operational leverage.
01.Experience
Operate production GPU inference infrastructure across 70+ Kubernetes clusters and roughly 20 cloud providers. Build clusters and their model-serving, networking, storage, autoscaling, GitOps, and telemetry stacks; lead P0 response and vendor escalations across large NVIDIA H100, B200, and A100 fleets.
Founded the platform SRE function and owned reliability, observability, cost efficiency, and production operations across multi-region AWS and EKS. Built the centralized VictoriaMetrics, Grafana, Loki, and Alertmanager platform and drove resource right-sizing that saved about $15K per month.
Contributed to planning a payment-platform migration from an on-premises data center to AWS, including secure and highly available architecture plus reusable Terraform and Terragrunt templates.
Operated large-scale AWS production environments supporting global PlayStation services, including the PS5 launch. Led a Datadog migration, moved legacy workloads to EKS, automated infrastructure, and built safer CI/CD pipelines that reduced deployment time by about 40%.
Designed AWS infrastructure and multi-region deployments for high-throughput financial systems, built Kubernetes platforms with KOPS and GitLab CI/CD, and improved observability for real-time transaction workloads.
Implemented cloud automation across AWS and OpenStack, deployed highly available architectures, and automated infrastructure and configuration with Terraform and Ansible.
02.Skills
- ›AWS
- ›GCP
- ›Azure
- ›OCI
- ›GPU neoclouds
- ›Terraform / CDKTF
- ›Atlantis
- ›Kubernetes (EKS, GKE, RKE2)
- ›Karpenter
- ›KEDA
- ›Helm
- ›Kustomize
- ›Flux CD
- ›ArgoCD
- ›Cilium
- ›Istio
- ›KServe
- ›vLLM
- ›Triton
- ›NVIDIA H100 / B200 / A100
- ›GPU Operator
- ›DCGM
- ›NCCL
- ›MIG
- ›InfiniBand
- ›Grafana
- ›VictoriaMetrics
- ›Prometheus
- ›Loki
- ›Alertmanager
- ›OpenTelemetry
- ›Datadog
- ›New Relic
- ›ELK
- ›incident.io
- ›P0 incident command
- ›SLOs / SLIs
- ›Error budgets
- ›Capacity planning
- ›Cost optimization
- ›Python
- ›Go
- ›Bash
- ›GitHub Actions
- ›Jenkins
- ›Harness
- ›PromQL / MetricsQL
- ›LogQL
03.Education
04.Certifications
I write about what I learn on the job.
SRE / Kubernetes / GPU platforms / Observability
Read my writing →