Kubernetes & platform engineering
Designing paved roads that help teams ship reliably across cloud-native environments.
I'm Ravindra Bhargava, an experienced AI infrastructure and platform engineering leader. I build secure, scalable cloud-native platforms, operate NVIDIA H100 and A100 GPU workloads on Kubernetes, and create MCP servers and automation for modern AI systems.
Explore my areas of focusAreas of focus
Designing paved roads that help teams ship reliably across cloud-native environments.
Building security into delivery systems, from source and dependencies to runtime.
Building Kubernetes platforms to schedule, scale, and operate NVIDIA H100 and A100 GPU workloads.
Applying reusable infrastructure patterns across GCP, AWS, and Azure.
Making systems easier to understand, operate, and optimize over their full lifecycle.
Writing
Future essays on the systems, practices, and tradeoffs behind AI infrastructure, cloud platforms, and automation.
Cloud observability
Lessons from designing a unified telemetry platform for applications running on Cloud Run and Google Kubernetes Engine.
Read articlePlatform engineering
A practical architecture for provisioning Google Kubernetes Engine with Terraform while keeping application delivery, security, and operations maintainable.
Read articlePlatform engineering
Principles for treating internal platforms as products, with clear interfaces and sensible defaults.
Topic previewSpeaking & community

Speaking
OpenTelemetry and GitOps
Mentoring
Supporting the next generation of AI practitioners
Community
Reviewing and recognizing thoughtful technical work
About
I specialize in AI infrastructure, cloud platforms, and Kubernetes, with hands-on experience managing NVIDIA H100 and A100 GPU workloads. My focus is creating technical foundations that help teams move quickly without compromising reliability, security, or trust.
This site is where I share practical lessons about Kubernetes, DevSecOps, multi-cloud architecture, observability, GPU platform operations, MCP server development, and automation at scale.