- Worked in the Kubernetes-focused DevOps leadership role supporting the India Technology Hub
- Focused on Kubernetes platform engineering, DevOps practices, and cloud-native platform reliability.
Profile Summary
AKS-focused DevOps and Site Reliability Engineer with 15+ years of experience building and managing secure, scalable, and highly available Azure platforms. Skilled in Kubernetes (AKS), Helm deployments, Terraform, Bicep, Docker, and CI/CD pipelines using Azure DevOps, GitHub Actions, and Jenkins. Experienced in monitoring and observability using Prometheus, Grafana, Splunk, Dynatrace, and Azure Monitor. Strong knowledge of DevSecOps, GitOps, RBAC, Azure Key Vault, ArgoCD, and Flux. Proven ability to improve platform reliability, automate infrastructure, handle production incidents, and lead cross-functional teams in enterprise environments.
Professional Experience
- Led DevOps & SRE delivery for large-scale Azure SaaS applications across enterprise clients.
- Designed, deployed, and operated production AKS clusters with upgrades, scaling, backups, and day-2 reliability operations.
- Implemented multi-environment IaC using Terraform and Bicep with reusable modules, workspaces, and remote state best practices.
- Collaborated with development teams to containerize workloads and standardize Kubernetes manifests and Helm charts.
- Implemented observability and alerting with Prometheus, Grafana, Splunk, and Dynatrace to improve platform visibility.
- Enforced Kubernetes security controls including RBAC, policy guardrails, and secure secret management practices.
- Led on-call incident response and reliability improvements, reducing MTTR for production services.
- Implemented GitOps delivery workflows using ArgoCD and Flux, standardizing Helm chart-based deployments across AKS environments with automated sync, rollback, and drift detection.
- Built and maintained CI/CD workflows using Azure DevOps, GitHub Actions, and Jenkins for automated build, test, and deployment processes.
- Architected multi-region highly available Azure environments and led on-premises to cloud migration initiatives.
- Supported enterprise cloud and hybrid infrastructure operations with a strong focus on Azure
- Implemented monitoring and alerting standards to improve early detection of service degradation
- Handled incident triage, escalation, and service recovery activities while maintaining SLA
- Managed IIS-hosted applications across virtualized and cloud environments
- Performed performance tuning and reliability optimization for critical workloads
- Collaborated with application, database, and infrastructure teams to resolve cross-stack issues