Overview
We are looking for a hands-on AVP Site Reliability Engineer for the CaaS Private on-premises Kubernetes platform built on Google Distributed Cloud. The role focuses on Kubernetes platforms, deep troubleshooting, automation, observability, incident response and continuous improvement across complex infrastructure and tenant workloads.
What you'll do
- Define, implement, and continuously improve service-level indicators, service-level objectives, alerting standards, and error budgets for the CaaS Private platform and its critical services.
- Build and maintain observability across metrics, logs, traces, alerts, and dashboards to provide insight into platform health, saturation, latency, capacity, and failure modes.
- Investigate and resolve complex production issues across Kubernetes, Linux, networking, storage, ingress, service mesh, node services, platform dependencies, and tenant workloads.
- Lead or coordinate incident response for platform-impacting events, ensuring timely mitigation, clear stakeholder communication, blameless post-incident reviews, and durable follow-up actions.
- Automate repetitive operational tasks, diagnostics, and remediation workflows to reduce toil, improve consistency, and accelerate recovery.
- Improve platform reliability, upgrade safety, resilience, capacity planning, performance, disaster recovery, and operational readiness.
- Improve runbooks, dashboards, alerts, support processes, and engineering standards based on recurring issues and operational learning.
- Partner with platform, network, security, storage, and application teams on release readiness, troubleshooting, change execution, documentation, and adoption of SRE practices.
- Contribute to reliability-focused platform enhancements and use system-level insights to improve platform stability and tenant experience.
- Support the global GDC/Kubernetes platform in a 24x7 follow-the-sun model, including on-call, weekend, public-holiday rotation, and early Monday coverage as required.
What you'll need
- A bachelor’s degree in a technical or engineering discipline.
- 10–12 years of hands-on experience in Site Reliability Engineering, Production Engineering, DevOps or a closely related infrastructure role.
- Strong hands-on Kubernetes expertise, including cluster operations, upgrades, troubleshooting, networking, storage, security, Helm, Operators, and workload lifecycle management in bare-metal or private-cloud environments.
- Demonstrated SRE mindset and experience with SLI/SLO/SLA management, error budgets, incident reduction, capacity planning, performance optimization, resilience engineering, and operational excellence.
- Strong Linux system administration and networking fundamentals, with the ability to troubleshoot complex infrastructure and distributed-system failures.
- Experience building and operating observability solutions using Prometheus, Grafana, Splunk, distributed tracing, logging, alerting, dashboards, and OpenTelemetry-style concepts.
- Strong automation and infrastructure scripting experience using Python, Bash, and Ansible.
- Experience operating and troubleshooting CI/CD and GitOps platforms such as Argo CD, Jenkins, GitHub or Bitbucket, and Artifactory.
- Solid understanding of incident management, root-cause analysis, operational readiness, virtualization, containerization, and distributed-system behavior under failure.
- Ability to communicate clearly, collaborate across teams, prioritise under pressure, and independently take complex problems from symptom to sustainable resolution.
- Ability to use AI tools to improve productivity and workflows while applying critical judgement and ensuring responsible, ethical use of data and AI-generated outputs.
Nice to have
- CKA or CKAD certification, or equivalent demonstrable Kubernetes expertise.
- Experience with Istio or Envoy, service-mesh observability, traffic management, Cilium, OPA Gatekeeper, admission controls, or policy-driven operational guardrails.
- Experience with software-defined or container-native storage, including Ceph/Rook, CSI, or comparable storage platforms.
- Familiarity supporting stateful services such as PostgreSQL, Kafka, MongoDB, or comparable platform dependencies.
- Practical experience with load testing, chaos testing, failure injection, alert tuning, self-healing automation, and disaster-recovery validation.
- Exposure to regulated or low-latency environments with strict uptime, change-control, compliance, time-synchronization, deterministic-performance, or SR-IOV workload requirements.
- Ability to read and understand Golang code while troubleshooting platform components.
Details
- Location: Bangalore, India.
- Support the global GDC/Kubernetes platform in a 24x7 follow-the-sun model.
- On-call, weekend, public-holiday rotation, and early Monday coverage are required as needed.
Read the full description and apply on the company’s own careers page.