Overview
Lead Site Reliability Engineer responsible for driving reliability efforts, building resilient systems, improving observability and automation, and collaborating on reliable, scalable, high-performance services.
What you'll do
- Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
- Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
- Lead incident response and post-incident reviews, including blameless postmortems, robust root cause analysis, and continuous improvement of systems.
- Automate incident detection and response using automated runbooks or predefined workflows.
- Write software as needed to support reliability or efficiency needs.
- Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
- Use capacity planning, forecasting, and performance testing to ensure systems scale effectively as the user base and load grow.
- Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
- Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.
- Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
- Get involved in chaos engineering initiatives.
- Participate on the on-call rotation.
- Drive advanced alerting and anomaly detection applied to metrics.
What you'll need
- 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
- Deep understanding of Linux systems, networking, and systems administration.
- Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
- Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
- Strong skills in at least one programming language, Python or Go, to write production level code.
- Strong skills in shell scripting using bash or similar.
- Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
- Experience with Chaos Engineering methodologies and tools, including Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.
- Reliability-focused mindset with the ability to balance fast product iterations and system stability.
- Solid understanding of SLOs, SLIs, and error budgets.
- Hands-on knowledge of CI/CD pipelines and infrastructure automation.
- Proven expertise in incident management, postmortems, and root cause analysis.
- Knowledge of modern deployment strategies, such as blue-green deployments and canary releases, and resiliency patterns, such as circuit breakers and retry mechanisms.
Nice to have
- Experience with distributed systems.
- Experience with statistical analysis applied to metrics.
- Familiarity with high-performance, low-latency systems.
- Strong problem solving skills.
- Experience as an on-call engineer.
- Experience writing production level code.
- Hands-on experience running Chaos Engineering drills and initiatives.
Details
- Location: Bengaluru, Karnataka, India.
Read the full description and apply on the company’s own careers page.