Z

Lead Site Reliability Engineer

Zeta Global
NewPosted yesterday

LOCATION

Bengaluru · Onsite

EXPERIENCE

3 - 5 Years

SALARY

Negotiable

SKILLS REQUIRED

ObservabilitySite Reliability EngineeringInfrastructure as CodeIncident response

Job description

Overview

Lead Site Reliability Engineer responsible for driving reliability efforts, building resilient systems, improving observability and automation, and collaborating on reliable, scalable, high-performance services.

What you'll do

  • Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
  • Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
  • Lead incident response and post-incident reviews, including blameless postmortems, robust root cause analysis, and continuous improvement of systems.
  • Automate incident detection and response using automated runbooks or predefined workflows.
  • Write software as needed to support reliability or efficiency needs.
  • Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
  • Use capacity planning, forecasting, and performance testing to ensure systems scale effectively as the user base and load grow.
  • Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
  • Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.
  • Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
  • Get involved in chaos engineering initiatives.
  • Participate on the on-call rotation.
  • Drive advanced alerting and anomaly detection applied to metrics.

What you'll need

  • 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
  • Deep understanding of Linux systems, networking, and systems administration.
  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
  • Strong skills in at least one programming language, Python or Go, to write production level code.
  • Strong skills in shell scripting using bash or similar.
  • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
  • Experience with Chaos Engineering methodologies and tools, including Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.
  • Reliability-focused mindset with the ability to balance fast product iterations and system stability.
  • Solid understanding of SLOs, SLIs, and error budgets.
  • Hands-on knowledge of CI/CD pipelines and infrastructure automation.
  • Proven expertise in incident management, postmortems, and root cause analysis.
  • Knowledge of modern deployment strategies, such as blue-green deployments and canary releases, and resiliency patterns, such as circuit breakers and retry mechanisms.

Nice to have

  • Experience with distributed systems.
  • Experience with statistical analysis applied to metrics.
  • Familiarity with high-performance, low-latency systems.
  • Strong problem solving skills.
  • Experience as an on-call engineer.
  • Experience writing production level code.
  • Hands-on experience running Chaos Engineering drills and initiatives.

Details

  • Location: Bengaluru, Karnataka, India.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Lead Site Reliability Engineer