Overview
Intermediate Site Reliability Engineer focused on maintaining and operating large-scale, distributed, fault-tolerant systems.
What you'll do
- Maintain and execute Infrastructure as Code (IaC) using Terraform across public cloud environments (GCP/AWS).
- Deploy, support, and troubleshoot containerized microservices on Kubernetes (GKE/EKS).
- Use GitHub action workflows and groovy scripts for automated incident remediation and task automation.
- Set up and maintain observability dashboards, alerts, and metrics using tools such as Datadog.
- Participate in a 24/7 follow-the-sun operational rotation to manage and resolve production incidents.
- Collaborate with development teams on outage analysis and preventative actions via postmortems.
What you'll need
- Bachelor’s degree in Computer Science, Information Technology, or a related technical discipline.
- 4–10 years of overall technical experience across DevOps, SRE, Systems Administration, or Software Engineering.
- 1+ years running and maintaining applications in a public cloud environment (GCP preferred, AWS acceptable).
- Working experience with IaC using Terraform and container orchestration using Kubernetes/Docker.
- Practical scripting skills in Terraform for infrastructure automation.
- Experience with CI/CD pipeline automation (e.g., Jenkins, GitLab CI).
Nice to have
- Certified Kubernetes Application Developer (CKAD) or GCP Associate Cloud Engineer.
- Familiarity with DevSecOps practices and automated security scanning tools within build pipelines.
Details
- Location: Pune, India.
- Work setting: Hybrid.
Read the full description and apply on the company’s own careers page.