Overview
Site Reliability Engineer II (Platform Engineering on the SRE track) for a fully AWS-hosted, multi-tenant SaaS platform. Build and operate monitoring, automation, and incident response systems to keep services reliable, observable, and secure.
What you'll do
- Build and maintain monitoring, alerting, and reliability tooling on an OpenTelemetry-based stack.
- Analyze production performance and error budgets to maintain SLIs and SLOs for owned services.
- Implement automated health checks, scaling rules, and self-healing mechanisms.
- Perform root cause analysis and contribute to post-incident reviews and permanent fixes.
- Build and maintain infra automation using Terraform across multi-account AWS organizations.
- Lead incident response and participate in a 24/7 PagerDuty on-call rotation.
What you'll need
- 2-4 years of hands-on experience in SRE, DevOps, or cloud engineering roles.
- Ability to operate production workloads on AWS (e.g., EC2, ECS/EKS, RDS, S3, IAM, VPC).
- Working proficiency with Terraform and Git-based CI/CD (GitHub Actions or similar).
- Solid scripting ability in Python or Bash for automation.
- Experience with observability practices and tools such as OpenTelemetry and CloudWatch.
- Networking and cloud security fundamentals.
- Bachelor's degree in Computer Science/Engineering or equivalent skills.
Nice to have
- AWS Certified SysOps Administrator, DevOps Engineer, CKA/CKAD, or Terraform Associate.
- Experience with multi-tenant SaaS or account-per-customer AWS architectures.
- Exposure to PagerDuty or equivalent incident management platforms.
- Experience operating AI/LLM-backed services or building agentic automation under governance controls.
Details
- Location: India.
- 24/7 PagerDuty on-call rotation is required.
Read the full description and apply on the company’s own careers page.