Overview
Cloud Site Reliability Engineer (SRE) L2 role within Infrastructure Engineering, also titled Tech Lead, focusing on reliability, scalability, and performance for cloud infrastructure.
What you'll do
- Design and operate fault-tolerant, highly available architectures across AWS, Azure, and GCP.
- Deploy, manage, and optimize cloud resources using Infrastructure as Code (IaC) with Terraform and Ansible.
- Implement monitoring, alerting, and logging using tools such as Splunk, Azure Monitor, Dynatrace, and AWS CloudWatch.
- Lead incident response including triage, mitigation, root cause analysis, and post-incident reviews.
- Drive capacity planning and scaling improvements using forecasting and autoscaling.
- Build automation and internal tooling to reduce manual toil using Python, PowerShell, or Bash.
- Collaborate with Security to implement secure infrastructure practices.
What you'll need
- 5+ years of hands-on programming/scripting experience in Python, PowerShell, Bash, or equivalent.
- Strong experience with AWS, Azure, or GCP and core services including VPC/VNet, IAM, serverless patterns, and managed Kubernetes services.
- Experience with containers and orchestration technologies including Docker and Kubernetes.
- Proficiency in IaC using Terraform and/or Ansible.
- Experience with observability stacks and operational monitoring using Splunk, Azure Monitor, Dynatrace, or AWS CloudWatch (or similar).
- Advanced knowledge of Windows and Linux/Unix environments including system administration and networking fundamentals.
- Proven incident response capability with troubleshooting skills under pressure.
Nice to have
- Cloud certifications such as AWS Certified Solutions Architect, Google Cloud Professional DevOps Engineer, or Azure DevOps Engineer.
- Experience with chaos engineering or resilience testing frameworks.
- Experience supporting multicloud and/or hybrid cloud deployments.
- Familiarity with SLOs, SLIs, and error budgets.
- Experience gathering operational feedback and driving improvement solutions using Azure services.
Details
- On-site work Monday through Friday.
- Work location listed as Bangalore OR Chennai.
Read the full description and apply on the company’s own careers page.