Overview
Serve as a Staff Site Reliability Engineer and Tech Lead for an SRE pod supporting the Guidewire Cloud Platform. Provide deep technical expertise, cross-team architectural influence, and technical leadership while driving reliable, self-healing cloud operations.
What you'll do
- Serve as the primary technical authority for the SRE pod.
- Own the technical design for pod deliverables, oversee production deployments, and act as the final point of escalation for complex technical impediments.
- Design and implement systemic improvements to AWS and Kubernetes-based environments.
- Lead the creation of frameworks that prevent incidents across the organization.
- Break down high-level requirements into actionable technical tasks and protect the pod’s sprint goals.
- Write code in AI, Python, Go, or Java, conduct code reviews, and promote Clean Code and testing standards.
- Triage and assign work.
- Lead the development of automated runbooks, self-healing systems, and advanced CI/CD pipelines using Terraform and GitOps.
- Coach and mentor P1–P5 engineers and provide constructive feedback to support professional growth.
- Lead responses to high-severity incidents and facilitate blameless post-mortems.
- Identify opportunities to integrate AI/ML and data-driven insights into the Datadog/Resolve observability stack.
- Participate in and provide technical oversight for a 24x7 on-call rotation.
- Act as a high-level escalation point for complex production incidents and drive post-incident analysis to eliminate recurring issues.
- Take ownership of the pod’s technical roadmap, identify opportunities for automation, and serve as a technical resource for the team.
- Lead major cross-pod initiatives such as application migrations or upgrades, observability standards, or deployment-velocity improvements.
- Influence global cloud strategy and shape the future of SRE at Guidewire.
What you'll need
- 10+ years of experience in SRE, DevOps, or Software Engineering.
- At least 3+ years in a Staff-level, Technical Lead, Manager or Sr. Manager capacity.
- Expert-level command of the AWS ecosystem and Kubernetes (EKS) orchestration.
- A proven track record of managing large-scale, distributed cloud environments.
- Advanced proficiency in Infrastructure as Code using Terraform/CloudFormation.
- Experience with configuration management using Ansible/Puppet within a GitOps framework.
- A strong software engineering foundation and the ability to write production-grade code in Python, Go, or Java or with tools such as Claude Code, CodeEx or CoPilot.
- Deep experience building comprehensive monitoring, SLIs/SLOs, and tracing strategies using tools like Datadog, Prometheus, or Splunk.
- Proven ability to manage “up and out,” communicate technical risks to stakeholders, and collaborate with Product and Engineering managers to align reliability goals with feature delivery.
- Experience managing on-call rotations.
- A “safety-first” mindset and understanding that reliability is the most important feature of any product.
Details
- Location: Bangalore, India.
- Participate in a 24x7 on-call rotation.
- Provide technical oversight for production incidents and high-level escalation support.
Read the full description and apply on the company’s own careers page.