Overview
Senior Software Engineer focused on SRE responsibilities to improve system reliability, scalability, and performance.
What you'll do
- Design and build platforms, tools, and frameworks to improve reliability, scalability, and performance.
- Define and implement SRE best practices including SLIs/SLOs, error budgets, and reliability metrics.
- Lead incident response, perform root cause analysis, and implement long-term fixes to prevent recurrence.
- Analyze system behavior to identify bottlenecks and saturation points and improve resilience.
- Collaborate with engineering teams to prioritize improvements and resolve production issues.
- Drive capacity planning, performance tuning, and cost optimization efforts.
- Provide technical leadership and mentorship across the engineering organization.
What you'll need
- 10+ years of software engineering experience with a focus on backend systems and distributed architecture.
- Experience building and operating Java-based systems using RESTful APIs, Spring Boot, and microservices architecture.
- Strong understanding of distributed systems concepts including fault tolerance, eventual consistency, and scalability.
- Proven experience with cloud platforms (AWS/Azure/GCP) and cloud-native architectures.
- Expertise in observability tools such as Prometheus, Grafana, ELK, or similar.
- Ability to define and manage SLIs, SLOs, and error budgets.
- Hands-on experience with incident management, root cause analysis (RCA), and postmortems.
Details
- Location: Bengaluru, India.
- Work style: office-first with three days per week in office for most roles.
Read the full description and apply on the company’s own careers page.