Overview
Site Reliability Engineer role focused on reliability, observability, performance, and operational excellence for application development.
What you'll do
- Own SRE across .NET and React applications, focusing on reliability and observability.
- Define and track SLOs, SLIs, and error budgets, partnering with development teams to respond when budgets are at risk.
- Design and maintain monitoring, alerting, logging, and observability solutions (e.g., Azure Monitor, Application Insights, Log Analytics, Grafana).
- Lead incident response, root cause analysis, and post-incident reviews with blameless postmortems and corrective follow-through.
- Use AI tools (e.g., GitHub Copilot and Copilot in Azure) to accelerate incident triaging and root cause analysis.
- Contribute to internal AI-driven tooling (e.g., chatops bots, runbook copilots, log-analysis assistants) to reduce engineering toil.
What you'll need
- Minimum 3 years of experience troubleshooting and debugging C# .NET, React, MS SQL, and APIs.
- Hands-on development experience with C# .NET and React is an added advantage.
- Required experience with Azure platform services (App Service, Functions, Storage, Queues, Key Vault, API Management, Application Insights, Log Analytics).
- Required experience with git source control, GitHub and Azure DevOps including CI/CD pipelines and GitHub Actions.
- Required understanding of deployment processes including blue/green, canary, and feature-flag-based release patterns.
- Experience reviewing and analyzing Infrastructure as Code (Bicep, ARM, or Terraform) is required.
- Experience working with SRE principles covering SLOs/SLIs/error budgets, incident management, and blameless postmortems.
Details
- Location: Mumbai - Lower Parel.
Read the full description and apply on the company’s own careers page.