Overview
Own end-to-end production incident response, driving detection through investigation, mitigation, resolution, recovery, and closure for critical incidents.
What you'll do
- Take end-to-end ownership of production incidents from detection through closure.
- Serve as primary point of contact and single point of accountability for assigned incidents.
- Lead P1/P2 and Sev0/Sev1 major incidents and coordinate incident bridge/war rooms.
- Perform incident triage and determine incident severity based on customer, business, and platform impact.
- Analyze failures and review logs, monitoring dashboards, alerts, metrics, and dependencies to investigate incidents.
- Drive root cause analysis and post-incident reviews, ensuring corrective and preventive actions are tracked to closure.
- Monitor production systems and work to improve alert quality, detection coverage, dashboards, and automation.
What you'll need
- 3+ years of experience in incident/major incident management, production support, technical operations, NOC, SRE, application support, or IT service management.
- Hands-on experience managing critical production incidents (P1/P2, Sev0/Sev1 or equivalent).
- Strong understanding of incident management, problem management, change management, ITIL, and ITSM processes.
- Experience working in SaaS, cloud, enterprise software, or large-scale production environments.
- Good understanding of application architecture, service dependencies, APIs, microservices, databases, distributed systems, and cloud technologies.
- Working knowledge of AWS, Azure, or GCP.
- Experience with monitoring/observability tools (e.g., Grafana, Prometheus, Kibana, Elasticsearch, Splunk, Dynatrace, AppDynamics or similar).
Details
- Work in a 24×7 operations environment.
Read the full description and apply on the company’s own careers page.