Overview
Software Engineer II (Platform & Infrastructure) for Abnormal AI’s PI team, focused on building and evolving shared observability infrastructure.
What you'll do
- Own and evolve monitoring, metrics, and alerting infrastructure across the Prometheus, Chronosphere, and Grafana stack.
- Build dashboards and manage metric pipelines at scale.
- Operate the PagerDuty alerting pipeline and drive cost-efficient observability across production environments (US, EU, GovCloud).
- Design platform features and developer tooling to reduce friction and deployment time.
- Drive SLAs and SLOs for critical shared infrastructure to improve resilience and cost efficiency.
- Participate in on-call rotations to triage, diagnose, and resolve production issues independently.
- Convert runbooks into automated solutions and improve system resilience through SLAs/SLOs and proactive bottleneck identification.
What you'll need
- 4+ years of hands-on backend engineering experience designing, building, and operating production-grade distributed systems.
- Strong proficiency in Python (for Airflow DAGs, platform services, and automation tooling).
- Working proficiency in Golang (for high-performance infrastructure components, metric pipelines, and platform services).
- Experience owning a service/platform end-to-end, from design through deployment, monitoring, and iteration.
- Ability to write technical design documents that articulate trade-offs and propose solutions for buy-in.
- Proven incident response capability from on-call experience diagnosing and resolving production issues.
- Strong testing discipline including unit and integration testing, with knowledge of what to test.
Details
- Location: Hybrid in Bangalore, India.
- Must Haves include 4+ years of backend engineering and distributed systems experience.
Read the full description and apply on the company’s own careers page.