Overview
Staff Engineer – L4 serving as a senior escalation point for complex, production-impacting issues in enterprise and hyperscale environments. The role supports the Infinia Core engineering team through root-cause analysis, incident leadership, product improvement, and AI-assisted diagnostics.
What you'll do
- Own critical customer escalations through root-cause analysis and mitigation.
- Lead incident war rooms and cross-functional response efforts.
- Debug logs and system patterns using AI-powered tools.
- Reproduce customer issues and propose product improvements or workarounds.
- Author runbooks, performance guides, and RCA documentation.
- Partner with Field CTOs, solution architects, and sales engineers.
- Deliver training to customer support and field engineering teams.
What you'll need
- 8+ years in enterprise storage, distributed systems, or cloud infrastructure support or engineering.
- Deep knowledge of S3, POSIX, NFS, storage performance, and Linux kernel internals.
- Scripting or coding experience with Python, Go, or C++.
- System, protocol, and application debugging experience using tools such as strace, tcpdump, or perf.
- Hands-on Linux troubleshooting experience.
- Experience with AI tools for diagnostics and reducing MTTR.
- Strong communication and executive reporting skills.
Nice to have
- Experience with DDN, VAST, Weka, or similar scale-out file systems.
- Familiarity with Prometheus, Grafana, ELK, or OpenTelemetry.
- Knowledge of replication, consistency models, and data integrity mechanisms.
- Exposure to RDMA, NVMe-oF, or high-performance networking stacks.
- Experience with Sovereign AI, LLM training environments, or autonomous system data architectures.
Details
- Location: Pune Office.
- Participate in an on-call rotation for after-hours support as needed.
Read the full description and apply on the company’s own careers page.