Overview
Senior Staff Site Reliability Engineer to lead infrastructure reliability and efficiency initiatives across on-prem and cloud systems.
What you'll do
- Lead initiatives to transform IT Compute Core Team architecture for new on-prem and cloud service offerings.
- Design, scale, and deploy core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP.
- Build for performance and reliability at global scale using automation, monitoring, high availability, capacity planning, and lifecycle management.
- Define and implement service efficiency metrics and optimize using software/hardware approaches (SR-IOV/DPU).
- Use eBPF and XDP for observability and DDoS mitigation.
- Collect/review system data for capacity planning, analyze capacity, and coordinate change implementation.
- Develop and maintain tools for data collection, analysis, and visualization for reporting/alerting/monitoring.
What you'll need
- Bachelor’s degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
- 15+ years of proven experience in compute platform engineering with a focus on automation.
- Experience designing and deploying containerization architectures and distributed systems infrastructure.
- Strong analytical skills to define and track key performance metrics.
- Experience developing tools for data analysis and performance profiling.
- Development with Terraform and configuration management tools.
- Proficiency in Go and/or Python and Linux OS/kernel internals.
- Understanding of network protocols and architectures (VLAN/VxLAN/SDN/BGP/Anycast).
Details
- Location: Bengaluru, India.
- Work mode: Hybrid (#LI-Hybrid).
Read the full description and apply on the company’s own careers page.