Overview
Staff Engineer for Equinix’s Reliability Engineering team operating a bare-metal Kubernetes platform across multiple metro locations.
What you'll do
- Provision bare-metal servers across metros using PXE/iPXE imaging and cloud-init, and harden the OS prior to cluster bootstrap.
- Bring up and operate Kubernetes clusters, including control-plane topology, etcd members, and worker node setup with CNI.
- Run routine cluster operations like node drains/replacements, certificate rotation, capacity checks, and alert remediation.
- Manage cluster and fleet configuration changes via GitOps to keep clusters consistent and drift-free.
- Execute rolling OS and Kubernetes upgrades across the fleet from tested runbooks and run restore drills.
- Own an on-call shift as first responder for incidents, build dashboards/alerts, and maintain runbooks.
What you'll need
- Hands-on Linux/Ubuntu administration and provisioning for bare metal.
- Familiarity with PXE/iPXE imaging (or equivalent), cloud-init, and OS hardening.
- Working knowledge of BMC/IPMI consoles, SMART data, and platform sensors.
- Production experience operating Kubernetes clusters on bare metal (RKE2 or kubeadm), including etcd basics and CNI configuration.
- Day-to-day use of infrastructure-as-code tools such as Ansible, Terraform, ArgoCD, and Fleet (or similar).
- Working knowledge of VLANs, BGP, bonded NICs, and SR-IOV for network-layer isolation.
- Experience executing OS and Kubernetes upgrades in production against an established runbook.
- Experience in on-call operations responding to production incidents and writing them up afterward.
- Automation mindset to improve runbooks by reducing repeated manual work.
Details
- Location: Bangalore Office BLS2.
- Team: Reliability Engineering; operates a bare-metal Kubernetes platform across multiple metro locations.
Read the full description and apply on the company’s own careers page.