Overview
Senior Staff Engineer to build, deploy, operate, and continuously improve a bare-metal Kubernetes platform running across multiple metro locations.
What you'll do
- Provision and harden Debian/Ubuntu on bare-metal nodes using PXE/iPXE imaging and cloud-init.
- Deploy and operate Kubernetes clusters with control-plane fault tolerance and correct failure-domain placement.
- Maintain a Layer 4 load balancer or keepalived VIP in front of kube-apiserver instances.
- Build and own the CI/CD pipeline for deploying the software-defined data plane with throughput/latency gates.
- Operate and troubleshoot the multi-metro network underlay (VLANs, BGP peering, bonding, connectivity, firewall rules, segmentation).
- Plan and execute rolling OS and Kubernetes upgrades with zero unplanned downtime and tested etcd snapshot/restore.
- Administer the centralized management plane (RBAC, registration, and policy enforcement) and manage GitOps-based fleet configuration.
What you'll need
- Hands-on Linux/Ubuntu provisioning at scale using PXE/iPXE imaging and cloud-init.
- Hardware fluency including BMC/IPMI, SMART/platform sensors, and diagnosing disk/memory/NIC/power faults.
- Experience deploying and operating RKE2 or kubeadm clusters on bare metal, including etcd operations and kube-apiserver HA.
- Production experience managing multiple clusters via Rancher Prime/Rancher or comparable fleet tools.
- Infrastructure CI/CD using GitOps or pipeline-driven deployment tools such as Ansible/Terraform/ArgoCD/Fleet.
- Hands-on networking with VLANs, BGP, bonded NICs, SR-IOV, and ability to troubleshoot across layers.
- Proven zero-downtime rolling upgrades of OS and Kubernetes across a multi-node fleet.
Details
- On-call rotation with incident response for control-plane, data-plane, and network events.
- Team operates a bare-metal Kubernetes platform across multiple metro locations.
Read the full description and apply on the company’s own careers page.