Overview
Staff+ Infrastructure Engineer for Cluster Infrastructure, owning the technical strategy for agent-driven compute cluster lifecycle management.
What you'll do
- Own the technical strategy and roadmap for agent-driven cluster lifecycle management (provisioning, updates, decommissioning).
- Partner across teams to ensure new compute capacity is ingested on time.
- Collaborate on physical build-out and cloud solutions for high-bandwidth inter-cluster connectivity.
- Work with security owners to provision clusters secure-by-default.
- Define strategy for cluster scalability, homogeneity, and fault tolerance.
- Drive operational-excellence practices including incident response, postmortems, and on-call health.
- Provide technical mentorship and coaching to engineers.
What you'll need
- Deep expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, AWS/GCP/Azure).
- Strong proficiency in at least one systems language (Rust, Go, or Python).
- IaC proficiency with Terraform.
- Track record leading complex multi-quarter technical initiatives across multiple teams/systems.
- Ability to build alignment across senior stakeholders and communicate at all levels.
Details
- Location-based hybrid policy: staff expected in an office at least 25% of the time.
- Visa sponsorship: offered, with reasonable efforts if an offer is made.
Read the full description and apply on the company’s own careers page.