Overview
Own the end-to-end operational lifecycle of datacenter servers, from provisioning and deployment through steady-state operation, maintenance, repair, and decommissioning. Help define automation, tooling, and trusted compute standards for large-scale, secure fleet operations.
What you'll do
- Build automation to support datacenter fleets at scale.
- Define and own the end-to-end system lifecycle strategy, including refresh and decommissioning.
- Maintain automation and operational procedures for lifecycle events like hardware failures and firmware upgrades.
- Partner with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle.
- Work with Networking to ensure end-to-end connectivity across all sites.
- Build and maintain tooling to track machine health, configuration, and operational status across the fleet.
What you'll need
- Hands-on experience with server hardware, including rack deployment, cabling, troubleshooting, and failure modes.
- End-to-end understanding of hardware lifecycle management, including asset tracking, provisioning, maintenance scheduling, and decommissioning.
- Proficiency in at least one programming language (Python, Rust, Go, or Java).
- Working knowledge of modern cloud infrastructure, including Kubernetes and large cloud providers (AWS, Azure, GCP).
- Ability to communicate clearly and build consensus with stakeholders.
- Comfort navigating ambiguity and progressing on complex, cross-functional problems.
Details
- Location-based hybrid policy: expected to be in one of the company offices at least 25% of the time.
- Visa sponsorship: sponsored and efforts made for offered candidates.
- Minimum education: Bachelor’s degree or equivalent education/training/experience.
- Required field of study: field relevant to the role as demonstrated through coursework, training, or experience.
Read the full description and apply on the company’s own careers page.