Overview
Director of Production Engineering (Remote) to lead DevOps/SRE teams for the availability, scalability, and security of a production environment.
What you'll do
- Hire and lead a globally distributed DevOps/SRE engineering team.
- Own the reliability and infrastructure roadmap for AWS-based production (EKS, RDS, related services).
- Run SecOps activities including vulnerability management, threat detection, incident response, and remediation.
- Define and track engineering OKRs for reliability, automation, and security outcomes.
- Lead observability and alerting best practices and automating alert triage/response to reduce MTTR.
- Drive Infrastructure-as-Code and CI/CD automation to reduce operational toil and improve velocity.
- Participate in an incident management on-call rotation to meet availability goals.
What you'll need
- 8-12 years of experience in DevOps, Site Reliability Engineering, or production infrastructure, including people management.
- Hands-on experience running production workloads on AWS including EKS and RDS.
- Demonstrated experience in security operations (SecOps) such as vulnerability management and incident response.
- 5+ years using observability platforms (e.g., Datadog, Prometheus, Grafana).
- Experience with Infrastructure-as-Code (e.g., Terraform, CloudFormation) and CI/CD automation.
- Proficiency in Go, Python, or Bash; day-to-day use of Git and test automation pipelines.
- Bachelor’s degree in Computer Science/Engineering/related field required (Master’s preferred).
Details
- Work mode: Remote; location: United States.
Read the full description and apply on the company’s own careers page.