Overview
Senior Software Engineer (MLOps/DevOps) to support and scale Roku’s ML infrastructure for the Advertising Performance team.
What you'll do
- Lead design and operations of scalable cloud infrastructure for ML workloads on AWS and GCP.
- Architect and improve CI/CD systems for ML models and platform services.
- Own low-latency infrastructure for real-time model inference, including KV store and vector databases.
- Define and enforce observability standards for ML systems, including monitoring and drift detection.
- Participate in on-call rotation, leading incident response and root-cause analysis.
- Partner with ML scientists and engineers to improve platform usability and MLOps/SRE practices.
- Champion operational excellence through automation, resilience, and disaster recovery planning.
What you'll need
- BS or MS in Computer Science, Engineering, or a related quantitative field.
- 8+ years of experience in DevOps, SRE, or ML infrastructure, including 4+ years supporting large-scale ML/AI systems.
- Strong programming skills in Python and/or Scala or Java.
- Deep experience with Kubernetes on GCP (GKE) and/or AWS (EKS).
- Expertise with NoSQL or low-latency data stores such as Aerospike.
- Hands-on experience with data and orchestration technologies such as Apache Spark, Apache Flink, Apache Airflow, and Kafka.
- Infrastructure-as-code experience with Terraform or similar tooling.
Details
- Hybrid work: teams work in the office Monday through Thursday; Fridays are generally flexible for remote work.
Read the full description and apply on the company’s own careers page.