Overview
Twilio is hiring a Machine Learning Engineer to build, evaluate, and maintain scalable, low-latency ML systems for real-time applications.
What you'll do
- Analyze business problems with stakeholders to clarify requirements and translate them into measurable ML problem statements.
- Design, implement, and maintain scalable ML solutions in production.
- Build reproducible ML workflows for data preparation, training, evaluation, and inference using orchestration and MLOps tooling.
- Implement monitoring and evaluation frameworks to improve data quality, model performance, latency, and cost.
- Collaborate with cross-functional teams to deliver resilient, scalable, and compliant ML-powered services.
- Own operational excellence including SLAs, on-call, incident response, customer feedback triage, and post-mortems.
- Drive engineering excellence through AI-assisted SDLC practices, code reviews, automated testing, and mentoring.
What you'll need
- 5+ years of experience building, deploying, and operating data and ML systems in production.
- Strong foundation in ML/AI using statistics, probability, and optimization for real-world problems.
- Proficiency in Python, Java, and SQL.
- Experience with workflow orchestration and data pipelines (e.g., Airflow, Kubeflow).
- Experience with ML lifecycle and MLOps tooling (e.g., MLflow, Metaflow, SageMaker; evaluation/observability tools).
- Working knowledge of containerization and cloud infrastructure including Docker and Kubernetes, plus a major cloud platform (AWS, GCP, or Azure).
- Understanding of distributed computing and streaming frameworks (e.g., Spark/EMR, Flink, Kafka Streams).
Details
- Location: Remote in the US, with occasional ad-hoc in-person team gatherings, off-sites, or customer meetings.
- Not eligible for hire in CA, CT, NJ, NY, PA, or WA.
- Travel may be required occasionally for in-person meetings.
Read the full description and apply on the company’s own careers page.