Overview
Staff ML Software Engineer responsible for platform subsystems supporting large-scale AI/ML workflows, with ownership spanning observability, cost efficiency, reliability, and developer experience.
What you'll do
- Design and operate observability, evaluation, and tooling subsystems for next-generation ML architecture.
- Build systems for anomaly detection, root cause analysis, and operational automation.
- Develop visibility into model behavior, training health, serving latency, and data quality.
- Drive cost optimization across training and serving infrastructure.
- Architect reliability improvements that reduce operational toil and improve on-call workflows.
- Coordinate the target architecture and migration path for the modernized AI/ML stack.
- Evaluate emerging infrastructure patterns and translate them into a technology roadmap.
What you'll need
- Significant experience designing, building, and operating production AI/ML systems at scale.
- Experience with training pipelines, model serving, and high-traffic online inference.
- Hands-on experience building orchestration or control subsystems for advanced agentic architectures.
- Deep Python expertise and working proficiency in Scala or Java.
- Experience improving AI/ML reliability, infrastructure costs, and operational scalability.
- Strong distributed systems background, including batch processing and real-time serving.
- Experience driving technical programs across teams without formal authority.
Nice to have
- Familiarity with LLM evaluation, trace, replay, observability, or debugging tools.
- Familiarity with feature stores, model serving platforms, and experiment frameworks.
- Experience migrating production AI/ML systems across technology generations.
- Experience in recommendation, search, or discovery systems.
Details
- Remote in the United States.
- Full-time role.
- Annual salary range: $600,000–$1,066,000.
Read the full description and apply on the company’s own careers page.