Overview
Build and operate scalable, reliable ML inference infrastructure for real-time AI applications in a remote ML Ops Engineer role.
What you'll do
- Build and operate production-grade model-serving infrastructure.
- Implement blue/green and canary deployment pipelines for ML models.
- Develop auto-scaling, multi-model serving, and request-routing systems.
- Optimize GPU utilization, memory, network throughput, and artifact storage.
- Design observability for inference latency, throughput, GPU usage, cost, and system health.
- Manage model registries and CI/CD pipelines for reproducible deployments.
- Own ML system operations, including production support and on-call responsibilities.
What you'll need
- 4+ years in ML Ops, Platform Engineering, SRE, or similar ML infrastructure roles.
- Hands-on experience with vLLM, TGI, Triton, or equivalent model-serving frameworks.
- Experience operating containerized GPU workloads in production.
- Experience with model registries, experiment tracking, and automated deployment pipelines.
- Proficiency in Python and infrastructure-as-code tools such as Terraform or Helm.
- Understanding of distributed systems, performance tuning, and reliability engineering.
- Ability to work independently in a remote-first environment.
Nice to have
- Experience with Kubeflow, MLflow, or KubeAI.
- Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference.
- Experience optimizing costs across GPU types and inference workloads.
- Background in early-stage startups or greenfield infrastructure projects.
- Experience building production systems from scratch.
Details
- Location: Ukraine.
- Work mode: Fully remote in the EMEA timezone.
- Fluent English required.
- Start date: ASAP.
Read the full description and apply on the company’s own careers page.