Overview
Build and operate scalable, reliable, and cost-efficient ML inference infrastructure for real-time AI applications in a distributed cloud environment.
What you'll do
- Build and operate production-grade model-serving infrastructure using vLLM, TGI, Triton, or equivalent.
- Implement blue/green and canary deployment pipelines for ML models.
- Develop auto-scaling, multi-model serving, and intelligent request-routing systems.
- Optimize GPU utilization, memory efficiency, network throughput, and model storage performance.
- Design observability for inference latency, throughput, GPU usage, cost, and system health.
- Manage model registries and automated CI/CD model deployments.
- Support the full ML system lifecycle, including production operations and on-call responsibilities.
What you'll need
- 4+ years of experience in MLOps, platform engineering, SRE, or similar ML infrastructure roles.
- Hands-on experience with model-serving frameworks such as vLLM, TGI, or Triton.
- Experience operating containerized GPU workloads in production.
- Experience with model registries, experiment tracking, and automated deployment pipelines.
- Proficiency in Python and infrastructure-as-code tools such as Terraform or Helm.
- Understanding of distributed systems, performance tuning, and production reliability engineering.
- Fluent English and ability to work independently in a remote-first environment.
Nice to have
- Experience with Kubeflow, MLflow, or KubeAI.
- Knowledge of GPU scheduling, CUDA/ROCm optimization, or multi-tenant inference systems.
- Experience optimizing costs across GPU types and inference workloads.
- Background in early-stage startups or greenfield infrastructure projects.
- Experience building production systems from scratch.
Details
- Fully remote in the EMEA timezone.
- Start date: ASAP.
Read the full description and apply on the company’s own careers page.