Overview
AI Platform Engineer to operate, automate, and improve production AI services, focusing on reliable deployment and operations.
What you'll do
- Build and operate platforms for deploying, monitoring, and maintaining production AI models and services.
- Support model-serving environments, inference pipelines, deployment workflows, and release automation.
- Develop automation for model packaging, promotion, validation, rollout, rollback, and lifecycle management.
- Implement monitoring, logging, alerting, and observability for AI service performance, latency, availability, error rates, and cost.
- Troubleshoot model serving, API behavior, environment configuration, infrastructure, and deployment pipelines.
- Partner with teams to move AI workflows and services from prototype to production.
What you'll need
- Experience in MLOps, AI platform engineering, DevOps, software engineering, or production ML operations.
- Experience deploying or operating AI/ML models and inference services (or API-based production systems).
- Strong scripting or programming skills using Python or similar languages.
- Experience with containers, CI/CD, cloud environments, monitoring, and operational automation.
- Understanding of model deployment concepts, lifecycle management, versioning, validation, and rollback.
- Experience troubleshooting production systems across application, model, infrastructure, and deployment layers.
Nice to have
- Experience with model-serving frameworks and MLOps tools (e.g., Kubernetes, Docker, ECS, EKS, MLflow, KServe, BentoML, Ray, Airflow).
- Experience supporting LLMs/agentic workflows/RAG systems or other applied AI services (e.g., speech, translation).
- Experience with AWS services for compute, storage, networking, security, monitoring, and deployment.
- Experience evaluating or integrating commercial and open-source AI platforms.
Details
- Location in the source ATS line: DE, Germany, Home Office.
Read the full description and apply on the company’s own careers page.