Overview
Senior Machine Learning Platform Engineer to design, build, and scale enterprise machine-learning and generative-AI platform capabilities for teams across model development to production operations.
What you'll do
- Design and build reusable ML and GenAI platform capabilities for experimentation, evaluation, deployment, and production operations.
- Create self-service platform services, APIs, and automation to abstract infrastructure complexity for developers.
- Build and maintain MLOps capabilities including experiment tracking, model/prompt registries, evaluation frameworks, and deployment workflows.
- Develop model-serving and inference capabilities for classical ML models, deep-learning models, and LLMs via REST/gRPC/event-driven interfaces.
- Implement GenAI/agentic platform capabilities such as prompt management, embeddings, vector search, RAG, tool integration, and agent frameworks.
- Ensure observability, operational monitoring, AI evaluation/quality gates, and security/governance controls within platform capabilities.
What you'll need
- 3–5 years of experience in machine learning engineering, ML platform engineering, MLOps, backend engineering, cloud engineering, or enterprise AI systems.
- Strong software-engineering fundamentals building production-grade distributed services, APIs, and reusable libraries.
- Strong programming skills in Python; Java or another enterprise language is preferred.
- Hands-on experience with Docker and Kubernetes (or equivalent containerization/orchestration).
- Experience designing REST APIs, microservices, and backend platform services.
- Experience with MLOps/GenAIOps concepts including experiment tracking, model lifecycle management, evaluation, deployment, monitoring, and reproducibility.
- At least one major cloud platform experience (AWS, Azure, or GCP).
Nice to have
- Experience building internal developer platforms, ML platforms, or self-service engineering platforms.
- Experience with Infrastructure as Code such as Terraform.
- Experience designing multi-tenant platform services with resource isolation, quotas, governance, and usage attribution.
- Familiarity with event-driven architectures, message queues, and asynchronous processing.
Details
- Location: Hyderabad, India.
Read the full description and apply on the company’s own careers page.