Overview
AI/MLOps SRE Lead Engineer to drive reliability, scalability, observability, and operational excellence across an AI/ML and cloud ecosystem.
What you'll do
- Drive service reliability, availability, and performance across multi-cloud environments using SLOs, SLIs, and error budgets.
- Design, build, and scale enterprise ML platform infrastructure across technologies including Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
- Develop AI-driven observability with anomaly detection, predictive analytics, and automated remediation.
- Implement and monitor LLM, SLM, RAG, and AI Agent platforms for performance, governance, operational efficiency, and scalability.
- Build Infrastructure as Code, CI/CD pipelines, and platform automation for operational resilience.
- Architect enterprise ChatOps solutions integrating operational events, AI workflows, observability, and automated remediation.
What you'll need
- 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology fields with enterprise-scale delivery.
- Hands-on experience operating across two or more major cloud platforms including AWS, GCP, or Azure.
- Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
- Proven end-to-end ML workflow experience across model training, deployment, experiment tracking, monitoring, and pipeline orchestration.
- Advanced proficiency in Infrastructure as Code tools such as Terraform, Pulumi, and AWS CDK plus modern CI/CD automation.
- Strong programming/scripting skills in Python, Go, or Bash.
- Experience building enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, and related tracing/metrics/logging.
Details
- Location: Hyderabad (Hybrid).
Read the full description and apply on the company’s own careers page.