Overview
Operate and improve enterprise AI platforms, cloud infrastructure, and production AI workloads across Azure and other model providers. The role covers platform reliability, deployment automation, security, observability, and developer self-service.
What you'll do
- Operate Azure AI platforms and application hosting across development, test, and production.
- Deploy and support copilots, assistants, chatbots, agents, RAG services, and machine learning applications.
- Manage containerized AI workloads, Kubernetes environments, cluster health, scaling, networking, and recovery readiness.
- Operate model-access gateways with routing, failover, entitlements, quotas, monitoring, and provider onboarding.
- Build monitoring, logging, alerting, runbooks, incident processes, and reliability improvements.
- Develop Terraform infrastructure, CI/CD pipelines, release controls, and automated operational checks.
- Implement secure networking, identity, secrets management, privileged access, and platform hardening.
What you'll need
- 8+ years in cloud platform engineering, infrastructure engineering, DevOps, platform operations, SRE, MLOps, or AI operations.
- Bachelor's degree in a relevant STEM discipline or equivalent practical experience.
- Hands-on Azure production operations, including Azure AI services, App Service, Functions, and container-based compute.
- Experience operating AI, machine learning, or generative AI platforms in controlled production environments.
- Experience with Docker, Kubernetes, cloud networking, private endpoints, DNS, NSGs, and secure architecture.
- Terraform, Azure DevOps or GitHub Actions, Python, and Bash or PowerShell skills.
- Knowledge of observability, incident management, AI/ML operations, security, governance, and compliance.
Nice to have
- Azure Kubernetes Service operations and production troubleshooting experience.
- Multi-provider model access using AWS Bedrock, Google Vertex AI, or gateway technologies such as LiteLLM.
- Experience with API Management, Helm or Kustomize, GitOps, Prometheus, or Azure Managed Grafana.
- Experience building internal developer platforms or self-service engineering capabilities.
Details
- Location: Bangalore, India.
Read the full description and apply on the company’s own careers page.