Overview
As an AI Platform and LLMOps Engineer, you will build and maintain infrastructure and tooling for AI adoption and customer-facing AI-powered solutions. You will own the operational backbone that lets agentic AI systems run reliably, observably, safely, and cost-effectively in production.
What you'll do
- Build and maintain operational tooling for LLM-powered systems, including prompt and pipeline versioning, evaluation harnesses, and CI/CD workflows tailored to non-deterministic AI outputs.
- Implement monitoring and observability for production LLM systems, tracking latency, token usage, cost per request, output quality, and drift over time.
- Design and run automated evaluation suites to catch regressions, hallucinations, and quality degradation before they reach customers.
- Manage RAG pipelines and vector store infrastructure, keeping retrieval sources fresh, accurate, and performant.
- Implement safety and compliance guardrails, including content filtering, PII redaction, and access controls, in line with enterprise data privacy and residency requirements.
- Own cost governance for LLM usage through caching strategies, model routing, and usage reporting to keep spend predictable as adoption scales.
- Collaborate with Product Managers and Senior Engineers to scope operational requirements and translate them into reliable systems.
- Participate in sprint planning, technical grooming, and retrospective discussions.
What you'll need
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering or a related field.
- 4+ years of experience in a software engineering role, with a solid background in Backend Engineering and DevOps fundamentals.
- Hands-on experience operating or supporting LLM-powered systems in production via APIs, RAG pipelines, frontier/open weight/fine-tuned models.
- Proficiency in a backend language such as Python, Java, TypeScript or Go.
- Familiarity with Docker and container orchestration, including Kubernetes.
- Experience implementing APIs, working with SQL and NoSQL databases, and writing automated tests.
- Working knowledge of at least one cloud platform, AWS, GCP, or Azure, across both managed and self-hosted services.
- Familiarity with infrastructure-as-code, such as Terraform, and CI/CD tooling, such as GitHub Actions or Jenkins.
- Comfort debugging production issues involving non-deterministic systems, with attention to detail and a bias towards reliability.
- Strong belief in engineering quality and building tooling that creates leverage for others.
Nice to have
- Experience with LLMOps-specific tooling such as LangFuse, LangSmith, Weights & Biases-style evaluation frameworks, or vector databases like Pinecone, Weaviate, pgvector.
- Working knowledge of RAG architectures, MCP implementation & governance, agent orchestration frameworks like LangGraph and LLM Gateways such as LiteLLM, Open Router, and enterprise AI ecosystems such as Vertex AI or Bedrock.
- Exposure to prompt management and versioning practices treated as code, along with model access management, deterministic guardrails, and evals.
- Understanding of AI governance considerations, including data privacy, residency, and compliance in enterprise AI deployments.
- Ability to work in a fast-paced, ambiguous environment with a proactive, ownership-driven mindset.
- Strong communication skills and comfort collaborating across engineering, product, and strategy functions.
Details
- Job location: Pune, India.
Read the full description and apply on the company’s own careers page.