Mastercard logo

Lead AI Ops Engineer

Mastercard
NewPosted today

LOCATION

Gurgaon · Onsite

EXPERIENCE

Not specified

TYPE

FullTime

SALARY

Negotiable

SKILLS REQUIRED

ObservabilityIncident responseRoot Cause AnalysisMachine LearningData GovernanceAPI Integration

Job description

Overview

The Lead AI Ops Engineer owns the production operation, monitoring, reliability, evaluation, and governance of one or more AI agents as they learn and change in production.

What you'll do

  • Own end-to-end production monitoring for one or more AI agents, tracking key metrics, KPIs, accuracy, latency, SLA breaches, guardrail triggers, and out-of-scope rates through dedicated observability dashboards.
  • Serve as L1/L2 production support by providing first-response monitoring, triage, and escalation to tech teams as new agentic services go live.
  • Identify and diagnose issues by monitoring evaluation scores and drift alerts, inspecting failing traces and patterns, and classifying root causes such as intent, tool, parameter, or hallucination errors.
  • Evaluate whether agent reasoning remains correct in production, context retrieval stays current, resolutions match ground truth, and the learning loop is drifting.
  • Monitor and manage post-production reliability, including API and authentication failures and issues originating from external AI endpoints, to protect revenue-generating, client-facing agents.
  • Partner with AI Engineering to prioritize fixes, run experiments on prompts, context, and routing in Dev, and validate improvements via evaluations before handing off proven changes for deployment.
  • Embed governance within the agent’s learning loop so approval workflows and audit logging travel with every production change.
  • Abide by Mastercard’s security policies and practices.
  • Ensure the confidentiality and integrity of the information being accessed.
  • Report any suspected information security violation or breach.
  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.

What you'll need

  • Experience monitoring or operating production AI/ML or agentic systems, including observability, evaluation, and drift detection practices.
  • Strong analytical skills to diagnose root causes across intent classification, tool selection, parameter errors, and hallucination patterns.
  • Understanding of AI/LLM observability concepts such as tracing, telemetry, and guardrails.
  • Understanding of evaluation engineering, including benchmarks, automated evaluations, human review, and regression testing.
  • Ability to work independently alongside existing Tech teams while owning a dedicated agent’s post-launch health.
  • Excellent communication skills to translate production signals into clear, prioritized asks for AI Product and AI Engineering teams.
  • Comfort working in a fast-evolving environment where the process itself, not just the throughput, is continuously operated on and improved.

Nice to have

  • Familiarity with governance, compliance, and responsible AI practices such as PII detection, access control, audit logging, and policy enforcement.

Details

  • Location: Gurgaon, India.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Lead AI Ops Engineer