Overview
Organize the deployment of AI workflows and architectures across diverse AI and technology platforms. Facilitate solution design across the enterprise technology stack to integrate with AI models and multi-provider platforms.
What you'll do
- Design, build, and operate enterprise observability platforms using Splunk, providing unified metrics, logging, and tracing across infrastructure tiers with OpenTelemetry instrumentation standards.
- Engineer and govern the infrastructure-as-code platform using Terraform or Pulumi, including module libraries, state management, policy-as-code, drift detection, and multi-cloud coverage across AWS and GCP.
- Build and maintain CI/CD and GitOps pipeline infrastructure using GitHub Actions and ArgoCD for infrastructure and application platform delivery.
- Operate the Internal Developer Portal, Backstage, for self-service provisioning.
- Own ITSM platform engineering, including ServiceNow workflow automation, CMDB reconciliation, PagerDuty/OpsGenie integrations, and automated incident creation pipelines across infrastructure towers.
- Engineer LLM-powered runbook automation integrating OpenAI, Anthropic Claude, or Google Gemini with infrastructure tooling to automate incident triage, resolution workflows, and knowledge base queries.
- Build agentic ITSM workflows using multi-agent orchestration to handle incident classification, runbook execution, and stakeholder communication with human-in-the-loop controls.
- Implement AI-powered alert correlation and noise reduction to group related alerts, suppress false positives, and surface actionable incident context to on-call engineers.
- Design RAG architectures for infrastructure knowledge bases backed by vector databases.
- Engineer secret management infrastructure using HashiCorp Vault, including dynamic secrets, PKI automation, and Kubernetes authentication across multi-cloud and on-premises environments.
- Mentor engineers across infrastructure towers, set platform engineering standards, drive tooling roadmaps, and act as the highest technical escalation point for platform and AI tooling issues.
What you'll need
- Minimum 7.5 years of experience in Agentic Orchestration.
- 15+ years in platform, tools, or infrastructure engineering.
- 15 years of full-time education.
- Mandatory certifications: Terraform Associate or Professional, Splunk Professional, AWS DevOps Engineer Pro or GCP DevOps Engineer, Certified Kubernetes Administrator (CKA), and HashiCorp Vault Associate.
- Experience with Splunk SLO engineering, OpenTelemetry, ELK/Splunk log pipelines, and distributed tracing.
- Experience with advanced Terraform modular design, policy-as-code, drift detection, and Ansible/Chef for configuration management.
- Experience with GitHub Actions, ArgoCD pipeline engineering, GitOps workflows, policy enforcement, and infrastructure deployment automation.
- Experience with ServiceNow workflow automation, CMDB integration, incident/problem/change API integration, and PagerDuty/OpsGenie connectivity.
- Experience with service catalogues, golden path templates, self-service provisioning, and plugin development for an Internal Developer Portal.
- Experience with HashiCorp Vault dynamic secrets, PKI automation, Kubernetes authentication, and multi-cloud secrets lifecycle.
- Experience with OpenAI, Anthropic Claude, or Google Gemini prompt engineering, RAG architecture, vector databases, and LLMOps.
- Experience with LangChain, LlamaIndex, or CrewAI agentic workflow design, multi-agent orchestration, tool calling, and human-in-the-loop workflows.
- Advanced Python experience for platform tooling development, LLM SDK integration, infrastructure automation, and API development.
- Experience with PCI DSS, SOX, and DORA platform access controls, audit logging, policy-as-code enforcement, and regulatory evidence production.
Nice to have
- Experience with Crossplane for Kubernetes-native cloud resource provisioning.
- Experience with eBPF-based observability, including Cilium, Pixie, or Hubble for deep infrastructure telemetry.
- Experience with AI/ML infrastructure, GPU node management, model serving platforms such as Triton or vLLM, or MLflow.
- Platform engineering or SRE consulting background within financial services.
Details
- Based at the Bengaluru office.
Read the full description and apply on the company’s own careers page.