Overview
Senior AI/ML Engineer role focused on building and operating production-grade AI and Generative AI solutions, including RAG, agentic systems, evaluation, deployment, and observability.
What you'll do
- Own AI/ML and GenAI solutions end to end across data pipelines, prompt/model workflows, APIs, evaluation, deployment, observability, and optimization.
- Design enterprise RAG platforms covering ingestion, chunking, embeddings, hybrid/vector search, reranking, citations, and access-aware retrieval.
- Build production agentic systems with tool/function calling, structured outputs, planning/memory, multi-agent orchestration, approvals, failure recovery, and auditable traces.
- Architect and operate MCP clients and servers with secure transports, authentication, least-privilege access, tenant isolation, and protections against prompt injection/unsafe tool execution.
- Define evaluation strategies and quality gates for accuracy, groundedness, safety, latency, and cost, and establish end-to-end observability.
- Build production Python services using GitHub CI/CD and cloud-native deployment on AWS or Azure with Docker, Kubernetes, and infrastructure as code.
What you'll need
- 5–9 years developing production software, data, ML, or AI solutions, including hands-on GenAI/LLM application delivery.
- Advanced Python with practical experience using data/ML libraries including pandas, NumPy, scikit-learn, PyTorch, or TensorFlow (or equivalent).
- Strong grasp of LLM and agentic architecture including prompting, context engineering, embeddings/RAG, tool/function calling, orchestration, evaluation, and human-in-the-loop controls.
- Proven experience designing APIs, distributed services, and event-driven or asynchronous workflows with secure tool execution for AI agents.
- Hands-on experience with Git/GitHub and CI/CD with automated build, test, security-scan, and deployment workflows.
- Strong experience with AWS or Azure and containerized deployment with Docker; Kubernetes and infrastructure-as-code experience expected.
- A current role-relevant AWS AI/ML certification or Microsoft Azure AI certification is mandatory.
- Hands-on production experience with Model Context Protocol (MCP), including consuming/developing MCP servers and implementing security, approvals, testing, and tracing.
- MLOps/LLMOps practices including experiment tracking, model/prompt versioning, tracing, evaluation, monitoring, and cost optimization.
Details
Read the full description and apply on the company’s own careers page.