T

Sr. Staff AI Architect

Thermo Fisher Scientific
NewPosted today

LOCATION

Bangalore · Hybrid

EXPERIENCE

15+ Years

TYPE

FullTime

SALARY

Negotiable

SKILLS REQUIRED

System DesignMachine LearningAPI IntegrationObservabilityRetrieval-Augmented Generation

Job description

Overview

The Senior Staff AI Architect provides enterprise-wide architectural, design, and technical leadership for Generative AI and agentic AI solutions across multiple Scrum teams. The role defines architecture, standards, and the roadmap for an enterprise agentic harness and runtime platform, while contributing hands-on to production-grade agentic systems, cloud-native platforms, RAG architectures, LLM integrations, and AI engineering practices.

What you'll do

  • Provide enterprise architecture leadership for Generative AI and agentic AI platforms across multiple teams.
  • Define the architecture and roadmap for a reusable agentic harness and runtime platform.
  • Own high-level and low-level system design, including component architecture, data flows, integration patterns, deployment topologies, and runtime interactions.
  • Design and evolve cloud-native, event-driven, API-first, and distributed architectures for AI-enabled products.
  • Define reference architectures, design standards, and best practices for agentic systems, multi-agent orchestration, RAG and knowledge-grounded agents, tool and function execution, agent memory and state management, human-in-the-loop workflows, and Agent-to-Agent communication.
  • Define architecture patterns for agent planning, task decomposition, reasoning, execution, retries, timeouts, compensation, and failure recovery.
  • Establish standards for agent creation, versioning, testing, evaluation, deployment, monitoring, and retirement.
  • Act as the go-to authority for architecture, design trade-offs, scalability decisions, and complex implementation challenges.
  • Ensure architectural decisions address performance, scalability, security, reliability, privacy, governance, cost, and observability.
  • Lead architecture reviews, design reviews, technical investigations, and architecture decision records.
  • Design and implement agentic systems using LangGraph, LangChain, and comparable orchestration frameworks.
  • Architect agent runtime and execution management, planning and task orchestration, tool registries, tool discovery, function calling, API integration, agent memory, context management, workflow state persistence, multi-agent collaboration, human approvals, guardrails, policy enforcement, execution tracing, auditability, and replay.
  • Design model-agnostic architectures supporting Azure OpenAI, Anthropic Claude, OpenAI-compatible APIs, and local models using Ollama.
  • Define model routing, fallback, provider abstraction, latency optimization, token management, and cost-control strategies.
  • Architect and implement RAG solutions covering document ingestion, chunking, embeddings, retrieval, reranking, grounding, and response synthesis.
  • Integrate agents with APIs, enterprise systems, scientific applications, structured data, knowledge graphs, and domain ontologies.
  • Apply prompt engineering, structured outputs, tool-calling, memory patterns, and context optimization techniques.
  • Define safe and reliable patterns for autonomous and semi-autonomous agent execution.
  • Define automated testing and evaluation strategies for agentic and GenAI systems, including task completion, tool-use accuracy, response correctness, grounding and retrieval quality, safety and policy compliance, robustness and recovery, latency, and cost.
  • Design regression testing for prompts, models, tools, workflows, and agent behaviors.
  • Define observability standards using execution traces, agent steps, tool calls, model interactions, latency, token usage, cost, and failure metrics.
  • Establish mechanisms for agent debugging, replay, inspection, and root-cause analysis.
  • Define secure deployment, monitoring, governance, and operational practices for AI systems.
  • Ensure production systems support reliability, scalability, disaster recovery, data protection, and compliance requirements.
  • Contribute to hands-on development using Python and modern backend frameworks such as FastAPI.
  • Build reference implementations and reusable platform components for the agentic harness.
  • Design and build maintainable and extensible APIs supporting AI, agent, and data-driven workloads.
  • Define service boundaries, integration contracts, data models, workflow interfaces, and event schemas.
  • Implement performance, scalability, security, and observability patterns for distributed AI services.
  • Guide the use of PostgreSQL, pgvector, Qdrant, event streaming, caching, and persistent workflow storage.
  • Promote clean code, automated testing, code reviews, CI/CD, and engineering quality standards.
  • Mentor and guide architects, staff engineers, and senior engineers on agentic architecture, distributed systems, system design, GenAI engineering, evaluation, and observability.
  • Influence technical direction across multiple Scrum teams and product areas.
  • Lead Communities of Practice focused on agentic AI, GenAI architecture, and AI engineering excellence.
  • Communicate through documentation, architecture diagrams, design reviews, and technical presentations.
  • Partner with product, security, platform, data, and engineering leadership to align agentic AI capabilities with business and scientific objectives.
  • Anticipate architectural risks and opportunities and guide the organization through technology evolution.

What you'll need

  • Bachelor’s degree in Engineering or Master’s degree in Computer Science with 15+ years of proven industry experience, including significant experience in software architecture, platform engineering, and technical leadership.
  • Proven experience defining architecture and technical strategy across multiple products or engineering teams.
  • 8+ years of experience building scalable backend systems and RESTful APIs using Python and FastAPI.
  • Strong knowledge of microservices, event-driven architecture, asynchronous processing, messaging, caching, and workflow orchestration.
  • Strong experience with API lifecycle management, authentication, authorization, service-to-service communication, and integration patterns.
  • Experience designing scalable and resilient solutions on Azure, AWS, or GCP.
  • Proficiency with Git, automated delivery pipelines, infrastructure automation, and release management.
  • Experience with pytest, unittest, integration testing, contract testing, and automated quality strategies.
  • Demonstrated ability to influence architecture and engineering decisions across multiple teams without relying solely on organizational authority.
  • Deep hands-on experience designing and implementing agentic AI systems and agentic harnesses.
  • Strong experience with LangGraph, LangChain, or comparable agent orchestration frameworks.
  • Practical experience with agent planning and task decomposition, tool and function calling, multi-agent coordination, agent memory and state management, human-in-the-loop workflows, guardrails and policy enforcement, retry and failure handling, and agent execution tracing and replay.
  • Experience integrating LLMs using Azure OpenAI, Anthropic Claude, OpenAI-compatible APIs, and local inference platforms such as Ollama.
  • Strong prompt engineering skills, including structured prompting, few-shot learning, prompt versioning, context management, and output validation.
  • Experience implementing RAG architectures, embeddings, semantic search, hybrid search, reranking, grounding, and citation patterns.
  • Hands-on experience with vector databases such as PostgreSQL with pgvector and Qdrant.
  • Experience with model routing, fallback strategies, token optimization, latency management, and LLM cost control.
  • Experience evaluating nondeterministic AI and agentic systems in production-like environments.
  • Knowledge of evaluation-driven development, golden datasets, regression testing, and quality metrics for agents.
  • Experience implementing observability for prompts, model calls, tool calls, agent steps, workflow state, latency, token usage, and cost.
  • Understanding of AI security, data privacy, access control, responsible AI, auditability, and governance.
  • Experience with LLMOps, model gateways, prompt management, model evaluation, monitoring, and deployment governance.
  • Experience designing and managing data stores and vector indexes supporting GenAI and RAG workloads using PostgreSQL, pgvector, and Qdrant.
  • Strong data engineering skills, including ETL pipelines, data ingestion, preprocessing, metadata management, and large-scale data processing.
  • Experience integrating structured data, knowledge graphs, and domain ontologies with agentic systems.
  • Expertise in Pandas, NumPy, and Python data-processing libraries.
  • 5+ years of experience working in Agile/Scrum or comparable product development environments.
  • Excellent written and verbal communication skills.
  • Ability to explain complex architecture and AI concepts clearly to technical and non-technical stakeholders.
  • Demonstrated experience mentoring senior engineers, architects, and technical leads.

Nice to have

  • Azure preferred, including AKS, managed identity, Azure AI services, eventing, security, and monitoring.
  • Experience with MCP, Agent-to-Agent communication protocols, or comparable tool and agent interoperability standards.
  • Experience with durable workflow engines or distributed orchestration platforms.
  • Familiarity with MLflow, Kubeflow, model gateways, prompt registries, evaluation platforms, and AI governance.
  • Exposure to scikit-learn, PyTorch, or TensorFlow.
  • E

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Sr. Staff AI Architect