F

Fullstack AI Quality Engineer – Applied & Agentic AI Systems

F. Hoffmann-La Roche Ltd
NewPosted today

LOCATION

Hyderabad · Onsite

EXPERIENCE

7+ Years

TYPE

FullTime

SALARY

Negotiable

SKILLS REQUIRED

Security TestingLLM EvaluationTest AutomationQuality Assurance (QA)Prompt Injection Testing

Job description

Overview

Roche is seeking a Fullstack AI Quality Engineer to test Generative and non-Generative AI applications and agentic systems. The role covers manual exploratory testing, automated UI/API/MCP testing, classical machine learning validation, and generative and agentic AI evaluation within a lean software team.

What you'll do

  • Design and implement testing strategies for LLM-based applications and agentic systems.
  • Evaluate model responses, RAG pipeline accuracy, system reliability, autonomous multi-agent behaviors, and tool-use accuracy.
  • Design and implement automated integration, performance, and AI behavior tests across Python backend services and TypeScript/React/Angular frontends.
  • Use, co-create, and maintain testing frameworks covering traditional software quality and AI-specific concerns such as output consistency, contextual accuracy, and ethical compliance.
  • Create and maintain test strategies, test cases, testing guidelines, and comprehensive test documentation for AI applications.
  • Identify and test edge cases involving ambiguous inputs, potential biases, and unexpected response patterns.
  • Design and execute tests for bias and fair treatment across different user groups and use cases.
  • Test LLM application vulnerabilities, including prompt injection, data leakage, and other AI-specific security concerns.
  • Conduct performance testing, including response time analysis, load testing, and resource utilization monitoring.
  • Stay current with AI and GenAI testing methodologies and best practices.
  • Clearly document findings and recommendations and communicate them in English at the level of C1+.

What you'll need

  • B.Sc., B.Eng., M.Sc., M.Eng. in Computer Science, Software Engineering, or equivalent degree with a strong background in testing methodologies.
  • 7+ years of experience in software testing.
  • At least 1 year focused on genAI and/or agentic applications.
  • Experience working in regulated environments with qualified infrastructure and validated applications.
  • Understanding of the difference between GxP and non-GxP products and the importance of CSV and ITSM processes in delivery.
  • Experience working with Agile development teams.
  • Passion for quality assurance and testing, particularly for AI and GenAI applications and agents.
  • Strong analytical and problem-solving skills with attention to detail.
  • Critical thinking and problem-solving when using AI prompts, including in NLP, data distributions, and edge cases.
  • Proactivity in identifying potential issues and suggesting improvements.
  • Advanced command of Python and standard testing frameworks such as pytest and allure, alongside practical familiarity with specialized AI evaluation tools.
  • High competence in TypeScript-driven test frameworks such as Playwright, Jest, and Cypress.
  • Proven expertise architecting and maintaining end-to-end automated testing suites bridging UI/API validation with asynchronous AI/LLM evaluation pipelines.
  • Hands-on experience integrating test suites within GitHub Actions or modern CI/CD pipelines.
  • Expertise implementing evaluation harnesses such as RAGAS, DeepEval, and Promptfoo.
  • Ability to test autonomous agentic loops, including tool-use accuracy, step-by-step reasoning verification, and goal-directed performance metrics.
  • Experience curating, managing, and versioning “Golden Datasets” for benchmarking model responses against ground truth in development and production environments.
  • Experience using tracing such as LangSmith and Arize Phoenix to monitor production inputs/outputs, identify drift, and capture user feedback.
  • Understanding of LLM architectures, RAG systems, agents, and common AI application failure modes.
  • Advanced proficiency writing maintainable Python code for test automation and orchestrating AI evaluation workflows.
  • Experience using Python to process large benchmarking datasets and automate interactions with LLM/AI services.
  • Ability to integrate Python scripts into CI/CD pipelines for automated quality checks and performance observability in agentic systems.

Nice to have

  • Practical experience testing applications on AWS and working with cloud-based AI services.
  • Experience with monitoring and observability tools, log analysis, and performance metrics tracking, including Grafana.
  • Experience with Roche mandatory ALM tools, including Jira and GitHub.
  • Daily use of AI coding assistants and autonomous agents such as Claude Code and Ona to accelerate full-stack development and testing cycles.
  • Experience conducting rigorous code reviews for human-written and AI-generated code.
  • Understanding of security testing methodologies, particularly for AI applications.
  • Proven understanding of AI quality assurance and testing standards within highly regulated industries.

Details

  • Located in Hyderabad, India.
  • Working hours are structured to capture golden-hours overlap with Central European Time, typically through the IST evening.
  • Shift: CET time zone.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Fullstack AI Quality Engineer – Applied & Agentic AI Systems