Overview
Roche is seeking a Fullstack AI Quality Engineer to test Generative and non-Generative AI applications and agentic systems. The role covers manual exploratory testing, automated UI/API/MCP testing, classical machine learning validation, and generative and agentic AI evaluation within a lean software team.
What you'll do
- Design and implement testing strategies for LLM-based applications and agentic systems.
- Evaluate model responses, RAG pipeline accuracy, system reliability, autonomous multi-agent behaviors, and tool-use accuracy.
- Design and implement automated integration, performance, and AI behavior tests across Python backend services and TypeScript/React/Angular frontends.
- Use, co-create, and maintain testing frameworks covering traditional software quality and AI-specific concerns such as output consistency, contextual accuracy, and ethical compliance.
- Create and maintain test strategies, test cases, testing guidelines, and comprehensive test documentation for AI applications.
- Identify and test edge cases involving ambiguous inputs, potential biases, and unexpected response patterns.
- Design and execute tests for bias and fair treatment across different user groups and use cases.
- Test LLM application vulnerabilities, including prompt injection, data leakage, and other AI-specific security concerns.
- Conduct performance testing, including response time analysis, load testing, and resource utilization monitoring.
- Stay current with AI and GenAI testing methodologies and best practices.
- Clearly document findings and recommendations and communicate them in English at the level of C1+.
What you'll need
- B.Sc., B.Eng., M.Sc., M.Eng. in Computer Science, Software Engineering, or equivalent degree with a strong background in testing methodologies.
- 7+ years of experience in software testing.
- At least 1 year focused on genAI and/or agentic applications.
- Experience working in regulated environments with qualified infrastructure and validated applications.
- Understanding of the difference between GxP and non-GxP products and the importance of CSV and ITSM processes in delivery.
- Experience working with Agile development teams.
- Passion for quality assurance and testing, particularly for AI and GenAI applications and agents.
- Strong analytical and problem-solving skills with attention to detail.
- Critical thinking and problem-solving when using AI prompts, including in NLP, data distributions, and edge cases.
- Proactivity in identifying potential issues and suggesting improvements.
- Advanced command of Python and standard testing frameworks such as pytest and allure, alongside practical familiarity with specialized AI evaluation tools.
- High competence in TypeScript-driven test frameworks such as Playwright, Jest, and Cypress.
- Proven expertise architecting and maintaining end-to-end automated testing suites bridging UI/API validation with asynchronous AI/LLM evaluation pipelines.
- Hands-on experience integrating test suites within GitHub Actions or modern CI/CD pipelines.
- Expertise implementing evaluation harnesses such as RAGAS, DeepEval, and Promptfoo.
- Ability to test autonomous agentic loops, including tool-use accuracy, step-by-step reasoning verification, and goal-directed performance metrics.
- Experience curating, managing, and versioning “Golden Datasets” for benchmarking model responses against ground truth in development and production environments.
- Experience using tracing such as LangSmith and Arize Phoenix to monitor production inputs/outputs, identify drift, and capture user feedback.
- Understanding of LLM architectures, RAG systems, agents, and common AI application failure modes.
- Advanced proficiency writing maintainable Python code for test automation and orchestrating AI evaluation workflows.
- Experience using Python to process large benchmarking datasets and automate interactions with LLM/AI services.
- Ability to integrate Python scripts into CI/CD pipelines for automated quality checks and performance observability in agentic systems.
Nice to have
- Practical experience testing applications on AWS and working with cloud-based AI services.
- Experience with monitoring and observability tools, log analysis, and performance metrics tracking, including Grafana.
- Experience with Roche mandatory ALM tools, including Jira and GitHub.
- Daily use of AI coding assistants and autonomous agents such as Claude Code and Ona to accelerate full-stack development and testing cycles.
- Experience conducting rigorous code reviews for human-written and AI-generated code.
- Understanding of security testing methodologies, particularly for AI applications.
- Proven understanding of AI quality assurance and testing standards within highly regulated industries.
Details
- Located in Hyderabad, India.
- Working hours are structured to capture golden-hours overlap with Central European Time, typically through the IST evening.
- Shift: CET time zone.
Read the full description and apply on the company’s own careers page.