Overview
Senior Software Engineer specializing in Generative AI and Large Language Models, focused on agentic systems, Retrieval-Augmented Generation (RAG), and AI evaluations.
What you'll do
- Design, build, and operate multi-agent workflows and tool-enabled agents with orchestration, state management, safety guardrails, and fallback strategies.
- Architect and maintain end-to-end RAG systems from document ingestion to retrieval, reranking, and answer synthesis.
- Evaluate and integrate LLMs and GenAI services across cost, performance, and privacy tradeoffs.
- Develop and optimize prompting strategies with automated prompt testing and regression tracking.
- Define and own evaluation frameworks for generative outputs, including automated metrics, LLM-as-judge, and human evaluation.
- Build and operate scalable, secure AI infrastructure on AWS and own the deployment lifecycle.
What you'll need
- Production experience designing multi-agent systems with tool use, memory/state management, and fault-tolerant routing.
- Hands-on experience building end-to-end RAG pipelines including chunking, embeddings, vector retrieval, and answer synthesis.
- Strong experience defining and running evaluation pipelines for generative AI, including hallucination mitigation and drift monitoring.
- Experience with foundation models and GenAI providers including AWS Bedrock, OpenAI, Anthropic, or Meta.
- Solid grounding in classical NLP techniques such as NER, text classification, intent detection, or topic modelling.
- Hands-on experience with Amazon Bedrock foundation model APIs, Bedrock Agents, Knowledge Bases, and Guardrails.
Nice to have
- Experience with Bedrock Model Evaluation.
- LLM-as-judge patterns are a plus.
- Familiarity with Langfuse, Arize, or Langsmith for tracking runs, metrics, and artefacts.
Details
- Location: Germany (Remote).
Read the full description and apply on the company’s own careers page.