Overview
Build reinforcement learning environments and verification systems that enable AI agents to discover, exploit, and remediate software vulnerabilities using deterministic, verifiable rewards.
What you'll do
- Design and deploy high-throughput RLVR environments for evaluating LLM action sequences.
- Build test harnesses, sandboxes, and verification engines for vulnerability research.
- Integrate automated analysis platforms and vulnerability feeds into RL environment pipelines.
- Architect isolated, reproducible environments for parallel exploitation and patching trajectories.
- Implement telemetry, verification algorithms, and trajectory logging to detect reward hacking.
- Build instrumentation and debugging tools for memory, process, and network monitoring.
- Benchmark AI models across fuzzing, exploit generation, and patch validation tasks.
What you'll need
- Understanding of RL training workflows for modern LLM systems, including RLVR.
- Proficiency in Python and C programming.
- Understanding of software vulnerabilities, binary exploitation, fuzzing, or program analysis.
- Experience with DevOps pipelines and reproducible builds using Docker, BuildKit, or Nix.
- Comfort working with Linux systems and low-level debugging.
- Experience building or working with benchmark environments, CTFs, or execution sandboxes.
Nice to have
- Rust experience.
- Experience designing reward functions, ground-truth verifiers, or automated grading engines.
- Background with sanitizers, compiler instrumentation, or automated exploit generation tools.
- Experience with cybersecurity benchmarks, CTFs, or open-source AI evaluation frameworks.
Details
- Remote, work-from-home 100% of the time.
- US national base salary range: $176,400–$242,550 yearly.
Read the full description and apply on the company’s own careers page.