Overview
Senior/Staff Machine Learning Research Engineer on Scale AI’s General Agents team, designing and deploying production-ready AI agents for enterprise use cases.
What you'll do
- Design and implement end-to-end agent systems combining LLM reasoning, tool use, memory, and control logic.
- Build scalable, reliable agent architectures deployable across customers with varying data and constraints.
- Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance and reliability in production.
- Collaborate with product managers, customers, data annotators, and engineering teams to translate enterprise requirements into agent designs.
- Productionize frontier agent techniques such as planning, multi-step reasoning/tool-use, and multi-agent patterns.
- Own deployment, monitoring, iteration, failure analysis, and continuous improvement based on real-world usage.
- Provide technical direction and architectural decisions for general agent development best practices at the Staff level.
What you'll need
- 5+ years of experience building and deploying machine learning or AI systems for real-world production use cases.
- Strong engineering fundamentals with a Bachelor’s and/or Master’s degree in Computer Science, Machine Learning, AI, or equivalent practical experience.
- Deep understanding of modern LLMs and prompt-, context-, and system-level optimization plus agentic system design.
- Proficiency in Python for production-quality, testable, maintainable code.
- Experience integrating models with external tools, APIs, databases, and services.
- Ability to operate in ambiguous problem spaces combining research approaches with pragmatic product constraints.
- Strong communication skills for customer-facing or cross-functional environments.
Nice to have
- Hands-on experience building AI agents using modern generative AI stacks (OpenAI APIs, commercial or open-source LLMs).
- Experience with agent frameworks, orchestration layers, or workflow systems (e.g., tool calling, planners, multi-agent setups).
- Familiarity with evaluation, monitoring, and observability for LLM-powered systems in production.
- Experience deploying ML systems in cloud environments and operating them at scale.
- Experience fine-tuning or adapting foundation models using SFT, RLVR, or LoRA.
- Interest in shaping general-purpose enterprise agents and their real-world impact.