Overview
Forward Deployed Engineer (FDE) focused on inference optimization and post-training for production AI teams, partnering with customer production AI teams as well as Solutions Architects.
What you'll do
- Select, configure, and optimize inference engines based on hardware, model architecture, and workload.
- Develop configuration updates to improve critical POCs, benchmarks, and customer deployment performance.
- Tune KV cache, apply speculative decoding, determine tensor parallelism, and choose quantization strategies to meet throughput and latency targets.
- Run hands-on RL training and optimize system design through LoRA, SFT, DPO, RLHF, and GRPO pipelines.
- Serve as primary technical point of contact for strategic accounts by monitoring and optimizing endpoint configurations and milestone progress.
- Provide field insights to influence the software and model roadmap.
What you'll need
- 5+ years of experience in a technical role with a focus on inference systems, open-source LLM deployment, or post-training workflows.
- Expert, hands-on experience with inference engines such as vLLM, TensorRT-LLM, and SGLang.
- Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, pipeline parallelism, and quantization techniques.
- Hands-on experience with fine-tuning and post-training pipelines including LoRA, SFT, DPO, RLHF, and GRPO.
- Strong Python skills and comfort working in production environments.
- Broad knowledge of state-of-the-art open-source models and judgment on model selection for specific customer use cases and hardware profiles.
Details
- Must be a permanent resident or citizen of Singapore.
Read the full description and apply on the company’s own careers page.