Overview
Build and operate the infrastructure and platforms powering ActAI’s AI capabilities, from model training and evaluation through deployment, inference, observability, and continuous improvement.
What you'll do
- Build and operate ML infrastructure and platforms for AI products.
- Design systems for model training, evaluation, deployment, inference, and experimentation.
- Optimise model-serving infrastructure for high-throughput and low-latency workloads.
- Develop pipelines for data preparation, training, evaluation, model release, and improvement.
- Build evaluation, benchmarking, monitoring, tracing, and alerting infrastructure.
- Identify ML-stack bottlenecks and improve reliability, scalability, latency, throughput, and cost.
- Partner with AI, research, and product teams to deliver production-ready infrastructure.
What you'll need
- Strong software engineering fundamentals and production-systems experience.
- Experience building ML infrastructure, platforms, or production machine-learning systems.
- Experience with model deployment, inference, evaluation, or data pipelines.
- Strong understanding of distributed systems and system reliability.
- Ability to write clean, maintainable, production-quality code.
- Python experience.
- Experience with PyTorch or JAX, cloud infrastructure, and ML/data pipelines.
Nice to have
- Experience with vLLM, SGLang, or TensorRT-LLM.
- Experience with GPU infrastructure, performance tooling, vector databases, or retrieval infrastructure.
Details
Read the full description and apply on the company’s own careers page.