Overview
MLOps Engineer to own the infrastructure that runs Saaf AI systems already in production for mortgage lending.
What you'll do
- Re-architect AI service deployment from single-host setups to horizontally scalable infrastructure.
- Design infrastructure for LLM-backed workloads including long-running requests and streaming responses.
- Own capacity planning, autoscaling, and unit economics for AI serving.
- Build environment promotion from development through pre-production to production with consistent, reproducible environments.
- Implement version-controlled, reviewable infrastructure/config/app logic and CI/CD with rollback and staged rollout.
- Create internal self-serve tooling so engineers can define, modify, and test AI workflow logic without touching deployment plumbing.
- Instrument, monitor, and operate the AI stack including latency/throughput, cost signals, alerting, and incident practice.
What you'll need
- Production ownership of containerized services under real traffic, including rollouts, autoscaling, failure isolation, and rollback.
- Infrastructure as code with declarative, reproducible environment setup.
- CI/CD depth including automated testing, promotion between environments, safe rollout, and recovery.
- Strong Python skills to read/change code, profile, and fix application logic.
- Operational judgment to debug across application, network, and infrastructure boundaries.
- End-to-end ownership experience from design through monitoring and follow-up when issues occur.
Details
- Location: India.
- Work mode: Remote-first with flexible hours.
Read the full description and apply on the company’s own careers page.