Overview
Staff Software Engineer to build a Kubernetes-native control plane for provisioning and operating a GPU inference fleet.
What you'll do
- Design and implement a provisioning state machine for the full host lifecycle.
- Build a self-service declarative API/controls so inference teams can request, scale, and tear down inference clusters.
- Automate self-healing by detecting degraded/failed nodes, draining them safely, and reintroducing healthy capacity.
- Own reliability features such as idempotency, retries, rollback, and drift detection.
- Partner with the inference/ML platform team to encode cluster needs as platform abstractions.
- Develop infrastructure code with testing, review, versioning, and CI/CD, and operate/support it in production.
What you'll need
- Strong software engineering background writing and testing real software in Go, Python, Rust, or similar.
- Experience with durable workflow orchestration tools (e.g., Temporal, Cadence) for long-lived manifest-driven workflows.
- Experience building control planes/orchestration systems that model state and reconcile over time (e.g., Kubernetes controllers).
- Experience with event-driven systems using message queues/event streams/pub-sub (e.g., Kafka, NATS, SQS).
- A product mindset, building internal platforms/APIs consumed by other engineering teams.
Details
- Work mode noted as “Remote in India”.
- Location listed as India.
Read the full description and apply on the company’s own careers page.