Overview
Site Reliability/Cloud Platform Engineer for operations, building and running reliable infrastructure for a multi-tenant, security-sensitive B2B cloud platform under US time zones.
What you'll do
- Provide operational support aligned with US time zones to ensure system reliability and availability.
- Design and build end-to-end automation for provisioning, upgrades, scaling, and decommissioning across AWS/GCP and deployment models.
- Write production-grade Go and Python services and CLIs to convert infrastructure operations into self-service platform capabilities.
- Own and evolve the Infrastructure-as-Code stack (Terraform/OpenTofu, Helm, GitOps/ArgoCD) and drive Kubernetes/IaC migrations with minimal customer impact.
- Operate and scale core platform services including Kubernetes clusters, Istio, Aerospike/PostgreSQL, Kafka, and GPU-backed inference workloads.
- Build observability and alerting, run on-call, lead incident response, and produce root-cause analyses with corrective action.
What you'll need
- 4+ years of experience in platform engineering, infrastructure engineering, DevOps, or SRE with real ownership of production cloud infrastructure.
- Strong software engineering skills in Go and/or Python.
- Hands-on production experience with Kubernetes.
- Infrastructure-as-Code experience with Terraform, Pulumi, or OpenTofu.
- Experience with at least one major public cloud (AWS or GCP or Azure).
- Track record of building tools/platforms for other engineers rather than only maintaining infrastructure.
Details
- Time zone alignment: US time zones (PST timezone mentioned in job title).
Read the full description and apply on the company’s own careers page.