Overview
Lead Platform Engineer (DevOps) responsible for operating and improving large-scale distributed, fault-tolerant systems with a Platform Engineering (SRE) approach.
What you'll do
- Manage system uptime across cloud-native (AWS, GCP) and hybrid architectures.
- Build infrastructure-as-code patterns that meet security and engineering standards (e.g., Terraform, cloud CLI scripting, cloud SDK).
- Develop CI/CD pipelines for build, test, and deployment using platform (Jenkins) and cloud-native toolchains.
- Create automated tooling to deploy service requests into production.
- Produce detailed runbooks to detect, remediate, and restore services.
- Triage complex distributed service issues and provide on-call support for high severity incidents.
- Lead availability reviews and own remediation actions from blameless postmortems to reduce repeat issues.
What you'll need
- BS in Computer Science or related field (or equivalent coding-related job experience required).
- 7+ years experience across software engineering, systems administration, database administration, and networking.
- 3+ years developing and/or administering software in public cloud (AWS, GCP, Azure).
- Experience monitoring infrastructure and application uptime/availability.
- Programming/scripting experience in languages such as Python, Bash, Java, Go, JavaScript, and/or node.js.
- Cross-functional knowledge spanning systems, storage, networking, security, and databases.
- Linux/Windows system administration automation/orchestration using Terraform, Chef, Ansible, and/or containers (Docker, Kubernetes, etc.).
Details
Read the full description and apply on the company’s own careers page.