Overview
Zafin is seeking a Cloud Site Reliability Engineer I (CSRE I) to drive reliability, scalability, and performance improvements across cloud infrastructure and applications.
What you'll do
- Manage resolution of complex technical issues across Zafin products and the Azure cloud environment.
- Design and implement operational enhancements to improve resiliency and system reliability.
- Perform root cause analysis (RCA) for high-severity incidents and reduce error recurrence.
- Represent the organization in external client escalation calls with expert guidance.
- Optimize cloud infrastructure for performance, scalability, and cost-effectiveness.
- Oversee monitoring solutions and integrate predictive analytics for proactive issue resolution.
- Develop automation strategies for operational workflows and incident responses.
- Maintain documentation of cloud architectures, processes, and incident management strategies.
What you'll need
- 8+ years of experience in cloud support, operations, or a related role.
- Advanced expertise in Microsoft Azure (preferred) or equivalent cloud platforms.
- Experience designing and scaling container orchestration systems such as AKS or OpenShift.
- Experience managing automated deployment pipelines, including Azure DevOps.
- Experience with enterprise monitoring platforms (e.g., Azure Insights, Grafana) and predictive analytics tools.
- Advanced scripting skills with PowerShell, Python, or similar languages.
- Experience in incident management and defining SLAs for global production environments.
- In-depth knowledge of database management, particularly Postgres.
- Bachelor’s degree in computer science, engineering, or related field (Master’s degree preferred).
Read the full description and apply on the company’s own careers page.