Overview
Moniepoint is hiring a Site Reliability Engineer (SRE) to help ensure systems run smoothly, improve observability and reliability, and reduce repetitive operational work.
What you'll do
- Participate in on-call rotations to detect and triage service and reliability issues across all environments.
- Act as Incident Commander during major incidents, coordinating cross-functional teams and providing stakeholder updates.
- Create and maintain dashboards and alerts.
- Instrument code with development teams to improve visibility.
- Develop automation to eliminate manual and repetitive operational tasks (toil) for applications and infrastructure.
- Implement and track SLIs and SLOs for service reliability.
- Investigate and resolve customer complaints escalated beyond L1/L2, especially performance and reliability issues.
What you'll need
- Minimum 3 years of experience supporting enterprise applications as an SRE or similar role.
- Proficiency in writing code in Java, Go, or Python.
- Understanding of distributed systems concepts, microservices architecture, and software design patterns.
- Hands-on experience with Kubernetes.
- Experience managing applications on a major cloud provider (GCP, AWS, or Azure) and troubleshooting container issues.
- Experience setting up dashboards in Grafana and using APM tools like Datadog, New Relic, or Signoz.
- Proficiency in SQL (e.g., PostgreSQL, MySQL) including writing complex queries to debug data issues.
Details
- Work mode: Remote.
- Location: South Africa.
Read the full description and apply on the company’s own careers page.