Overview
As a Senior DevOps Engineer, you will evolve the infrastructure, delivery systems, and operational practices that enable engineering teams to build and run reliable software. You will partner with software engineers to improve developer experience and strengthen the reliability, security, performance, and cost efficiency of Karat’s hosted SaaS platform.
What you'll do
- Own and evolve Karat’s AWS SaaS infrastructure, ensuring services are secure, scalable, reliable, observable, and cost-efficient.
- Design, improve, and operate CI/CD pipelines using CircleCI and related tooling, improving deployment safety, speed, repeatability, and developer experience.
- Build and maintain observability capabilities including metrics, logs, traces, dashboards, and actionable alerting using Datadog and related tools.
- Apply Site Reliability Engineering principles to define and improve service reliability, availability, performance, capacity planning, incident response, root-cause analysis, and operational learning.
- Influence engineering-wide technical decisions and delivery practices through strong partnership, practical standards, and clear communication.
What you'll need
- 5+ years of experience in DevOps, Site Reliability Engineering, infrastructure engineering, platform engineering, or a closely related discipline.
- Significant hands-on production experience with AWS; this is a required qualification.
- Demonstrated experience designing, operating, and improving CI/CD systems using CircleCI, GitHub Actions, Jenkins, GitLab CI, or another major CI/CD platform.
- Strong experience with Docker and containerized application environments.
- Strong Linux, networking, security, and cloud-infrastructure fundamentals.
- Practical experience applying SRE principles to production systems, including observability, alerting, incident response, root-cause analysis, capacity planning, and reliability improvement.
- Hands-on experience with a leading telemetry and observability platform, such as Datadog, New Relic, Dynatrace, Grafana Cloud, or Splunk.
- Experience designing useful monitoring and alerting systems that reduce noise, support effective incident response, and improve service ownership.
- Experience with cloud cost management and optimization, including the ability to make pragmatic tradeoffs among cost, reliability, performance, and engineering velocity.
- Experience partnering with application-engineering teams to establish shared infrastructure and operational practices.
- Clear written and verbal English communication skills, including the ability to explain complex infrastructure and reliability issues to varied technical and non-technical audiences.
- Comfort working with globally distributed teams and regularly collaborating with colleagues in the United States.
Nice to have
- Experience with CircleCI is strongly preferred.
- Datadog experience is strongly preferred.
Details
- The position is available only to candidates residing in Bengaluru (formerly known as Bangalore), India.
- The team operates 100% remotely.
- The schedule must overlap with U.S. business hours.
- Applicants must submit materials that are 100% in English.
Read the full description and apply on the company’s own careers page.