A

Observability Platform Engineer

American Express Global Business Travel
NewPosted today

LOCATION

Bangalore · Onsite

EXPERIENCE

5+ Years

TYPE

FullTime

SALARY

Negotiable

SKILLS REQUIRED

ObservabilitySystem MonitoringIncident responsedistributed tracingAlertingDashboarding

Job description

Overview

Design, deploy, and optimize observability platforms that support system reliability and performance. The role works with engineering, development, and operations teams on monitoring, incident response, and observability best practices.

What you'll do

  • Design and deploy observability platforms using tools such as ELK Stack, New Relic, Datadog, and Alertsite.
  • Develop and maintain monitoring strategies, dashboards, and alerting rules to ensure system reliability and performance.
  • Collaborate with engineering teams to instrument applications and infrastructure for comprehensive observability.
  • Fix complex system issues using observability data and provide actionable insights.
  • Establish best practices for logging, metrics collection, and distributed tracing.
  • Optimize observability infrastructure for cost-efficiency and performance.
  • Conduct training and knowledge-sharing sessions with development and operations teams.
  • Participate in on-call rotations and incident response activities.
  • Continuously evaluate and recommend new observability tools and technologies.

What you'll need

  • 5+ years of experience in platform engineering, DevOps, or systems engineering roles.
  • Hands-on expertise with at least two of the following platforms: ELK Stack, New Relic, Datadog, or Alertsite.
  • Strong understanding of monitoring, logging, metrics, and alerting concepts.
  • Validated experience creating and maintaining monitoring dashboards and visualizations.
  • Hands-on experience implementing synthetic monitoring and end-to-end transaction monitoring.
  • Knowledge of Application Performance Monitoring (APM) concepts and implementation.
  • Knowledge of Real User Monitoring (RUM) and digital/browser/mobile app observability.
  • Knowledge of SLI/SLO definition and measurement methodologies.
  • Familiarity with MTTA, MTTR, MTTD, and other incident metrics.
  • Proficiency in Python, Bash, or similar scripting languages.
  • Experience with AWS, Azure, or GCP cloud platforms.
  • Knowledge of Docker and Kubernetes containerization and orchestration technologies.

Nice to have

  • Experience with multiple observability platforms.
  • Familiarity with Infrastructure as Code using Terraform or CloudFormation.
  • Background in incident management and on-call operations.
  • Certifications in relevant cloud or observability platforms.
  • Experience with Application Performance Monitoring (APM) tools.
  • Knowledge of security and compliance monitoring.

Details

  • Location: Bangalore, India.
  • Participate in on-call rotations and incident response activities.
  • Observability platforms and technical skills include ELK Stack, New Relic, Datadog, Alertsite, Amplitude Analytics, Splunk or similar, Python, Bash, Go or similar languages, AWS, Azure, GCP, Docker, Kubernetes, Prometheus, Grafana, CloudWatch, Azure Monitor, SQL, NoSQL, TCP/IP, DNS, and HTTP/HTTPS protocols.

Read the full description and apply on the company’s own careers page.

Stay safe

Hiring on Abekus is free for applicants

We never charge a fee, and employers are prohibited from doing so. If a recruiter asks for payment, please report them right away.

Observability Platform Engineer