Overview
Join the Digital Applications team as a Mid-Level Data Analyst – Observability & Telemetry to monitor, analyze, and optimize the performance, reliability, and health of customer-facing digital applications. Work with development, DevOps, and Site Reliability Engineering teams to implement observability practices and turn telemetry data into actionable insights, dashboards, and alerts.
What you'll do
- Assist application teams in instrumenting digital applications with OpenTelemetry APIs, SDKs, and auto-instrumentation agents to collect metrics, logs, and traces.
- Design, build, and maintain real-time operational dashboards and visualization tools in Splunk using SPL or ELK using Kibana.
- Analyze telemetry data to identify performance trends, system bottlenecks, latency issues, and operational anomalies across distributed systems.
- Support production support and SRE teams during incident investigations using distributed tracing and log correlation for rapid root cause analysis.
- Ensure telemetry data is clean, structured, and compliant with enterprise logging standards and data retention policies.
- Partner with software engineers to map end-to-end user journeys and ensure monitoring coverage across HTTP/S APIs, microservices, and messaging queues such as Kafka.
- Maintain documentation of telemetry configurations, dashboard designs, alerting thresholds, and monitoring best practices.
- Turn technology stacks and application designs into code across development platforms, including iOS, Android, web/Angular, and services.
- Develop, design, construct, test, and implement secure, stable, testable, and maintainable code.
- Analyze application systems and programming activities, including feasibility studies, time and cost estimates, and implementation of new or revised systems and programs.
- Build and maintain integrated project development schedules that account for dependencies, SDLC approaches, constraints, and contingency for unplanned delays.
- Negotiate features and priorities and help teams and customers reach consensus.
- Improve team development processes to accelerate delivery, drive innovation, lower costs, and improve quality.
- Automate code quality, code performance, unit testing, and build processing in CI/CD pipelines using RTC, Jenkins, and RLM.
- Assess risk in business decisions, drive compliance with applicable laws, rules, and regulations, and escalate, manage, and report control issues transparently.
What you'll need
- Bachelor’s degree in Computer Science, Information Technology, Data Analytics, or a related field, or equivalent practical experience.
- Bachelor’s/University degree or equivalent experience.
- 5+ years in an Apps Development role.
- 3 to 5 years of professional experience in data analysis, systems monitoring, application support, SRE, or a related technical role.
- Minimum 1+ years of hands-on experience with OpenTelemetry or Global Cloud Observability frameworks, including tracing, metrics, and logs.
- Significant hands-on experience of 2+ years using Splunk or the ELK Stack.
- Experience with Splunk search queries, alerts, and dashboards, or Elasticsearch, Logstash, and Kibana.
- Familiarity with containerized environments such as Kubernetes or Red Hat OpenShift and basic cloud infrastructure concepts.
- Strong problem-solving skills to correlate logs, metrics, and traces to diagnose complex system behaviors.
- OpenTelemetry SDKs and Collectors, Splunk SPL, and ELK Stack tools including Kibana and Elasticsearch.
- Splunk Search Processing Language, Elasticsearch Query DSL, and SQL.
- Basic scripting skills in Python, Bash, or PowerShell for data parsing and automation tasks.
- Understanding of microservices, REST APIs, and event-driven architectures such as Apache Kafka.
- Ability to read and understand basic application code in Java, Node.js, or Go to assist with telemetry instrumentation.
- Strong analytical and quantitative skills; data driven and results-oriented.
- Experience running high traffic, distributed, cloud based services.
- Experience affecting large culture change.
- Experience leading infrastructure programs.
- Skill working with third party service providers.
- Excellent written and oral communication skills.
- Strong team player able to work effectively across cross-functional engineering and operations teams.
- Ability to explain technical performance data to technical and non-technical stakeholders.
- High precision in configuring alerts and analyzing system metrics to prevent false positives and negatives.
- Demonstrated execution capabilities.
Details
- Location: Pune, Maharashtra, India.
- Employment type: Full time.
Read the full description and apply on the company’s own careers page.