Overview
Design, implement, and optimize observability solutions for Charles River services across enterprise SaaS and on-premises environments. The role covers application code, telemetry instrumentation, observability platform configuration, dashboards, alerting, automation workflows, proactive monitoring, root-cause analysis, and automation-assisted remediation.
What you'll do
- Own and advance observability enablement for key Charles River services, including implementation standards, telemetry patterns, dashboards, alert quality, and adoption support.
- Design and implement tooling that improves visibility into service health, performance, availability, and client-impacting issues.
- Implement, configure, and optimize Dynatrace monitoring solutions across SaaS and on-premises environments.
- Build and maintain Dynatrace Dashboards Gen 3 for unified visibility, executive reporting, and operational insights.
- Use Dynatrace SaaS, Grail, DQL, Davis AI, OneAgent, OpenTelemetry, synthetic monitoring, and real user monitoring where appropriate to improve observability coverage.
- Develop reusable implementation patterns and standards that enable consistent telemetry collection across critical services.
- Define log-based alerts, service-level indicators, trace-log correlation, batch metrics, and alert-quality standards for production services.
- Integrate Dynatrace with external systems, including Azure Event Hub or similar event-publishing platforms, to support downstream workflows and operational reporting.
- Improve root-cause analysis by correlating application, infrastructure, log, trace, metric, and user-experience signals.
- Conduct performance impact assessments using custom service detection rules, distributed traces, logs, and service health data.
- Partner with SRE, product, technical support, and engineering teams to ensure observability coverage for critical services and client-facing workflows.
- Support automation-assisted remediation by connecting observability insights to approved automation workflows where appropriate.
- Lead training sessions for SaaS, technical support, SRE, and engineering teams on Dynatrace usage, dashboard interpretation, and observability best practices.
- Develop documentation, onboarding materials, reference implementations, and dashboard templates for observability tools.
- Promote adoption of observability standards across teams by translating technical telemetry into actionable operational guidance.
What you'll need
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
- 12+ years of experience designing, developing, implementing, or operating enterprise software systems and applications.
- 8+ years of hands-on experience in observability, application performance monitoring, synthetic monitoring, telemetry instrumentation, dashboards, alerting, and trend analysis.
- Hands-on experience with Dynatrace SaaS, DQL, Logs on Grail, Dashboards Gen 3, and related Dynatrace capabilities.
- Strong Java or JVM-based engineering experience, including the ability to understand application behavior, instrumentation needs, and production performance characteristics.
- Experience with distributed systems, microservices, APIs, cloud-native architectures, and production support practices.
- Experience with log ingestion, trace correlation, service detection, metrics, and real user monitoring.
- Ability to collaborate with engineering, SRE, SaaS operations, technical support, product, and client-facing teams to translate observability data into actionable insight.
- Excellent written and verbal communication skills, with the ability to explain technical findings clearly to both engineering and operational stakeholders.
Nice to have
- Experience with React or other frontend frameworks, especially for frontend observability, real user monitoring, and user-experience diagnostics.
- Experience with Azure Event Hub, Power BI, Grafana, Prometheus, or other visualization and telemetry platforms.
- Knowledge of Kubernetes, containerized workloads, and data pipeline observability.
- Familiarity with automation, DevOps, CI/CD, and infrastructure-as-code tools and practices, including Terraform, Jenkins, Git, Ansible, and API integrations.
- Experience with OpenTelemetry, observability as code, SLOs, SLIs, error budgets, and alert tuning.
- Familiarity with Citrix environments, agentless deployments, .NET, Charles River IMS, enterprise SaaS, or financial services technology environments.
Details
- Location: Hyderabad, India.
Read the full description and apply on the company’s own careers page.