01Responsibilities
- Define and lead the enterprise observability architecture across logs, metrics, traces, events, and telemetry pipelines.
- Establish OpenTelemetry as a core observability standard across applications, services, APIs, and infrastructure.
- Design and implement distributed tracing patterns using OpenTelemetry, including context propagation, trace correlation, sampling, and service dependency mapping.
- Build reusable instrumentation standards, libraries, SDKs, and reference implementations for engineering teams.
- Improve adoption of Observability as Code, using tools such as Terraform, CI/CD pipelines, and automation frameworks.
- Design scalable telemetry ingestion, processing, routing, enrichment, and cost-optimization strategies.
- Integrate OpenTelemetry with enterprise observability platforms such as Splunk, Dynatrace, Cribl, Prometheus, Grafana, Datadog, and related tools.
- Create self-service onboarding patterns, APIs, and automation to help product and platform teams adopt observability.
- Partner with SRE, DevOps, platform engineering, product engineering, cloud, security, and operations teams to define observability standards and guardrails.
- Establish Measurements, governance models, best practices, and architectural patterns for enterprise observability.
- Evaluate current and latest observability technologies and guide the roadmap for tooling modernization.
- You will report to Director.
Summary:
This is a senior architecture role for someone who can lead the modernization of observability across the enterprise. You will not only understand observability tools but will define standards, build reusable frameworks, automate adoption, and make OpenTelemetry a foundational capability across the organization.
Qualifications
Required Experience:
- 1015+ years of experience in software engineering, platform engineering, SRE, DevOps, cloud infrastructure, or observability architecture.
- We require hands-on experience with OpenTelemetry.
- 5+ years of experience implementing distributed tracing, metrics, logs, and telemetry pipelines at scale.
- 5+ years of experience with OpenTelemetry components, including SDKs, Collector, exporters, processors, semantic conventions, and instrumentation libraries.
- Experience with observability tools such as Dynatrace, Splunk, Cribl, Prometheus, Grafana, Datadog, or similar platforms.
- Coding experience in Python, Go, Java, or similar languages.
- 5+ years of experience with infrastructure as code and automation tools such as Terraform, Ansible, GitHub Actions, Jenkins, or similar CI/CD platforms.
- Experience with cloud-native architecture, microservices, Kubernetes, APIs, and hybrid cloud environments.
- Define enterprise standards, engineering teams, and increase adoption across large, distributed organizations.
- 5+ years of experience building internal observability platforms or telemetry frameworks.
- Experience with OpenTelemetry Collector deployment patterns in Kubernetes, cloud .