Overview
Lead the technical direction, architecture, and delivery of Temporal’s distributed replication systems across open source and cloud environments. Own complex initiatives spanning design, implementation, rollout, operations, reliability, and scalability.
What you'll do
- Set the architecture and technical direction for the OSS replication stack.
- Design and implement protocols for high-availability namespaces and cross-cluster replication.
- Enable migration across Temporal clusters, including cloud and self-hosted scenarios.
- Drive scalability initiatives involving multi-cell namespaces, load distribution, and hot spots.
- Define consistency, ordering, idempotency, recovery, performance, and operational guarantees.
- Lead design reviews, mentor engineers, and guide technical practices across teams.
- Debug production issues and contribute to incident response and follow-up improvements.
What you'll need
- Experience designing and delivering complex production distributed systems.
- Deep knowledge of replication, partitioning, consistency, fault tolerance, durability, concurrency, and failure recovery.
- Experience defining architecture, protocol behavior, invariants, and trade-offs.
- Experience debugging concurrency bugs, data inconsistencies, partial failures, and performance bottlenecks.
- Production coding proficiency in Go, including concurrent code.
- Strong written and verbal communication skills.
- Ability to influence technical direction across teams and mentor engineers.
Nice to have
- Experience designing or maintaining replication protocols or data-plane infrastructure.
- Experience with multi-cluster or multi-region architectures.
- Familiarity with database internals, log-based replication, or event-sourced systems.
- Contributions to large open source projects or distributed systems infrastructure.
Details
- Remote in the United States.
Read the full description and apply on the company’s own careers page.