Description
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting). Provides oversight for KPIs, telemetry, and moderately complex dashboards; directs design of functional/correctness requirements, fault‑injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on‑call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change‑management plans enabling protected patching, updates, and rollbacks.
Responsibilities
Key Responsibilities
System Design & Architecture – System Scalability:
- Manages the development and implementation of scalable distributed systems and components across multiple teams, including the effective use of distributed state management tools.
- Oversees code and/or system optimization efforts for large-scale data processing and high-throughput requirements within and across teams to support hyper-scale systems.
- Guides teams to define scalability requirements for owned components and ensures design and implementation requirements are met.
- Manages the use of data plane platforms to effectively handle large-scale data retrieval, storage, and processing.
- Ensures team accurately designs performance and load testing.
System Design & Architecture – System Reliability Design:
- Manages the strategy for building fault-tolerant components and systems capable of withstanding in-service updates by guiding t