Leads a multple teams to implement strategies for the architecture and delivery of interdependent, scalable distributed systems that meet organizational and customer demands. Orchestrates cross-group optimization for high‑throughput, large‑scale data processing; aligns stakeholders on scalability requirements; and oversees elastic designs and effective use of data plane platforms. Provides strategic oversight for fault‑tolerant, in‑service‑upgradable architectures, sets direction for partition‑aware design choices, and leads initiatives to harden networks via load‑shedding, throttling, and rate‑limiting. Establishes expectations for formal verification and peer reviews, and sets SLO‑aligned durability and availability standards across the department. Drives KPI and telemetry strategies; directs creation of complex dashboards and alerting for proactive health assurance; and ensures functional/correctness validation, data replication, and synchronization meet organizational needs.
Guides organization‑wide incident management and operational readiness, eliminating customer maintenance windows and ensuring consistent SOPs. Provides strategic security guidance (encryption, access controls), oversees remediation and compliance documentation, and sponsors automation (IaC) and change‑management alignment so systems can be safely patched, updated, and rolled back at scale.
Responsibilities
Key Responsibilities
System Design & Architecture – System Scalability:
- Implements strategies across multiple teams or groups for the architecture and design of interdependent scalable distributed systems, including the use of distributed state management tools, ensuring organizational and system demands are met.
- Spearheads code and/or system optimization initiatives for large-scale data processing and high-throughput requirements across multiple areas, driving improvements that support hyper-scale systems.
- Facilitates collaborations to define sys