At ModMed, we re not just building software-we re reimagining the healthcare experience. Founded in 2010 by a practicing physician and a successful tech entrepreneur, we took a radically different approach: we hired doctors and taught them how to code. This "for doctors, by doctors" philosophy has allowed us to create an AI-enabled, specialty-specific cloud platform that places patients at the center of care.
Job Description Summary
The Role: Technical Visionary for Reliability and Developer Velocity. As a Principal Site Reliability Engineer at ModMed, you are a primary architect of our technical future. You don't just solve problems; you anticipate the needs of a global healthcare platform years in advance. You will think big to define the standards for reliability, scalability, developer velocity, and security that allow our doctors to provide world-class care without interruption. This is a high-impact technical leadership role where you will innovate boldly, then make things happen, acting as a vital bridge between engineering, product,
and business goals. By aligning passion with purpose, you will build a resilient infrastructure that serves as the backbone for up-to-date medicine and consistently creates customer delight.
Primary Responsibilities
- Core Infrastructure Strategy & Architecture: Architectural Strategy & Technical Governance: Define and execute the long-term architectural vision for our global AWS cloud ecosystem. Design high-performance, fault-tolerant, and cost-optimized multi-region topologies capable of scaling elastically to meet massive transaction volumes while enforcing rigorous cloud hygiene.
- Systemic Observability Ecosystems: Institutionalize "Observability as a Culture" across the enterprise. Architect unified, global telemetry standards using Datadog and OpenTelemetry, transforming raw logs, metrics, and distributed traces into actionable, predictive insights that elevate system reliability across al