24 Sep
|
ZEISS VisioGen
|
India
24 Sep
ZEISS VisioGen
India
About ZEISS VisioGen
ZEISS VisioGen is building the future of AI-powered healthcare by developing human-verified clinical communication solutions for ophthalmology practices. We combine cutting-edge Generative AI with clinical expertise to build scalable, production-ready applications that enhance patient engagement and improve healthcare delivery.
Role Overview
We are looking for a Support Lead to build, manage, and be the ultimate escalation point for a 24x7 operations support team covering the Visiogen platform — HiApp (Human Intelligence App), our primary and most business-critical application, serving as the core platform Optometrists (ODs) use to monitor and respond to patient conversations across multiple clinics; Customer Portal; Operations Portal; and clinic chatbot integrations — all running on Zeiss's Azure infrastructure. This is a hands-on leadership role: you will manage a team of engineers across shifts, but you are also expected to personally step in and solve the issue when the team is stuck. You are the technical backstop as much as the people manager.
Key Responsibilities
Team Leadership & Shift Management
- Build, staff, and manage a 24x7 rotational schedule across a team in active rotation
- Ensure backup engineers remain current on the platform and cycle into the active rotation periodically, so coverage never depends on a single point of failure
- Mentor and upskill the team on HiApp, Customer Portal, Operations Portal, Azure infrastructure, and CI/CD processes
- Act as the final escalation point when the on-shift team cannot resolve an issue — personally driving it to resolution
- Own hiring, onboarding, and performance management for the support team
Incident & Escalation Ownership
- Own the incident management process end-to-end: severity classification (P1/P2/P3), coordination with engineering, and SLA adherence
- Lead major incident / outage bridges, making real-time judgment calls under pressure
- Drive post-incident root cause analysis and ensure corrective actions are actually implemented, not just documented
- Communicate incident status and business impact clearly to engineering and business stakeholders
Technical Oversight
- Maintain deep working knowledge across the full Azure stack: AKS, App Gateway (WAF_v2), PostgreSQL, Redis, Function App (EP1), ACR Premium, Log Analytics/App Insights, Key Vault, Private Endpoints, Storage, and Traffic Manager
- Oversee troubleshooting of React, .NET, and Python (Gen AI) service issues across HiApp and the Customer Portal
- Oversee the Function App-based data ingestion pipeline into the vector database, and troubleshoot failures in that pipeline
- Review and validate production releases and deployments to prevent avoidable incidents
Process & Continuous Improvement
- Establish and refine monitoring, alerting, and on-call processes to reduce time-to-detect and time-to-resolve
- Build and maintain runbooks and knowledge base content from real incidents — while still expecting the team to solve novel problems beyond the runbook
- Partner with Operations and Customer Success teams on clinic onboarding, and remove recurring friction in onboarding or chatbot integration
- Report on support metrics (SLA adherence, incident volume, MTTR, recurring issues) to leadership
What we’re looking for
Experience
Prior experience as a Forward Deployed Engineer, Technical Lead, or Support/SRE Lead, with direct, hands-on ownership of production issues—not purely a management role.
Cloud & Infrastructure
Strong hands-on expertise across AKS, App Gateway (WAF_v2), Private Endpoints, Traffic Manager, and Azure resource management at an architectural level.
Data Layer
Solid troubleshooting depth in PostgreSQL and Redis, including performance and scaling issues.
Serverless / AI Pipeline
Experience with Azure Function Apps (EP1) and data ingestion into vector databases.
DevOps & Release
Strong background in Azure DevOps, CI/CD pipelines, ACR (Premium), and production release governance.
Application Stack
Working fluency across React, .NET, and Python (Gen AI) services to lead root cause investigations—not just delegate them.
Leadership
Demonstrated experience managing or leading a 24x7 / shift-based support team, including scheduling and escalation ownership.
Problem-Solving A proven, specific track record of personally resolving hard, customer-specific, or one-off issues—this matters more than any certification.
Preferred Qualifications
1. Experience leading support or forward-deployed teams for a SaaS or platform product with multiple customer-facing applications
2. Experience in healthcare, clinic, or patient-facing platforms is a plus, not a requirement
3. Azure or Kubernetes certifications are welcomed but are not a substitute for hands-on incident-resolution experience
4. Experience setting up or maturing an on-call / 24x7 support function from the ground up
What Success Looks Like
1. The 24x7 rotation runs smoothly with consistent SLA adherence and low unplanned escalation to engineering leadership
2. Major incidents are resolved quickly, with explicit communication and solid root cause analysis
3. The team is capable of independently solving customer-specific issues, with the Lead only pulled in for the genuinely hard cases
4. Clinic onboarding and chatbot integration friction decreases over time due to process and technical improvements
📌 Support Lead (24x7 Operations Support) (India)
🏢 ZEISS VisioGen
📍 India