07 Aug
|
Vee Healthtek
|
Salem
07 Aug
Vee Healthtek
Salem
Role Overview
We are seeking a Senior Site Reliability Engineer to lead the reliability, scalability, and performance of our production systems.
This role blends software engineering, systems thinking, and operational excellence to ensure highly available and resilient systems. You will take end-to-end ownership of production reliability, lead complex incident resolution, and drive systemic improvements that eliminate recurring issues.
You will work closely with engineering teams to design for reliability from the ground up, not just operate systems after deployment.
Key Responsibilities
1. System Reliability & Availability
- Own availability, latency, and reliability of mission-critical systems
- Define and operationalize:
- Service Level Indicators (SLIs)
- Service Level Objectives (SLOs)
- Drive continuous improvements in system uptime and resilience
- Build safeguards for failover, disaster recovery, and high availability
2. Incident Management & Operational Leadership
- Lead high-severity incident response with clear ownership and direction
- Perform deep technical triage across:
- Application (.NET Core, Angular)
- Web server (IIS)
- Database (SQL Server)
- Cloud platforms (AWS, Azure)
- Ensure rapid service restoration with minimal business impact
- Drive blameless postmortems with actionable outcomes
3. Root Cause Analysis & Problem Elimination
- Go beyond symptom fixes to identify true root causes
- Identify patterns across incidents and eliminate entire classes of failures
- Drive long-term fixes through:
- Code changes
- Configuration improvements
- Automation
- Ensure RCA learnings are implemented and tracked to closure
4. Observability & Monitoring Excellence
- Design and evolve monitoring and observability strategy
- Leverage:
- Zoho Site24x7
- APM tools (AppDynamics, Dynatrace, New Relic, etc.)
- Improve:
- Alert quality (high signal, low noise)
- System visibility (logs, metrics, traces)
- Enable faster detection (MTTD) and diagnosis of issues
5. Performance Engineering & Scalability
- Identify and resolve system bottlenecks across:
- Application layers
- IIS configuration
- SQL Server performance
- Lead capacity planning and load readiness
- Optimize performance for high throughput and low latency systems
6. Automation & Toil Reduction
- Identify repetitive operational tasks and eliminate them through automation
- Build tools/scripts using:
- C#, PowerShell, or similar
- Automate:
- Recovery workflows
- Health checks
- Incident diagnostics
- Improve operational efficiency and reduce manual effort
7. Application & System Debugging
- Perform deep debugging of .NET Core applications hosted on IIS
- Analyze issues across:
- Angular UI API backend database flow
- Diagnose:
- Memory leaks, thread contention
- API failures and latency
- Integration breakdowns
- Use logs, APM traces, and metrics for precise root cause isolation
8. Cloud & Infrastructure Reliability
- Operate and troubleshoot systems in AWS and Azure environments
- Diagnose issues related to:
- Compute, storage, networking
- Autoscaling and resource limits
- Contribute to resilient architecture and failover strategies
9. Database Reliability Engineering (SQL Server)
- Troubleshoot and optimize:
- Slow queries, indexing issues
- Locking/blocking scenarios
- Improve database performance and reliability
- Ensure data integrity and efficient resource usage
10. Engineering Collaboration & Production Readiness
- Partner with development teams to:
- Improve system design and fault tolerance
- Embed observability into applications
- Ensure production readiness before releases
- Influence engineering standards for reliability and scalability
Required Qualifications
Experience
- 8+ years in SRE, DevOps, or Production Engineering
- Proven experience supporting high-availability, production systems
- Strong track record in incident resolution and system reliability improvement
Technical Skills
Application Stack
- Strong experience with:
- .NET Core
- Applications hosted on IIS
- Solid understanding of Angular-based frontends and API integrations
Cloud & Infrastructure
- Hands-on experience with:
- AWS and Azure
- Robust understanding of distributed systems and cloud architecture
Database
- Advanced experience with Microsoft SQL Server
- Expertise in query tuning and performance troubleshooting
Monitoring & Observability
- Experience with:
- Zoho Site24x7 (or equivalent)
- APM tools (AppDynamics, Dynatrace, New Relic, etc.)
- Strong ability to interpret logs, metrics, and traces
Programming & Automation
- Proficiency in:
- C#, PowerShell, or Python
- Experience building automation tools and operational scripts
Preferred Qualifications
- Experience with CI/CD pipelines and release automation
- Knowledge of Docker and Kubernetes
- Experience with infrastructure-as-code practices
- Familiarity with resilience engineering and chaos testing
- Cloud certifications (AWS / Azure) are a plus
Core Competencies
- Strong systems thinking and problem-solving ability
- Deep debugging and analytical skills
- Ability to lead during high-pressure production incidents
- High ownership and accountability mindset
- Effective cross-team communication and collaboration
Success Measures
- Improved system uptime and SLO adherence
- Reduced:
- MTTR (Mean Time to Resolve)
- MTTD (Mean Time to Detect)
- Elimination of recurring incident patterns
- Increased automation and reduced operational toil
- Improved monitoring quality and observability maturity
What This Role Drives
- Highly reliable, scalable, and resilient systems
- Faster and more effective incident resolution
- Reduced operational overhead through engineering
- Stronger alignment between development and production reliability
- Continuous improvement in system performance and stability
What We Expect from You
- You own problems end-to-end, regardless of where they originate
- You fix systems, not just incidents
- You proactively identify risks before they impact production
- You build automation and solutions that scale beyond manual effort
You raise the reliability bar for the entire engineering organization
Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Site Reliability Engineer (Salem)
🏢 Vee Healthtek
📍 Salem