Site Reliability Engineer (Gurugram)

Site Reliability Engineer (Gurugram)

16 Sep
|
FinThrive
|
Gurugram

16 Sep

FinThrive

Gurugram

Role & responsibilities

Site Reliability Engineer / Cloud Engineer

SRE & Reliability Engineering

- Managed production environments ensuring high availability and reliability of cloud-hosted applications
- Led incident response, performed deep root cause analysis, and implemented preventive measures to reduce recurrence
- Improved system resilience through proactive monitoring and performance tuning strategies

Azure Application & Platform Engineering

- Designed and supported application architectures using:
- Azure App Services and App Service Plans
- Azure App Service Workplace v3 (ASEv3) for isolated, high-scale workloads
- Azure Application Gateway (WAF-enabled) for L7 traffic management
- Azure Front Door for global traffic routing and failover
- Implemented secure and scalable cloud networking patterns, optimizing latency and throughput

Automation & Toil Reduction

- Identified repetitive operational tasks and reduced manual effort through automation-first solutions
- Developed automation using:
- Terraform / Bicep / ARM templates
- Azure Automation (Hybrid Workers)
- Azure Functions for event-driven workflows
- Leveraged AI-assisted tools (GitHub Copilot, Copilot) to accelerate scripting and automation development, while ensuring strict validation for enterprise use

Observability & Monitoring

- Built and enhanced observability using:
- Azure Monitor, Application Insights, Log Analytics
- Created KQL-based queries and dashboards for proactive issue detection
- Reduced false alerts by optimizing alert thresholds and improving signal quality

Performance & System Optimization

- Analyzed application performance across distributed systems to identify bottlenecks
- Implemented improvements through:




- Scaling strategies (horizontal & vertical)
- Network optimization (AGW / Front Door tuning)
- Backend service improvements

Collaboration & Engineering Enablement

- Partnered with SRE, CloudOps, and development teams to design resilient systems
- Contributed to runbooks, documentation, and operational standards
- Enabled engineering teams by improving platform reliability and deployment pipelines

Key Achievements
- Reduced manual operational effort by X% through automation initiatives
- Improved system availability to 99.X% by strengthening monitoring and failure handling mechanisms
- Decreased incident resolution time by X% via enhanced observability and streamlined runbooks
- Optimized application performance using Front Door and AGW tuning, reducing latency by X%

Preferred candidate profile Cloud & Platform Engineering

- Microsoft Azure (Preferred)
- Understanding and experience in developing Azure function Apps, Azure logic Apps
- Understanding of event triggers, event hub, service bus.
- Azure Landing zones, Azure Cloud Adoption Framework, Azure Well Architectured Framework
- Application Hosting: App Services, App Service Plans, ASEv3
- Networking: Azure Application Gateway (AGW), Azure Front Door, VNet, NSGs, Load Balancing
- Cloud Architecture: High Availability, Fault Tolerance, Scalability Patterns

Incident Management and RCA
- Incident Management,



P1 troubleshooting, Change Management
- Experienced in leading RCA and representing on the weekly call
- SLA / SLO / Error Budget concepts
- System Performance Optimization & Capacity Planning
- Toil Reduction through Automation

Infrastructure as Code & Automation
- Terraform, Azure Bicep, ARM Templates
- Azure Automation (Hybrid Workers)
- Azure Functions (Serverless automation)
- API-based automation and orchestration

Observability & Monitoring
- Azure Monitor, Log Analytics Workspace, Grafana, Site 24x7 (or similar SaaS based synthetic monitoring tool)
- Application Insights
- KQL (Kusto Query Language)
- Alert tuning and signal-to-noise optimization

AI-Enabled Productivity (Not as Skill)
- Leveraging GitHub Copilot / Microsoft Copilot for:
- Code acceleration and script generation
- Automation development support
- Troubleshooting and log analysis assistance
- Proven track record of workforce optimization leveraging AI tools.
- Applying validation frameworks to ensure secure, accurate, and production-grade outputs

DevOps & Integration
- CI/CD using Azure DevOps
- Deep understanding on version control
- API integrations (REST, Postman, SoapUI)
- Source control and release management
- Bachelors Degree in Computer Science / Engineering or related field

Preferred/Additional Experience
- Experience with microservices and distributed architectures
- Exposure to low-code automation platforms
- Working knowledge of AWS cloud services

Preferred/Additional Certifications
- AZ-104 Azure Administrator
- AZ-700 Designing and Implementing Microsoft Azure Networking Solutions
- AZ-400 Microsoft Certified: DevOps Engineer Expert

📌 Site Reliability Engineer (Gurugram)
🏢 FinThrive
📍 Gurugram

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (gurugram) / gurugram

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (gurugram) / gurugram