Site Reliability Engineer (Gurugram)

Site Reliability Engineer (Gurugram)

27 Sep
|
FinThrive
|
Gurugram

27 Sep

FinThrive

Gurugram

Site Reliability Engineer (SRE) – Cloud & Platform Engineering
Professional Summary
Site Reliability Engineer with 1–3+ years of experience in operating and optimizing cloud-native platforms with a strong focus on Azure environments. Working knowledge in building highly available, scalable, and secure systems using Infrastructure as Code (IaC) and automation-first practices.
Experienced in managing application hosting architectures including Azure App Services, ASEv3, Application Gateway (AGW) and Azure Front Door, ensuring high performance and resilience across distributed systems.
Demonstrates an automation mindset by leveraging up-to-date engineering tools and AI-assisted development platforms (e.g., GitHub Copilot, Microsoft Copilot) to accelerate delivery, reduce operational toil, and improve reliability standards — with careful validation of outputs for security and production readiness.
Core Competencies
Cloud & Platform Engineering

Microsoft Azure (Preferred)
Understanding and experience in developing Azure function Apps, Azure logic Apps
Understanding of event triggers, event hub, service bus.
Application Hosting: App Services, App Service Plans, ASEv3

Incident Management and RCA

Incident Management, P1 troubleshooting, Change Management
Experienced in leading RCA and representing on the weekly call
SLA / SLO / Error Budget concepts
System Performance Optimization & Capacity Planning
Toil Reduction through Automation

Infrastructure as Code & Automation

Terraform, Azure Bicep, ARM Templates
Azure Automation (Hybrid Workers)
Azure Functions (Serverless automation)
Azure Logic Apps, Automation through Runbooks
API-based automation and orchestration

Observability & Monitoring

Azure Monitor, Log Analytics Workspace, Grafana, Site 24x7 (or similar SaaS based synthetic monitoring tool)




Application Insights
KQL (Kusto Query Language)
Alert tuning and signal-to-noise optimization

AI-Enabled Productivity (Not as Skill)

Leveraging GitHub Copilot / Microsoft Copilot for:

Automation development support
Troubleshooting and log analysis assistance
Proven track record of workforce optimization leveraging AI tools.

Applying validation frameworks to ensure secure, accurate, and production-grade outputs

DevOps & Integration

CI/CD using Azure DevOps
Understanding on version control
API integrations (REST, Postman, SoapUI)
Source control and release management

Professional Experience
Site Reliability Engineer / Cloud Engineer
SRE & Reliability Engineering

Managed production environments ensuring high availability and reliability of cloud-hosted applications
Understands incident response, performed deep root cause analysis, and implemented preventive measures to reduce recurrence
Improved system resilience through proactive monitoring and performance tuning strategies

Azure Application & Platform Engineering

Designed and supported application architectures using:

Azure App Services and App Service Plans
Azure App Service Environment v3 (ASEv3) for isolated, high-scale workloads
Azure Application Gateway (WAF-enabled) for L7 traffic management
Azure Front Door for global traffic routing and failover

Implemented secure and scalable cloud networking patterns,



optimizing latency and throughput

Automation & Toil Reduction

Identified repetitive operational tasks and reduced manual effort through automation-first solutions
Developed automation using:

Terraform / Bicep / ARM templates
Azure Automation (Hybrid Workers)
Azure Functions for event-driven workflows

Leveraged AI-assisted tools (GitHub Copilot, Copilot) to accelerate scripting and automation development, while ensuring strict validation for enterprise use

Observability & Monitoring

Built and enhanced observability using:

Azure Monitor, Application Insights, Log Analytics

Created KQL-based queries and dashboards for proactive issue detection
Reduced false alerts by optimizing alert thresholds and improving signal quality

Performance & System Optimization

Analyzed application performance across distributed systems to identify bottlenecks
Implemented improvements through:

Scaling strategies (horizontal & vertical)
Network optimization (AGW / Front Door tuning)
Backend service improvements

Collaboration & Engineering Enablement

Partnered with SRE, CloudOps, and development teams to design resilient systems
Contributed to runbooks, documentation, and operational standards
Enabled engineering teams by improving platform reliability and deployment pipelines

Education

Bachelor’s Degree in Computer Science / Engineering or related field

Preferred/Additional Experience

Experience with microservices and distributed architectures
Exposure to low-code automation platforms
Working knowledge of AWS cloud services

Preferred/Additional Certifications

AZ-900 Azure Fundamentals
AZ-104 Azure Administrator
AZ-700 Designing and Implementing Microsoft Azure Networking Solutions

📌 Site Reliability Engineer (Gurugram)
🏢 FinThrive
📍 Gurugram

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (gurugram) / gurugram

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (gurugram) / gurugram