Job Description Site Reliability Engineer (SRE) L2
Position: Site Reliability Engineer (SRE) L2
Experience: 5–6 Years
Location: Noida / Hybrid
Employment Type: Full-Time
Role Summary
We are seeking an experienced Site Reliability Engineer (SRE) – L2 with strong expertise in Microsoft Azure, Dynatrace, and Application Monitoring to support business-critical cloud applications. The ideal candidate will possess hands-on experience in Azure PaaS services, observability platforms, incident response, root cause analysis, and production support. The role requires proactive monitoring, troubleshooting, and ensuring high availability and reliability of cloud-hosted applications.
Key Responsibilities
- Monitor, troubleshoot, and support production applications hosted on Microsoft Azure.
- Perform end-to-end troubleshooting of application and infrastructure incidents.
- Investigate alerts and production issues using Azure Monitor, Application Insights, Log Analytics, and Dynatrace.
- Analyze application logs and telemetry using Kusto Query Language (KQL) and Dynatrace Query Language (DQL).
- Monitor and troubleshoot Azure API Management (APIM), Azure Functions, Service Bus, and Azure-native services.
- Identify the root cause of application failures by tracing requests across APIs, Azure Functions, messaging services, and backend components.
- Configure and manage alerts, dashboards, and monitoring rules to ensure proactive incident detection.
- Utilize Dynatrace features such as Smartscape, Problems & Events, Distributed Tracing, Synthetic Monitoring, and Davis AI for performance analysis.
- Participate in production on-call rotations and provide support for P1/P2 incidents.
- Lead technical troubleshooting during major incidents and collaborate with development and infrastructure teams for timely resolution.
- Prepare detailed Root Cause Analysis (RCA) reports and recommend preventive measures.
- Work closely with DevOps, Application Development, and Cloud Infrastructure teams to improve platform reliability and observability.
- Support continuous service improvements through automation and operational excellence initiatives.
Required Technical Skills
Microsoft Azure
- Azure Monitor
- Application Insights
- Log Analytics
- Kusto Query Language (KQL)
- Azure API Management (APIM)
- Azure Functions
- Azure Service Bus
- Azure Alerts & Action Groups
- Azure Portal
Monitoring & Observability
- Dynatrace
- Problems & Events Feed
- Smartscape
- Distributed Tracing
- Synthetic Monitoring
- Dynatrace Query Language (DQL)
Alternative tools (acceptable):
- Current Relic
- Datadog
Incident Management
- Production Support
- P1/P2 Incident Handling
- Major Incident Management
- Root Cause Analysis (RCA)
- Problem Management
- SLA Management
- On-call Support
Technical Knowledge
- REST APIs
- HTTP/HTTPS
- JSON
- OAuth/JWT Authentication
- API Request/Response Flow
- Azure Functions Lifecycle
- Microservices Architecture
- Event-driven Architecture
- Messaging using Azure Service Bus
Preferred Skills
- Understanding of CI/CD pipelines
- Azure DevOps
- Git
- PowerShell or Bash scripting
- Basic networking (DNS, TCP/IP, Load Balancer)
- ITIL Foundation
If Interested kindly share your resume at
[email protected]
📌 Site Reliability Engineer (Noida)
🏢 NLB Services
📍 Noida