Lead SRE & Support Engineer (Hyderabad)

Lead SRE & Support Engineer (Hyderabad)

19 Aug
|
Providence College of Engineering
|
Hyderabad

19 Aug

Providence College of Engineering

Hyderabad

Lead SRE & Support Engineer

We are seeking a highly motivated Lead Site Reliability Engineer (SRE) & Support Engineer (3P) to join the Cybersecurity Product (Software) Engineering Team. This role is responsible for maintaining the reliability, availability, and operational performance of critical cloud-native services running in Microsoft Azure currently, with potential to extend to AWS, GCP Cloud Service Providers. This is a full-time night-shift role (9:00 PM to 6:00 AM IST) supporting US business hours. While office presence requirements can be discussed, candidates must be based in Hyderabad and available to work from the Hyderabad location as needed.

The successful candidate will partner with software engineers, platform engineers, data engineers, and cybersecurity stakeholders to support production services, manage incidents, improve observability, automate repetitive work, and drive continuous operational improvement. The ideal candidate combines strong troubleshooting skills with an ownership mindset and practical experience using AI-assisted engineering tools.

About the Team

The Cybersecurity Product Engineering Team designs, develops, and operates strategic platforms that support security posture visibility, operational intelligence, enterprise risk management, and secure engineering outcomes. The team works at the intersection of cybersecurity, cloud engineering, data engineering, site reliability engineering, and artificial intelligence. Reliability and operational excellence are central to how we deliver secure, resilient, scalable services to global stakeholders.

About the Role

We are seeking a highly motivated Lead Site Reliability Engineer (SRE) Support Engineer (3P) to join the Cybersecurity Product (Software) Engineering Team. This role is responsible for maintaining the reliability, availability, and operational performance of critical cloud-native services running in Microsoft Azure currently, with potential to extend to AWS, GCP Cloud Service Providers.

Technical

- Own and manage DevOps pipelines for the platform across application code, data, RAG, and ML workloads; build and enhance pipelines as needed to improve reliability, automation, and deployment efficiency.
- Provide hands-on Production Support for product with clear, timely incident updates that summarize impact, investigation status, mitigation, and next actions.
- Drive incidents through containment, recovery, validation, closure, and handoff when cross-shift follow-up is required.
- Contribute to root cause analysis and track corrective and preventive actions to completion.




- Participate in a 24x7 support and on-call model as required by business and service needs.

Site Reliability Engineering

- Improve service reliability, resilience, scalability, and operational effectiveness through engineering-led support practices.
- Identify recurring failure patterns, reliability risks, and opportunities to reduce manual operational toil.
- Contribute to Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, and service health reviews.
- Partner with engineering teams on operational readiness, recovery procedures, capacity considerations, and platform hardening.
- Use incident and service-health learnings to recommend durable technical and process improvements.

Azure-Native Observability

- Use Azure Monitor, Application Insights, Log Analytics workspaces, Azure Alerts, and Azure Service Health to monitor and diagnose services.
- Investigate system behavior through metrics, logs, distributed traces, dependency maps, queries, and correlated events.
- Build and improve actionable dashboards, alert rules, health views, and operational reports.
- Tune monitoring and alerting to improve signal quality, reduce alert fatigue, and shorten detection and recovery cycles.
- Contribute to observability standards and consistent telemetry practices across supported services.

Data Platform Operations

- Support the operational reliability of cloud-based data services and analytics workloads.
- Troubleshoot Snowflake connectivity, access, workload, query performance, and operational issues within the scope of the support role.
- Investigate data ingestion, transformation, reporting, and pipeline failures in collaboration with data engineering teams.
- Validate data availability and operational recovery after incidents, releases, or maintenance activities.
- Use SQL to investigate data issues, validate processing outcomes, and support incident diagnosis.

Automation and Operational Excellence

- Develop and maintain Python, PowerShell, or shell-based automation for health checks, evidence gathering, diagnostics, and routine support activities.
- Create reusable utilities and workflow improvements that reduce manual effort and improve response consistency.




- Identify opportunities for safe self-service and self-healing capabilities with appropriate controls and auditability.
- Improve support processes through standardization, measurable outcomes, documentation, and continual learning.
- Contribute operational feedback to backlog prioritization and engineering improvement plans.

AI-Assisted Operations

- Use approved enterprise AI assistants and copilots to accelerate troubleshooting, knowledge retrieval, scripting, documentation, and incident summarization.
- Apply effective prompt engineering techniques to produce explicit, context-aware operational outputs and refine results through validation.
- Use AI assistance to summarize logs and telemetry, detect patterns, propose hypotheses, and organize evidence while independently validating conclusions.
- Use AI-assisted coding tools to draft or improve scripts, tests, queries, and automation with appropriate review and secure coding practices.
- Create and improve runbooks, knowledge articles, incident timelines, and post-incident documentation using AI-assisted workflows.
- Understand foundational AIOps concepts such as anomaly detection, event correlation, alert enrichment, and intelligent triage.
- Protect confidential, personal, security-sensitive, and regulated information when using AI tools, following organizational data-handling requirements.
- Recognize AI limitations, including inaccurate or incomplete outputs, and apply human review before operational use.

Required Qualifications

- Bachelors degree in Computer Science, Information Technology, Engineering, Cybersecurity, or a related technical discipline, or equivalent practical experience.
- 6-9 years of experience in Site Reliability Engineering, Production Support, Application Support, Platform Operations, or Cloud Operations.
- Hands-on experience supporting enterprise-scale production applications and services in Microsoft Azure. With experience/ exposure to other Cloud Service Providers (AWS, GCP).
- Experience with incident response, service restoration, root cause analysis, change support, and production release validation.
- Working knowledge of Azure-native monitoring and observability services.
- Hands-on SQL and Snowflake support or troubleshooting experience.
- Ability to automate operational work using Python, PowerShell, or shell scripting.
- Working knowledge of REST APIs, JSON, authentication and authorization concepts.

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Lead SRE & Support Engineer (Hyderabad)
🏢 Providence College of Engineering
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: lead sre & support engineer (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: lead sre & support engineer (hyderabad) / hyderabad