Senior Site Reliability Engineer (Bengaluru)

Senior Site Reliability Engineer (Bengaluru)

04 Aug
|
Anicca Data Science Solutions
|
Bengaluru

04 Aug

Anicca Data Science Solutions

Bengaluru

Position Overview

Anicca Data is seeking an experienced Senior Site Reliability Engineer (Senior SRE) to support Microsoft’s cloud-based platforms and enterprise services. The ideal candidate will have strong experience in reliability engineering, Microsoft Azure, Kubernetes, DevOps automation, infrastructure as code, application development, monitoring, and production troubleshooting. The Senior SRE will work closely with software engineering, infrastructure, operations, and product teams to improve system availability, scalability, performance, security, and operational efficiency.

This role requires participation in rotational shifts to provide continuous support for business-critical systems. Key Responsibilities

Design, implement, and maintain highly available, scalable, and reliable cloud services on Microsoft Azure.

Monitor production systems and proactively identify reliability, performance, capacity, and availability risks.

Troubleshoot complex application, infrastructure, network, deployment, and cloud-service issues.

Participate in incident response, root-cause analysis, and post-incident review activities.

Define and track service-level indicators, service-level objectives, availability targets, and operational metrics.

Develop automation scripts and tools to reduce manual operational activities and improve engineering productivity.

Build and maintain CI/CD pipelines using Azure DevOps Services.

Create and manage infrastructure through infrastructure-as-code practices.

Deploy, configure, operate, and troubleshoot containerized applications running on Kubernetes.

Develop and maintain applications, utilities, and automation solutions using C# and object-oriented programming principles.

Write and maintain YAML-based build, release, deployment, and configuration pipelines.

Use Kusto Query Language to analyze telemetry, application logs, operational data, and system performance.

Build monitoring dashboards, operational reports, and automated workflows using Azure services and Microsoft Power Automate.

Perform reliability, integration, regression, performance, and production-readiness testing.

Conduct application and infrastructure reviews to identify reliability gaps and opportunities for automation.





Collaborate with development teams to improve application observability, fault tolerance, resiliency, and deployment practices.

Support release management, setting configuration, production deployments, and rollback activities.

Create technical documentation, operational procedures, troubleshooting guides, and incident-response runbooks.

Participate in rotational shifts, on-call support, and escalation management for critical production systems.

Mentor junior engineers and promote SRE, DevOps, automation, testing, and cloud-engineering best practices.

Required

Qualifications

Bachelor’s or master’s degree in Computer Science, Information Technology, Engineering, or a related discipline.

6+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering, or Production Support.

Strong experience with Microsoft Azure and Azure-hosted cloud services.

Hands-on experience with Kubernetes administration, deployment, monitoring, and troubleshooting.

Strong understanding of site reliability engineering principles, high availability, scalability, resiliency, and disaster recovery.

Experience designing and maintaining CI/CD pipelines using Azure DevOps Services.

Hands-on experience with infrastructure-as-code tools and practices.

Strong programming experience in C# and object-oriented programming.

Experience developing and debugging applications using Visual Studio.

Proficiency in scripting and task automation.

Experience writing and maintaining YAML-based pipeline definitions and configuration files.

Strong working knowledge of Kusto Query Language for log analysis and operational troubleshooting.

Experience with application, infrastructure, integration, and reliability testing.

Strong troubleshooting, analytical, and problem-solving abilities.





Experience working with monitoring, alerting, logging, and observability platforms.

Ability to work effectively in rotational shifts and respond to production incidents.

Strong communication, documentation, collaboration, and stakeholder-management skills.

Preferred

Qualifications

Microsoft Azure, Kubernetes, DevOps, or Site Reliability Engineering certifications.

Experience supporting large-scale Microsoft or enterprise cloud environments.

Experience with Azure Monitor, Application Insights, Log Analytics, and Kusto-based telemetry platforms.

Knowledge of PowerShell, Python, Bash, or similar scripting languages.

Experience with Terraform, Bicep, ARM templates, or other infrastructure-as-code technologies.

Experience developing automation workflows using Microsoft Power Automate.

Understanding of distributed systems, microservices, REST APIs, networking, identity, and cloud security.

Experience with chaos engineering, capacity planning, performance optimization, and disaster-recovery testing.

Familiarity with IT service-management processes, incident management, change management, and problem management.

Core Technical

Skills

Object-Oriented Programming, C#, Visual Studio, Microsoft Azure, Cloud Services, Kubernetes, Site Reliability Engineering, Azure DevOps Services, CI/CD, Infrastructure as Code, YAML, Kusto Query Language, Scripting and Automation, Microsoft Power Automate, Testing and Quality Engineering, Monitoring and Observability, Troubleshooting, Incident Management, Problem Solving Shift Requirements The selected candidate must be willing to work in rotational shifts, including early-morning, evening, night, weekend, or on-call schedules, depending on Microsoft client support and operational requirements.

Desired Candidate

Profile The successful candidate will be a hands-on reliability engineer who combines software development, cloud infrastructure, DevOps automation, and production-support expertise. The candidate should be comfortable handling critical incidents, analyzing complex system behavior, automating repetitive processes, and collaborating with globally distributed Microsoft and engineering teams.

📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Anicca Data Science Solutions
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (bengaluru) / bengaluru