Company Profile
Our client is a leading global provider of technology driven talent and Tech solutions, delivering technology-driven staffing, business services and managed workforce diverse industries. The organization specializes in helping enterprises build agile, scalable, and future-ready solutions through
AI-enabled talent acquisition, workforce management, and digital HR solutions. With a strong international presence, it supports organizations in optimizing workforce productivity, improving operational efficiency, and accelerating business growth.
The organization serves diverse sectors including BFSI, Retail, Telecommunications, Manufacturing,
Information Technology, Global Capability Centres (GCCs), Healthcare, and other key industries, helping clients accelerate growth, improve operational efficiency, and drive business transformation.
Job Profile: Site Reliability Engineer
Location: Chennai / Hyderabad / Bangalore
Work Mode – Hybrid (3 days onsite mandatory)
Preferred experience: 8+ years The Role:
We are looking for an experienced SRE Engineer with strong expertise in Microsoft Azure, Kubernetes,
Infrastructure as Code, CI/CD, automation, monitoring, and production operations. The candidate will be responsible for managing cloud infrastructure, building and maintaining CI/CD pipelines, supporting containerized workloads, improving platform reliability, and driving automation and modernization initiatives. The role requires strong hands-on experience in production environments, along with ownership of deployments, infrastructure lifecycle, monitoring, troubleshooting, and on-call support.
Responsibilities:
- Own and manage end-to-end CI/CD pipelines, deployments, release processes, and production
releases.
- Design, implement, and maintain Azure cloud infrastructure and manage its complete lifecycle.
- Develop and maintain Infrastructure as Code using Terraform.
- Manage containerized applications and workloads using Kubernetes and Azure Kubernetes
Service (AKS).
- Implement and maintain CI/CD solutions using GitHub, GitHub Actions, and Octopus Deploy.
- Manage production monitoring, alerting, incident response, and platform reliability.
- Work with ELK Stack, Prometheus, and Grafana for monitoring, logging, dashboards, and
observability.
- Support ML platform operations and the Kubernetes ecosystem, including Kubeflow, KServe,
Istio, and EvidenceAI.
- Manage web hosting environments using IIS on Windows Server.
- Work with SQL Server and support database-related infrastructure and operational
requirements.
- Develop automation scripts using Python, Bash, and PowerShell to reduce manual effort and
operational toil.
- Identify and implement opportunities to reduce CI/CD toil through automation.
- Actively contribute to platform modernization, cloud migration, and infrastructure improvement
initiatives.
- Troubleshoot infrastructure, deployment, application, and production issues and drive them to
resolution.
- Participate in production support, on-call rotations, incident management, and root-cause
analysis.
- Ensure high availability, uptime, performance, scalability, and reliability of production instances.
- Establish and improve DevOps best practices around automation, deployment, monitoring,
security, and operational reliability.
- Collaborate with development, QA, architecture,
and operations teams to ensure smooth and
- reliable application delivery.
Must - Have Qualifications:
- Minimum 8 years of hands-on experience in DevOps / Cloud / Infrastructure Engineering.
- Strong hands-on experience with Microsoft Azure.
- Experience supporting ML platforms / MLOps infrastructure.
- Experience with Kubeflow, KServe, Istio, and EvidenceAI.
- Solid experience with Terraform / Infrastructure as Code (IaC).
- Hands-on experience with Kubernetes and AKS.
- Working knowledge of SQL Server.
- Strong experience in CI/CD pipelines using GitHub, GitHub Actions, and/or Octopus Deploy.
- Experience with IIS and Windows Server administration.
- Strong production experience in deployment, release management, troubleshooting, and
production support.
- Experience with monitoring and observability using Prometheus, Grafana, ELK Stack, or
- equivalent tools.
- Experience with Azure migration and platform modernization projects.
- Strong scripting/automation skills in Python, Bash, and/or PowerShell.
- Experience managing cloud infrastructure lifecycle and production environments.
- Experience with on-call support, incident management, uptime, and platform reliability.
- Solid understanding of DevOps automation and CI/CD toil reduction.
Preferred Qualifications:
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Knowledge of networking, security, identity, and access management in Azure.
- Knowledge of SRE / reliability engineering practices.
- Experience with high-availability and scalable production architectures.
- Experience with root-cause analysis, capacity planning, performance optimization, and disaster
recovery.
Application Method
- Apply on LinkedIn or email your resume to:
[email protected]
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 SpeedMart
📍 Bengaluru