Who we are:
At UST, we help the world’s best organizations grow and succeed through transformation. Bringing together the right talent, tools, and ideas, we work with our client to co-create lasting change. Together, with over 30,000 employees in 30+ countries, we build for boundless impact—touching billions of lives in the process. Visit us at .
Job Description
Role Summary
We are seeking an experienced Site Reliability Engineer (SRE) with robust hands-on expertise in cloud infrastructure, platform reliability, automation, observability, and production support. This role focuses on improving service reliability through SLIs/SLOs and error budgets, reducing operational toil through automation, and partnering with global engineering teams to ensure resilient, scalable, and secure platforms.
Key Responsibilities
Reliability Engineering
- Define, measure, and report SLIs, SLOs, and error budgets for critical services.
- Drive service reliability improvements and systematically reduce operational toil through automation.
- Own capacity planning, performance tuning,
and scalability initiatives.
- Lead blameless postmortems, root cause analyses, and corrective action tracking.
Platform Reliability & Automation
- Design and maintain CI/CD pipelines using Azure DevOps and Jenkins.
- Operate and manage Azure cloud infrastructure and Kubernetes platforms.
- Deploy and support containerized applications using Docker, Kubernetes, and Helm.
- Automate infrastructure provisioning and configuration using Terraform and Ansible.
- Manage artifacts and repositories using JFrog Artifactory.
Observability & Production Support
- Implement monitoring, logging, and observability using Prometheus, Grafana, Loki, and OpenTelemetry.
- Provide L2/L3 production support, incident management, troubleshooting, and RCA.
- Participate in on-call rotations supporting critical production services.
- Support PostgreSQL, Redis, and RabbitMQ environments, including high availability, backups, re
📌 Lead SRE (Bengaluru)
🏢 UST
📍 Bengaluru