Role Description
Who we are:
At UST, we help the world’s best organizations grow and succeed through transformation. Bringing together the right talent, tools, and ideas, we work with our client to co-create lasting change. Together, with over 30,000 employees in 30+ countries, we build for boundless impact—touching billions of lives in the process. Visit us at .
Job Description
Role Summary
We are seeking an experienced Site Reliability Engineer (SRE) with robust hands-on expertise in cloud infrastructure, platform reliability, automation, observability, and production support. This role focuses on improving service reliability through SLIs/SLOs and error budgets, reducing operational toil through automation, and partnering with global engineering teams to ensure resilient, scalable, and secure platforms.
Key Responsibilities
Reliability Engineering
Define, measure, and report SLIs, SLOs, and error budgets for critical services.
Drive service reliability improvements and systematically reduce operational toil through automation.
Own capacity planning, performance tuning,
and scalability initiatives.
Lead blameless postmortems, root cause analyses, and corrective action tracking.
Platform Reliability & Automation
Design and maintain CI/CD pipelines using Azure DevOps and Jenkins.
Operate and manage Azure cloud infrastructure and Kubernetes platforms.
Deploy and support containerized applications using Docker, Kubernetes, and Helm.
Automate infrastructure provisioning and configuration using Terraform and Ansible.
Manage artifacts and repositories using JFrog Artifactory.
Observability & Production Support
Implement monitoring, logging, and observability using Prometheus, Grafana, Loki, and OpenTelemetry.
Provide L2/L3 production support, incident management, troubleshooting, and RCA.
Participate in on-call rotations supporting critical production services.
Support PostgreSQL, Redis, and RabbitMQ environments, including high availability, backups, replica
📌 Lead SRE (Bengaluru)
🏢 UST
📍 Bengaluru