Site Reliability Engineer (SRE) – Cloud Platform
About the role
We are looking for a Site Reliability Engineer (SRE) to help us run, scale, and continuously improve the mission-critical Cloud business built around our successful Risk Management solution. You will be part of an agile, product-aligned and globally distributed SRE team, working hand in hand with our Cloud Development organization to ensure our services are reliable, observable, secure, and operated at scale.
This role sits at the intersection of software engineering, cloud infrastructure, and operations.
If you enjoy solving complex technical problems, automating everything that can be automated, and taking responsibility for how systems behave in production, this role is for you.
What you’ll work on
• Ensuring and continuously improving the reliability, availability, performance, and scalability of our cloud services
• Working closely with Cloud Development teams to design for operability, reliability, and resilience
• Managing infrastructure using Infrastructure as Code (Terraform, Ansible)
• Running and improving CI/CD pipelines and deployment processes (Harness)
• Designing and maintaining monitoring, alerting, and observability (CloudWatch and related tools)
• Reducing operational toil through automation (deployments, upgrades, onboarding, incident response)
• Participating in incident handling and post-incident reviews, turning learnings into concrete improvements
• Contributing to standards, runbooks, and reusable patterns used across the global SRE team
What we’re looking for
• Hands-on experience with AWS in real production environments
• Practical experience with Kubernetes (EKS) and Docker
• Experience with Terraform, Harness and Ansible (or similar tooling)
• Familiarity with monitoring and logging concepts (CloudWatch, metrics, alerts)
• Valuable understanding of distributed systems and database technology (ideally Microsoft SQL Server) from an operational standpoint
• Solid tro
📌 Site Reliability Engineer (Pune)
🏢 FIS
📍 Pune