We are seeking a Cloud Site Reliability Engineer SRE to drive the reliability scalability and performance of our cloudbased infrastructure The ideal candidate combines software engineering expertise with advanced systems operations skills to maintain highly available systems while reducing operational toil This role involves automation monitoring capacity planning incident response and cloud platform management across a dynamic distributed environment
As a Cloud SRE you will work closely with Engineering Architecture DevOps and security teams to ensure seamless service experiences for our customers while contributing to platform design and operational efficiency
Position Requirements
Our Engineers play a citical role in the success of our clients and are expected to effectively communicate our recommended solutions in a consultative role for each client Therefore a successful candidate will possess a high degree of selfmanagement personal accountability robust communication skills and teamwork The ability to interact engineer and communicate collaboratively at the highest technical levels with customers vendors partners and all members of staff is required
Key Responsibilities
System Reliability Availability Design and maintain faulttolerant highavailability architectures across AWS Azure and GCP Implement redundancy load balancing and automated failover strategies
Cloud Infrastructure Management Deploy manage and optimize cloud resources using IaC tools such as Terraform Ansible
Monitoring Observability Implement monitoring ing and logging frameworks using Splunk Azure monitor Dynatrace AWS cloud watch or similar to detect and resolve issues proactively
Incident Management Lead realtime incident response rootcause analysis and postmortems to continuously improve uptime and resilience