Role Overview :
As a Senior Site Reliability Engineer based in Dubai, you will serve as the backbone of our mission-critical infrastructure, ensuring that our high-scale platforms remain resilient, performant, and secure. You will work closely with cross-functional engineering teams, product managers, and stakeholders to bridge the gap between development and operations, transforming manual processes into automated, self-healing systems. Your daily contributions will directly influence the end-user experience by minimizing downtime and optimizing resource utilization, ultimately driving the scalability and reliability of our business operations in a fast-paced, global market.
Key Responsibilities :
- Architect and maintain robust, scalable cloud infrastructure across AWS, Azure, and GCP to ensure high availability for global customer traffic.
- Design and implement end-to-end CI/CD pipelines that accelerate deployment velocity while maintaining rigorous quality and security standards.
- Establish comprehensive observability frameworks using Prometheus, Grafana, and Datadog to proactively identify and resolve performance bottlenecks before they impact users.
- Manage centralized logging and security analytics via ELK/OpenSearch, New Relic, and Splunk to provide actionable insights into system health and threat detection.
- Lead incident response efforts and conduct blameless post-mortems to continuously improve system architecture and operational resilience.
- Mentor junior engineers on best practices in infrastructure-as-code and site reliability engineering principles to foster a culture of technical excellence.
Required Skillset :
- Demonstrated expertise in managing multi-cloud environments (AWS, Azure, GCP) with a deep understanding of cloud-native architecture and cost-optimization strategies.
- Proven ability to engineer automated CI/CD workflows that streamline development lifecycles and reduce time-to-market for complex software releases.
- Advanced proficiency in monitoring and alerting ecosystems, specifically leveraging Prometheus, Grafana, Datadog, and New Relic to maintain sub-millisecond latency and high uptime.
- Strong analytical skills in log aggregation and data visualization using ELK/OpenSearch and Splunk to troubleshoot distributed systems effectively.
- Exceptional communication skills with the ability to articulate complex technical risks and solutions to non-technical stakeholders and executive leadership.
- A collaborative mindset that thrives in a hybrid work environment, demonstrating the flexibility to support in office operations in Dubai while coordinating with distributed global teams.
- A Bachelors degree in Computer Science, Information Technology, or a related field, complemented by a track record of managing large-scale production environments with 5 - 10 years of professional experience.
📌 Senior Site Reliability Engineer - Cloud Infrastructure (India)
🏢 Good
📍 India