24 Sep
|
Majid Al Futtaim
|
Gurugram
24 Sep
Majid Al Futtaim
Gurugram
Site Reliability Engineer ( SRE I) | MAF Retail
About us ::
Majid Al Futtaim is an Emirati-owned, diversified lifestyle conglomerate operating across the Middle East, Africa and Asia. The Group started from one man’s vision to transform the face of shopping, entertainment, and leisure to ‘Create Great Moments For Everyone, Everyday’.
- Founded in 1992, we’re pioneers in shopping malls, communities, retail, and leisure across 15 international markets.
- We operate 25 shopping malls, 13 hotels, and 4 mixed-use communities, including icons like Mall of the Emirates and City Centre Malls.
- Carrefour? Yep, that’s us! We brought Carrefour to the region in 1995 and now run 375+ Carrefour stores across 17 countries, serving 750,000+ customers daily.
But that’s just the beginning. We’re leading the charge in digital innovation, with a strong focus on e-commerce and personalized customer experiences. Here are some of our cool projects:
- Scan & Go, Carrefour NOW, and even Tally the Robot—the first of its kind in the Middle East!
- We’re also driving sustainability and a customer-first culture with cutting-edge digital solutions.
Why should you join us? We’re a family of 250+ in India, and we’re growing rapid. With us, you’ll experience:
- Infinite tech exposure & mentorship
- Live case problem-solving with real impact
- Hackdays and continuous learning through tech talks
- Fun, collaborative work environment that’s more sincere than serious
Key Responsibilities:
- Cloud Infrastructure Management : Manage, deploy, and monitor highly scalable and resilient infrastructure using Microsoft Azure .
- Containerization & Orchestration : Design, implement, and maintain Docker containers and Kubernetes clusters for microservices and large-scale applications.
- Automation : Automate infrastructure provisioning, scaling, and management using Terraform and other Infrastructure-as-Code (IaC) tools.
- CI/CD Pipeline Management :
Build and maintain CI/CD pipelines by using Github Actions for continuous integration and delivery of applications and services, ensuring high reliability and performance.
- Monitoring & Incident Management : Implement and manage monitoring, logging, and alerting systems to ensure system health, identify issues proactively, and lead incident response for operational challenges.
- Kafka & Apigee Management : Manage and scale Apache Kafka clusters for real-time data streaming and Apigee for API management.
- Scripting & Automation : Utilize scripting languages (e.g., Python , Bash , PowerShell , etc.) to automate repetitive tasks, enhance workflows, and optimize infrastructure management.
- Collaboration : Work closely with development teams to improve application architecture for high availability, low latency, and scalability.
- Capacity Planning & Scaling : Conduct performance tuning and capacity planning for cloud and on-premises infrastructure.
- Security & Compliance : Ensure security best practices and compliance requirements are met in the design and implementation of infrastructure and services.
Required Skills & Qualifications:
- Experience: 2+ years of relevant experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, Cloud Operations, or a similar role.
- Cloud Expertise: Hands-on experience with Microsoft Azure services, including AKS, storage, networking, identity, and security fundamentals.
- Containerization & Orchestration: Working experience with Docker and Kubernetes, including deployments, services, pods, configuration, logs, and troubleshooting.
- Infrastructure as Code (IaC):
Working knowledge of Terraform for provisioning and managing cloud infrastructure; familiarity with reusable modules and remote state is desirable.
- CI/CD: Experience working with CI/CD pipelines using GitHub Actions, Azure DevOps, or similar tools.
- Messaging: Basic understanding of Kafka and/or cloud messaging services such as Azure Service Bus and Event Hubs.
- API Management: Familiarity with API gateways and API management platforms such as Apigee, Azure API Management, or Kong is desirable.
- Scripting & Automation: Working knowledge of Python, Bash, PowerShell, or similar scripting languages for operational automation.
- Monitoring & Logging: Experience with observability tools such as New Relic, Prometheus, FluentBit etc..
- Version Control: Good working knowledge of Git and source-control platforms such as GitHub or Bitbucket.
- Problem-Solving: Strong troubleshooting fundamentals and the ability to investigate infrastructure and application issues using logs, metrics, traces, and system data.
- Collaboration & Communication: Ability to work effectively in a cross-functional environment, document operational knowledge, and communicate technical issues clearly.
PREFERRED SKILLS
- Kubernetes & Container Operations: Good understanding of Kubernetes fundamentals including Pods, Deployments, Services, ConfigMaps, Secrets, resource requests/limits, health probes, HPA, ingress, and basic troubleshooting of containerized workloads.
- Monitoring & Observability: Familiarity with monitoring and observability concepts including metrics, logs, dashboards, alerting, and basic distributed tracing using tools such as New Relic, Prometheus, Grafana, FluentBit etc ..
- SRE & Incident Management: Understanding of core SRE practices including SLIs/SLOs, availability and reliability monitoring, incident response, on-call operations, runbooks, root-cause analysis, and post-incident reviews .
📌 Site Reliability Engineer I (Gurugram)
🏢 Majid Al Futtaim
📍 Gurugram