30 Jul
|
Eurofins GSC IT DC
|
Bengaluru
30 Jul
Eurofins GSC IT DC
Bengaluru
Job Description
Job Title: Cloud Engineer (SRE) – Azure Platform
Location: Bangalore (Hybrid, 3 Days WFO)
Reports To: SRE Manager
Experience Level: 4 to 8 Years
Job Summary
We are looking for a skilled Site Reliability Engineer (SRE) to join our Cloud Engineering and Operations team. The ideal candidate will be responsible for ensuring high availability, performance, and reliability of our cloud-hosted systems, particularly in Microsoft Azure environments. This role combines software engineering practices with operational excellence to build scalable, automated, and resilient infrastructure. You’ll work closely with developers, security teams, and platform engineers to implement best practices, reduce toil, and proactively manage incidents and risks across production environments.
Key Responsibilities
Kubernetes Platform Engineering
- Operate, maintain, and optimize production Azure Kubernetes Service (AKS) clusters.
- Perform Kubernetes cluster lifecycle management including upgrades, scaling, capacity planning, and patching.
- Troubleshoot production issues involving Kubernetes clusters, networking, storage, and containerized applications.
- Improve Kubernetes platform reliability, performance, availability, and operational efficiency.
- Implement Kubernetes security and governance best practices.
- Support production deployments and collaborate with development teams to ensure application reliability.
Automation & Platform Engineering
- Design and implement automation solutions to eliminate repetitive operational tasks.
- Develop reusable automation frameworks, self-service capabilities, and operational tooling.
- Automate infrastructure provisioning, deployment, configuration, monitoring, and maintenance processes.
- Continuously identify opportunities to improve operational efficiency through engineering and automation.
- Contribute to an Automation First culture across the platform engineering team.
Cloud Platform Operations
- Manage and support Azure cloud infrastructure and platform services.
- Ensure production environments remain secure, highly available, and operationally efficient.
- Perform capacity planning, performance tuning, and platform optimization.
- Support production releases and infrastructure changes.
Monitoring, Reliability & Incident Management
- Design, maintain, and improve platform monitoring, alerting, and observability.
- Respond to production incidents, troubleshoot platform issues, and restore service quickly.
- Perform root cause analysis (RCA) and implement permanent corrective actions.
- Participate in on-call rotations and incident management activities.
- Continuously improve platform reliability using SRE best practices.
Infrastructure & Security
- Implement Infrastructure as Code (IaC)
for Azure platform provisioning.
- Support identity, access management, and security controls across Azure and Kubernetes.
- Ensure infrastructure complies with enterprise security standards and governance requirements.
- Support disaster recovery, backup validation, and business continuity initiatives.
Collaboration & Documentation
- Collaborate with Development, DevOps, Security, Architecture, and Infrastructure teams.
- Create and maintain operational documentation, runbooks, architecture diagrams, and knowledge articles.
- Share technical knowledge and contribute to engineering best practices and continuous improvement initiatives.
Required Skills & Experience
Kubernetes (Mandatory)
- 4–8 years of experience supporting enterprise production environments.
- Strong hands-on experience administering production Kubernetes clusters, preferably Azure Kubernetes Service (AKS).
- Strong understanding of Kubernetes architecture and core components.
- Experience managing Deployments, StatefulSets, DaemonSets, Jobs, CronJobs, ConfigMaps, Secrets, Services, and Ingress.
- Experience with Kubernetes networking, storage, autoscaling, scheduling, and resource management.
- Experience troubleshooting Kubernetes cluster, application, and networking issues.
- Hands-on experience with Helm Charts and Kubernetes manifests.
- Good understanding of Docker and container technologies.
Automation (Mandatory)
- Strong scripting experience using PowerShell and Python.
- Proven experience developing operational automation and self-service solutions.
- Experience automating Azure operations using Azure CLI, Azure SDKs, and REST APIs.
- Strong Infrastructure as Code (IaC) experience using Bicep.
- Experience building and maintaining CI/CD pipelines using Azure DevOps.
- Experience integrating automation with monitoring, alerting, and operational workflows.
Microsoft Azure
- Strong knowledge of Microsoft Azure platform services.
- Experience managing Azure resources in enterprise production environments.
- Hands-on experience with:
- Azure Kubernetes Service (AKS)
- Virtual Machines
- Virtual Networks
- Network Security Groups (NSGs)
- Load Balancers
- Application Gateway
- Azure Firewall
- Storage Accounts
- Azure Key Vault
- App Services
- Azure Monitor
- Log Analytics
- Application Insights
- Azure DNS
- VPN Gateway
- Private Endpoints
- Good understanding of Azure networking and hybrid connectivity concepts.
Monitoring & Observability
- Experience implementing monitoring and alerting using Azure Monitor, Log Analytics, and Application Insights.
- Experience with Prometheus and Grafana is preferred.
- Experience creating dashboards, alerts, operational metrics, and health monitoring.
Identity & Security
- Strong understanding of Microsoft Entra ID (Azure AD).
- Experience implementing Azure RBAC and Kubernetes RBAC.
- Experience managing Managed Identities, Service Principals, and Workload Identity.
- Experience using Azure Key Vault for secrets management.
- Good understanding of cloud security best practices and platform hardening.
SRE & Operations
- Experience supporting production cloud platforms with high availability requirements.
- Strong troubleshooting and analytical skills.
- Experience with incident management, root cause analysis, and postmortems.
- Familiarity with SRE principles including reliability engineering, operational excellence, and continuous improvement.
- Experience with capacity planning, performance optimization, and cost optimization.
Preferred Qualifications
- Microsoft Certified: Azure Administrator Associate
- Microsoft Certified: Azure DevOps Engineer Expert
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
- Microsoft Certified: Azure Solutions Architect Expert (preferred)
Preferred Qualities
- Robust engineering mindset with a passion for automation.
- Hands-on production Kubernetes experience rather than deployment-only exposure.
- Passion for building scalable, resilient, and secure cloud platforms.
- Excellent troubleshooting and problem-solving abilities.
- Strong ownership and accountability.
- Documentation-oriented with attention to operational excellence.
- Self-driven, proactive, and collaborative.
- Continuous learner with interest in cloud-native technologies and platform engineering.
Tech Stack
Cloud
- Microsoft Azure
Container Platform
- Azure Kubernetes Service (AKS)
- Kubernetes
- Docker
- Helm
Automation & Scripting
- PowerShell
- Python
- Azure CLI
- REST APIs
- YAML
Infrastructure as Code
- Bicep
DevOps & Source Control
- Azure DevOps
- Git
Monitoring & Observability
- Azure Monitor
- Log Analytics
- Application Insights
- Prometheus (Preferred)
- Grafana (Preferred)
Identity & Security
- Microsoft Entra ID
- Azure RBAC
- Kubernetes RBAC
- Azure Key Vault
- Managed Identities
- Workload Identity
Engineering Practices
- Site Reliability Engineering (SRE)
- Infrastructure as Code (IaC)
- CI/CD
- Automation Engineering
- GitOps
- Incident Management
- Postmortems
- Continuous Improvement
- Zero Trust Security
📌 Azure Site Reliability Engineer (Bengaluru)
🏢 Eurofins GSC IT DC
📍 Bengaluru