Role description
We are seeking a motivated and detail-oriented Junior Site Reliability Engineer (SRE) to join our Cloud Operations team. In this role, you will support the reliability, availability, performance, and operational excellence of cloud-native platforms hosted on Microsoft Azure.
Key Responsibilities
Cloud Operations & Platform Support:
Support the day-to-day operations of Microsoft Azure cloud infrastructure and services.
Assist in maintaining and supporting Azure Kubernetes Service (AKS) settings.
Incident Management & Problem Resolution:
Participate in production incident response, triage, troubleshooting, escalation, and post-incident reviews.
Azure Virtual Machines
Azure Kubernetes Service (AKS)
Kubernetes pods, nodes, services, ingress controllers
Azure Storage services and connectivity
Cloud networking and platform dependencies
Utilize Azure logs, metrics, monitoring dashboards, and diagnostic tools to identify root causes and implement resolutions.
Monitoring & Observability:
Respond to s, investigate anomalies,
and assist in reducing recurring operational issues.
Support performance analysis and system optimization initiatives.
Automation & DevOps:
Assist with Infrastructure as Code (IaC) deployments using Terraform.
Support CI/CD pipeline execution and troubleshooting using Azure DevOps.
Automate repetitive operational tasks using:
Azure Automation Runbooks
Azure Logic Apps
PowerShell and Python scripts
Perform diagnostics, administration, and remediation activities using Azure CLI.
Kubernetes & Workload Management:
Support containerized workloads running on AKS.
Assist with deployment, troubleshooting, and maintenance of:
Kubernetes Deployments
Services
Jobs
CronJobs
Collaborate with engineering teams to improve platform reliability and operational efficiency.
Continuous Improvement:
Learn and adopt Site Reliability Engineering best practices focused on reducing operational toil a
📌 Junior Site Reliability Engineer Pune (India)
🏢 UST
📍 India