Role description
We are seeking a motivated and detail-oriented Junior Site Reliability Engineer (SRE) to join our Cloud Operations team. In this role, you will support the reliability, availability, performance, and operational excellence of cloud-native platforms hosted on Microsoft Azure.
Key Responsibilities
Cloud Operations & Platform Support:
- Support the day-to-day operations of Microsoft Azure cloud infrastructure and services.
- Assist in maintaining and supporting Azure Kubernetes Service (AKS) environments.
Incident Management & Problem Resolution:
- Participate in production incident response, triage, troubleshooting, escalation, and post-incident reviews.
- Azure Virtual Machines
- Azure Kubernetes Service (AKS)
- Kubernetes pods, nodes, services, ingress controllers
- Azure Storage services and connectivity
- Cloud networking and platform dependencies
- Utilize Azure logs, metrics, monitoring dashboards, and diagnostic tools to identify root causes and implement resolutions.
Monitoring & Observability:
- Respond to s, investigate anomalies,
and assist in reducing recurring operational issues.
- Support performance analysis and system optimization initiatives.
Automation & DevOps:
- Assist with Infrastructure as Code (IaC) deployments using Terraform.
- Support CI/CD pipeline execution and troubleshooting using Azure DevOps.
- Automate repetitive operational tasks using:
- Azure Automation Runbooks
- Azure Logic Apps
- PowerShell and Python scripts
- Perform diagnostics, administration, and remediation activities using Azure CLI.
Kubernetes & Workload Management:
- Support containerized workloads running on AKS.
- Assist with deployment, troubleshooting, and maintenance of:
- Kubernetes Deployments
- Services
- Jobs
- CronJobs
- Collaborate with engineering teams to improve platform reliability and operational efficiency.
Continuous Improvement:
- Learn and adopt Site Reliability Engineering best practices focused on reducing operational toil a
📌 Junior Site Reliability Engineer (Pune)
🏢 UST
📍 Pune