- Strong Linux administration skills
- Good understanding of Windows Server environments
- Hands-on experience with AWS core services (EC2, RDS, VPC, IAM)
- Scripting experience (Python / PowerShell / Bash)
- Experience with monitoring and log analysis tools
- Basic networking knowledge (DNS, Load Balancer, Firewall concepts)
- Understanding of incident and change management processes
Preferred Skills
- Experience with automation tools (Ansible, Terraform)
- Familiarity with CI/CD pipelines
- Exposure to hybrid infrastructure environments
- Knowledge of SLO/SLI concepts
- Experience handling legacy applications
Key Behavioural Competencies
- Solid troubleshooting mindset
- Automation-first thinking
- Ownership of production stability
- Clear documentation discipline
- Ability to work under pressure in production environments
- Collaborative approach across engineering and operations teams
Career Growth Opportunity
This role is part of a strategic SRE transformation initiative. High performers will have opportunities to:
- Lead automation initiatives
- Own reliability domains (Cloud / On-Prem / Observability)
- Transition into Senior SRE / Reliability Lead roles
Why Join This Role
- Be part of a structured reliability transformation
- Move from reactive support to engineering-led operations
- Work across hybrid infrastructure (Cloud + Data Center)
- Contribute to measurable uptime and cost optimization goals
Role Overview
We are building a reliability-driven operations model and are looking for a Site Reliability Engineer (L2) to support and engineer production resilience across a hybrid infrastructure (AWS + On-Prem).
This role is ideal for a strong L2 operations engineer who wants to move beyond reactive support and contribute to automation, reliability engineering,
and production stability at scale.
The primary objective of this role is to reduce MTTR, eliminate repetitive incidents through automation, and improve uptime across business-critical systems.
L2:4-6 years of experience in IT Operations / Production Support / SRE
Key Responsibilities
Production Reliability Incident Management
- Manage L2 production incidents across AWS and on-prem environments
- Perform structured Root Cause Analysis (RCA)
- Drive permanent fixes to recurring issues
- Participate in on-call rotation 247
- Critical Incident Management
Automation Engineering
- Develop scripts (Python / PowerShell / Bash) to automate repetitive tasks
- Contribute to self-healing automation initiatives
- Improve monitoring coverage and reduce alert noise
- Maintain and enhance operational runbooks
Infrastructure Cloud Operations
- Support AWS services (EC2, RDS, VPC, IAM, CloudWatch)
- Manage Linux and Windows servers
- Assist in patching, backup validation, and system health management
- Collaborate on Infrastructure-as-Code initiatives (Terraform/Ansible preferred)
Monitoring Observability
- Work with monitoring tools (Zabbix/ Datadog / AppDynamics / Splunk / Prometheus or similar)
- Tune alerts to reduce false positives
- Create dashboards for service visibility
- Track SLO/SLI metrics for Tier-1 applications
Continuous Improvement
- Identify high-frequency operational issues and propose automation
- Contribute to reliability KPIs (MTTR, uptime, incident recurrence)
- Support capacity planning and performance optimization
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Support Analyst (Pune)
🏢 Data Axle Solutions
📍 Pune
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.