Job Summary
We are seeking a Linux L3 Engineer to join Linux Operations team. This is a hands-on technical leadership role responsible for managing, supporting, and optimizing Linux servers in production environments across hosted, cloud (AWS/Azure), and remote infrastructure settings. You will drive operational excellence, incident resolution, automation, and continuous improvement for critical business systems.
Key Responsibilities
Incident & Problem Management
- Resolve high-priority incidents and major incidents (MIs) with end-to-end ownership and minimal escalation
- Perform advanced root-cause analysis and implement permanent, sustainable fixes
- Lead critical incident response, stabilization, and recovery efforts under pressure
- Provide explicit technical direction and minimize business impact during outages
Ownership & Accountability
- Take end-to-end ownership of service requests, incidents, and operational tasks
- Ensure proper validation, documentation, and closure of all work items
- Track MTTR (Mean Time to Resolution) and drive continuous improvement
- Maintain high availability and compliance standards
Process Excellence & SOP Development
- Create, review, and maintain Standard Operating Procedures (SOPs) for recurring operational tasks
- Identify and document gaps in procedures and driving standardization across teams
- Support audit and compliance requirements through documented processes
- Contribute to operational runbooks and disaster recovery procedures
Automation & Infrastructure as Code
- Develop Ansible playbooks, shell scripts, and automation to reduce manual effort
- Support infrastructure-as-code initiatives and configuration management
- Identify repetitive tasks and drive automation to improve reliability and efficiency
- Support deployment and orchestration initiatives (Docker, Kubernetes, orchestration tools)
Change & Risk Management
- Execute changes within ServiceNow/change management frameworks
- Perform impact analysis and risk assessment for infrastructure changes
- Coordinate with stakeholders during maintenance windows and service outages
- Validate blackout configurations and alert suppression during changes
Mentorship & Team Leadership
- Act as a technical reference point and mentor for junior engineers
- Guide team members during complex troubleshooting and escalations
- Share knowledge through documentation, training, and collaborative problem-solving
- Contribute to team capability development and cross-training
Client & Stakeholder Management
- Communicate effectively with internal teams and external clients
- Provide clear updates on incident status and resolution timelines
- Build trust through reliability, transparency, and technical credibility
- Support SLA compliance and service level agreements
Required Qualifications
Technical Skills
- 8+ years of production Linux/Unix administration and support experience
- Deep knowledge of RHEL, CentOS, Ubuntu (including LTS versions and lifecycle management)
- Advanced troubleshooting skills: kernel, networking, storage,
performance tuning
- Experience with patching, security updates, and vulnerability remediation
- Strong shell scripting (Bash) and Ansible automation experience
- Proficiency with systemd, RPM/APT package management, and kernel modules
- Hands-on experience with SSH, sudo, LDAP/Active Directory authentication
- Solid understanding of Linux security (SELinux, AppArmor, firewalls, file permissions)
Cloud & Infrastructure
- Experience with AWS EC2, GCP and Azure VMs in production environments
- Understanding of managed services and infrastructure as a service (IaaS)
- Experience with DNS, routing, load balancing concepts
Monitoring & Logging
- Experience with Datadog, Azure Monitor, Prometheus, or similar monitoring platforms
- Ability to configure alerts, dashboards, and log aggregation
- Understanding of metrics, logs, traces (observability fundamentals)
Operational Skills
- Strong incident response and problem-solving under pressure
- Change management and risk assessment experience
- Documentation and SOP writing skills
- Ability to work in 24/7 on-call rotations (if required)
Soft Skills
- Excellent communication and technical documentation abilities
- Strong time management and prioritization in multi-tasking environment
- Team player with mentoring mindset
- Detail-oriented with focus on quality and compliance
Shift Timings
24/7 rotational shift and on-call support
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr Linux Systems Engineer (Pune)
🏢 Ensono
📍 Pune