03 Aug
|
Knitt Labs
|
Thiruvananthapuram
03 Aug
Knitt Labs
Thiruvananthapuram
Job Description
Role Overview
Provide second-level operational support for a containerized Open Stack Private Cloud deployed using Kolla-Ansible. Focus on advanced troubleshooting, service restoration, root cause analysis, platform validation, and coordination with engineering and infrastructure teams to ensure high availability and operational stability.
Key Responsibilities
- Provide second-level support for incidents and service requests escalated by the L1 Operations
team.
- Troubleshoot Open Stack services including Nova, Neutron, Keystone, Glance, Cinder, Horizon,
Placement, Heat, and Octavia.
- Diagnose and resolve issues related to compute, networking, storage, authentication, scheduling,
and API services.
- Troubleshoot Docker containers deployed through Kolla-Ansible and perform service recovery
following operational procedures.
- Perform advanced Linux administration, including troubleshooting CPU, memory, disk, filesystem,
networking, processes, and service failures.
- Analyze applications, containers, and system logs to identify root causes and restore services.
- Validate VM lifecycle operations, image management, volume operations, floating IP connectivity,
and Open Stack API availability.
- Monitor platform health using Grafana, Prometheus, and Alert manager, and investigate
infrastructure alerts.
- Validate the health of controllers, compute, storage, and network nodes after maintenance
activities.
- Perform troubleshooting of Cinder and Ceph storage services, validate volume creation,
attachment, snapshot operations, storage health, and collect diagnostic information for incident
analysis.
- Monitor storage capacity and backend storage health, coordinate with Infrastructure teams for
storage-related incidents, and validate storage services after maintenance activities.
- Perform hardware health validation using server management interfaces such as iDRAC, iLO, or
IPMI, and correlate hardware events with platform issues.
- Collect hardware diagnostic information and coordinate with Infrastructure, Data Center, and
OEM teams for hardware-related incidents.
- Support planned maintenance activities, platform validation, and change implementations.
- Prepare Root Cause Analysis (RCA) reports and update operational documentation and
knowledge base articles.
- Mentor L1 engineers and provide technical guidance during incident resolution.
Technical Skills
- Strong Linux administration (RHEL/Rocky Linux/Ubuntu)
- Good understanding of Open Stack architecture and core services
- Hands-on experience with Docker container operations
- Working knowledge of Kolla-Ansible-based Open Stack environments
- Experience with Grafana, Prometheus, and Alert manager
- Robust understanding of TCP/IP networking, VLANs, VXLAN, DNS, and SSH
- Experience with Open Stack CLI and basic API troubleshooting
- Knowledge of virtualization technologies (KVM/QEMU/libvirt)
- Basic understanding of HA Proxy, MariaDB/Galera, RabbitMQ, and Ceph
- Experience with ITSM tools such as JIRA or Service Now
- Basic shell scripting (Bash/Python) is an added advantage
Soft Skills
- Strong analytical and troubleshooting skills
- Excellent verbal and written communication
- Customer-focused approach
- Ability to work independently in a 24×7 rotational shift environment
- Good documentation and reporting practices
- Strong ownership and accountability
- Team collaboration and mentoring skills
- Effective incident management and escalation discipline
Preferred Certifications
- RHCE (Red Hat Certified Engineer)
- RHCSA (Red Hat Certified System Administrator)
- LFCS (Linux Foundation Certified System Administrator)
- Open Stack Foundation Certification (Preferred)
- Docker Certified Associate (Preferred)
📌 OpenStack Support- Private Cloud Operations (Thiruvananthapuram)
🏢 Knitt Labs
📍 Thiruvananthapuram