Qualification: Bachelor's degree in CS, IT, Engineering or related field
Role Overview
We are hiring a Platform Engineer to support a 24×7 AI platform operations workplace. The role involves platform monitoring, incident troubleshooting, recovery, escalation and operational support across Kubernetes/OpenShift-based platforms.
Key Responsibilities
• Monitor platform/application dashboards, alerts and operational systems.
• Perform L1 incident detection, troubleshooting and recovery.
• Troubleshoot basic Linux, networking, DNS, ports and connectivity issues.
• Check Kubernetes/OpenShift nodes, pods, deployments and services.
• Review logs and collect diagnostics for escalation.
• Execute approved runbooks, SOPs and GitOps-based recovery procedures.
• Restart/redeploy workloads and verify service recovery.
• Create incident tickets, provide status updates and coordinate with L2/L3 teams.
• Maintain proper shift handovers and operational documentation.
Must-Have Skills(with these skill sets, must have 4+yrs of hands-on real projects experience)
• Platform Monitoring
• Incident Troubleshooting & Recovery
• Kubernetes & Red Hat OpenShift
• Experience with Linux command-line
• Experience with Networking: IP, DNS, ports & connectivity
• Experience with Containers knowledge
• Experience in IT Operations / Infrastructure / Cloud / Application Support
• Ability to follow SOPs, runbooks and technical procedures
• Robust troubleshooting and analytical skills
• Excellent communication & stakeholder management
• Willingness to work in a 24×7 shift environment
• Minimum 2 years of stability in an organization preferred