01 Oct
|
Paramaah It Services
|
Bengaluru
01 Oct
Paramaah It Services
Bengaluru
: OpenShift L3 Eng (with OpenShift AI Experience)
Position Title
Senior Systems Engineer / L3 Support Engineer Red Hat OpenShift & OpenShift AI
Experience Required
12 - 15 Years (with at least 4 - 5 years focused on OpenShift WITH AI platform engineering)
Location
Remote- with short-duration travel requirements to customer locations overseas (EMEA)
Employment Type
Full time
About the Role
We are looking for an experienced OpenShift L3 Engineer to join our Platform Engineering / Infrastructure team. This role requires deep expertise in designing, deploying, and troubleshooting Red Hat OpenShift Container Platform (OCP) environments at scale, along with hands-on experience in OpenShift AI (formerly Red Hat OpenShift Data Science) to support MLOps and AI/ML workload deployments. The ideal candidate will act as the highest-level escalation point for platform issues, drive automation initiatives, and collaborate with architecture, security, and data science teams to build robust, production-grade container and AI platforms.
Key Responsibilities
OpenShift Platform Engineering (L3 Support)
- Serve as the final escalation point (L3) for complex OpenShift cluster issues, performing deep root-cause analysis (RCA) on platform, networking, storage, and application-layer failures.
- Design, deploy, and manage highly available OpenShift clusters (on-prem, bare metal, VMware, and/or hybrid cloud AWS/Azure/GCP ROSA/ARO).
- Perform end-to-end cluster lifecycle management: installation (IPI/UPI), upgrades, scaling, node management, and disaster recovery.
- Manage and troubleshoot core OpenShift components: etcd, API server, SDN/OVN-Kubernetes, Ingress/Router, Operators, Machine Config Operator (MCO), and Cluster Autoscaler.
- Own and optimize persistent storage integrations (ODF/Ceph, NFS, CSI drivers) and container networking (Multus, NetworkPolicies, Service Mesh/Istio).
- Implement and maintain platform security: SCCs, RBAC, image scanning, compliance operators, and integration with enterprise IAM/SSO (LDAP, OAuth, Keycloak).
- Drive CI/CD pipeline integration using OpenShift Pipelines (Tekton), GitOps (ArgoCD), and Jenkins.
- Monitor platform health using Prometheus, Grafana, Alertmanager, and OpenShift Logging (EFK/Loki stack); define SLIs/SLOs and proactive alerting.
OpenShift AI / MLOps Enablement
- Deploy, configure, and manage OpenShift AI components: Data Science Projects, Workbenches (Jupyter), Model Serving (KServe/ModelMesh), Pipelines (Kubeflow Pipelines/Data Science Pipelines), and Distributed Workloads (Ray/CodeFlare).
- Support GPU-enabled workloads: configure and troubleshoot NVIDIA GPU Operator, Node Feature Discovery (NFD), and GPU resource scheduling/quotas.
- Manage model lifecycle infrastructure: model registry integration, model serving runtimes (vLLM, TGI, ONNX, Triton), and inference endpoint scaling/autoscaling.
- Collaborate with Data Science/ML Engineering teams to onboard AI/ML workloads, troubleshoot pipeline failures, and optimize resource utilization (CPU/GPU/memory) for training and inference jobs.
CLI / Platform Operations
- Strong hands-on experience with OpenShift and Kubernetes CLI-based administration.
- CLI-based troubleshooting of clusters, nodes, pods, deployments, services, routes, networking, storage, and application workloads.
- CLI-based management and troubleshooting of OpenShift AI workloads and platform components.
- Strong Linux command-line administration and troubleshooting skills.
• Experience using CLI tools for production incident analysis, diagnostics, and root cause analysis.
Preferred Skills
- OpenShift Virtualization
- OpenShift Data Foundation (ODF)
- Tekton Pipelines
- Red Hat Advanced Cluster Management (ACM)
- Velero/OADP
- Service Mesh (Istio)
- GPU Operator
- AI/ML platform troubleshooting
• CI/CD and GitOps for AI/ML workloads
Additional Key Responsibilities
- Implement monitoring, logging, alerting, and capacity planning using Prometheus, Grafana, Alertmanager, and Loki/EFK.
- Automate operational and platform tasks using Ansible, Bash, and Python.
- Support GitOps deployments and configuration management using Argo CD.
- Support backup, disaster recovery, performance tuning, and platform optimization.
- Perform root cause analysis for production OpenShift and OpenShift AI incidents.
- Prepare technical documentation, runbooks, and SOPs.
- Collaborate with infrastructure, security, DevOps, data science, ML engineering, and application teams.
• Provide L3 production support in a 24x7 environment.
Certifications (Preferred)
- Red Hat Certified Specialist in OpenShift Administration
- RHCSA / RHCE
- Red Hat OpenShift AI-related certification or training
• CKA or CKADITIL-based incident/problem management experience.
Education Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
Soft Skills
- Strong analytical and problem-solving mindset with ability to work under pressure during critical incidents.
- Ability to mentor junior engineers and contribute to knowledge-sharing culture.
- Self-driven with the ability to manage multiple priorities in a fast-paced environment.
- Strong documentation habits and stakeholder communication skills.
What We Offer
- Opportunity to work on cutting-edge container and AI/ML infrastructure platforms.
- Exposure to enterprise-scale OpenShift and OpenShift AI deployments.
- Competitive compensation, certification sponsorship, and continuous learning opportunities.
- Collaborative environment working alongside architecture, DevOps, security, and data science teams.
Paramaah is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
📌 OpenShift AdminLevel (with OpenShift AI & CLI Experience)) (Bengaluru)
🏢 Paramaah It Services
📍 Bengaluru