05 Aug
|
HCL Technologies
|
Bengaluru
05 Aug
HCL Technologies
Bengaluru
SME
Experience: 7 to 8 years
Location: Bengaluru, India
Skills: HPC, Site Reliability Engineering, SRE, Linux, Kubernetes, AWS, Azure, GCP, observability, incident management, ticketing systems, Bash, Python
Job Summary
We are looking for an experienced HPC Site Reliability Engineer (SRE) with strong cloud and Kubernetes experience to support and operate large scale High Performance Computing (HPC) environments. The ideal candidate will have hands on experience working on HPC projects, managing Linux-based systems, supporting observability and monitoring workflows, and handling incident and ticket management in production environments. This role requires close collaboration with application, platform, and infrastructure teams to ensure high availability, performance, and reliability of HPC workloads running on AWS, Azure, or GCP platforms.
Key Responsibilities
• Support and operate HPC clusters and workloads in on prem and/or cloud environments.
• Ensure availability, performance, and reliability of HPC systems and services.
• Troubleshoot complex issues related to compute, storage, networking, and schedulers in HPC environments.
• Perform capacity planning, performance tuning, and root cause analysis.
• Manage and operate Kubernetes clusters supporting HPC or compute intensive workloads.
• Work with cloud platforms (AWS / Azure / GCP) for deploying and operating HPC infrastructure.
• Support containerized HPC workloads and hybrid architectures.
• Implement best practices for scalability, resiliency, and cost optimization in cloud environments.
• Work with observability and monitoring tools (e.g., Prometheus, Grafana, cloud-native monitoring).
• Analyze monitoring dashboards, metrics, and alerts to proactively identify issues.
• Perform incident monitoring, ticket analysis, and reporting based on historical and real time data.
• Drive improvements in alerting, dashboards,
and operational visibility using insights from previous roles and production experience.
• Handle production incidents, participate in on-call rotations, and ensure timely resolution.
• Review and analyze incident tickets, identify trends, and propose preventive actions.
• Conduct post-incident reviews (PIRs) and document root cause and corrective actions.
• Collaborate with cross functional teams to resolve complex incidents.
• Develop and maintain automation scripts to reduce manual effort and improve reliability.
• Use Linux shell scripting (Bash) and/or Python for operational automation.
• Automate health checks, reporting, and routine operational tasks for HPC systems.
Skill Requirements
• 7–8 years of experience in HPC, SRE, or Platform Operations roles.
• Proven experience working on HPC projects (compute intensive workloads, clusters, schedulers).
• Strong Linux system administration skills.
• Hands on experience with Kubernetes in production environments.
• Experience with at least one cloud platform like AWS or Azure or GCP.
• Robust understanding of observability, monitoring, and alerting workflows.
• Experience handling incident management and ticketing systems.
• Proficiency in scripting using Bash or Python for automation and operational tasks.
Other Requirements
• 7–8 years of experience in HPC, SRE, or Platform Operations roles.
• Proven experience working on HPC projects (compute intensive workloads, clusters, schedulers).
• Strong Linux system administration skills.
• Hands on experience with Kubernetes in production environments.
• Experience with at least one cloud platform like AWS or Azure or GCP.
• Strong understanding of observability, monitoring, and alerting workflows.
• Experience handling incident management and ticketing systems.
• Proficiency in scripting using Bash or Python for automation and operational tasks.
📌 SME (Bengaluru)
🏢 HCL Technologies
📍 Bengaluru