16 Sep
|
Esconet Technologies
|
Delhi
16 Sep
Esconet Technologies
Delhi
Company: Esconet Technologies Location: Delhi/NCR – Okhla Phase 1, New Delhi (On-site) Experience: 5+ Years About the Role Esconet Technologies is looking for a highly motivated HPC System Integrator/Administrator with a strong passion for cluster administration, system integration, and validation of HPC clusters. In this role, you will influence the overall integration, delivery, and management of largely open-source HPC products and solutions — spanning Intel and AMD processors, NVIDIA GPUs, storage, Infini Band, and Linux software. A solid understanding of parallel processing (problem decomposition and work distribution), parallel programming (MPI, Open MP), and computer architecture is essential for this role. Minimum Qualifications Bachelor's degree in Computer Science, Computer Engineering, Computational Science, equivalent mathematical sciences, or a related field, with 3+ years of relevant work experience; or an equivalent combination of education, training, and experience 3+ years of experience with software development in Linux 3+ years of experience with HPC clusters and systems integration Ability to manage the AI stack up to the framework level Experience with GPU cluster capabilities and storage integration (LUSTRE/Bee GFS, cluster storage) Working knowledge of object-based storage, networking, Infini Band, solution designing, and technical documentation Key Responsibilities Install, configure, fine-tune, and troubleshoot multi-vendor, multi-site Linux HPC servers Build and deploy open-source software as well as vendor/partner software Diagnose and resolve system operational issues quickly and effectively Verify full operation of systems, including network, systems,
and storage performance Configure scheduling and queuing systems Assist technical support teams with questions and issues encountered by customers Coordinate with vendors to resolve hardware and software problems Document system administration procedures for routine and complex tasks (wikis) Maintain and monitor the security of HPC systems and servers Design and configure HPC/Kubernetes clusters with NVIDIA GPU support, including MIG (Multi-Instance GPU) and AI Factory deployments Deploy and manage cluster file systems such as Lustre, including OSS (Object Storage Server) and MDS (Metadata Server) node configuration Build and maintain effective working relationships with coworkers, managers, and clients Travel as needed for on-site cluster installation or maintenance (limited) Desired Skills & Experience Building, configuring, and administering Linux distributions — Rocky Linux, Ubuntu, Cent OS, RHEL, and SUSE Windows Server administration Expert knowledge of parallel/distributed file systems such as Lustre or IBM GPFS Robust knowledge of networking and cluster-based distributed computing, including Infini Band switch configuration Experience deploying open-source and commercial HPC platforms HPC cluster architecture design and configuration; AI cluster and Kubernetes cluster experience Strong scripting skills — Bash, Python, Perl; Ansible for automation is a plus Experience building dashboards (e.G., Flask-based) for cluster monitoring is a plus Skilled in diagnosing and debugging complex HPC hardware/software issues, with proposed workarounds Strong troubleshooting and root-cause analysis capabilities Key Skill Areas at a Glance Cluster Application & Design: HPC Cluster Application, HPC Cluster Architecture Designing, HPC Cluster Configuration AI/GPU Infrastructure: AI Cluster, Kubernetes Cluster, NVIDIA GPU, MIG, AI Factory Storage: Cluster File System (Lustre), Cluster Storage (OSS & MDS Nodes), Bee GFS, Object-Based Storage Networking: Infini Band, Infini Band Switch, Networking & Solution Design Automation/Scripting: Python, Bash, Perl, Ansible, Flask Dashboards OS Administration: Rocky Linux, Ubuntu, Cent OS, RHEL, SUSE, Windows Server
📌 Hpc Engineer (Delhi)
🏢 Esconet Technologies
📍 Delhi