22 Aug
|
Net Connect
|
Bengaluru
22 Aug
Net Connect
Bengaluru
Key Responsibilities
- Administer and maintain a leading technology company and AI Factory environments, ensuring optimal uptime and performance.
- Manage compute nodes, GPU clusters, and InfiniBand networks, including hardware like a leading technology company DL380a, DL325, and Cray XD670.
- Perform configuration, patching, version upgrades, and firmware updates across hardware and software layers.
- Proactively monitor system health using tools like DCGM, NetQ, Grafana, and Exivity dashboards.
- Lead root cause analysis and corrective action plans to prevent recurring issues and maintain operational documentation.
- Oversee automation for provisioning, scaling, and patch management using Ansible and AWX.
- Operate a leading technology company Ezmeral Unified Analytics, Data Fabric, and AI Essentials platforms, supporting AI/ML workloads.
- Implement and maintain security protocols using Keycloak for authentication and role-based access control.
Overview
- a leading technology company is a forward-thinking IT services organization, dedicated to providing expert advice, integration, and acceleration of customer outcomes in the digital transformation landscape.
- We are seeking a Senior PCAI AI Factory Expert who will play a critical role in managing and optimizing our AI infrastructure platforms.
- This position calls for a deep understanding of AI, HPC, and GPU-accelerated environments, with the goal of ensuring operational stability and continuous improvement of our large-scale Private Cloud for AI (PCAI) and AI Factory deployments.
- The role is pivotal in overcoming IT complexities to match the speed of business opportunities with the right technology deployments.
- Join our cutting-edge team in Bangalore and help redefine the future by transforming insight into innovation.
- The Senior PCAI AI Factory Expert will be responsible for the administration and management of a leading technology company and AI Factory environments, ensuring optimal performance and uptime.
- The role involves overseeing complex compute nodes, GPU clusters, and network infrastructures.
- The expert will also lead operational monitoring, incident management, and lifecycle management of large-scale AI platforms.
- Additionally, the position requires collaboration with global teams to optimize automation workflows and drive continuous improvement in service response times.
- You will be a key player in ensuring the smooth operation of AI/ML workloads on containerized clusters and supporting advanced AI platform components.
- Your role will require a balance of technical expertise and leadership skills to conduct enablement sessions, manage operational dashboards, and ensure SLA compliance.
- This is an opportunity to work on next-generation AI infrastructure operations, contributing to service innovation and continuous improvement initiatives in AI infrastructure management.
- As part of a leading technology company, you will have the chance to work with cutting-edge technology and be part of a global team managing a leading technology company’s AI Factory and PCAI platforms, supporting large-scale AI workloads.
- We value strong analytical skills and the ability to lead operations improvement initiatives and mentor support engineers.
Requirements
- 8+ years of IT infrastructure administration experience, with 3+ years in AI/HPC or GPU-based environments.
- Proven experience in platform operations, monitoring, and lifecycle management of enterprise-grade AI environments.
- Strong knowledge of NVIDIA GPU stack, InfiniBand NDR, and Spectrum-X switches.
- Hands-on experience with virtualization and container platforms, including vSphere, RHEL, and Kubernetes.
- Expertise in automation tools such as Ansible, AWX, and a leading technology company Performance Cluster Manager.
- Familiarity with AI/ML tools and frameworks including TensorFlow, PyTorch, Spark, and Kubeflow.
- Preferred certifications include a leading technology company ASE / Master ASE, NVIDIA Certified Professional, and RHCE / Kubernetes Administrator.
📌 Senior AI Infra Engineer (Bengaluru)
🏢 Net Connect
📍 Bengaluru