Senior AI Infra Engineer (Bengaluru)

Senior AI Infra Engineer (Bengaluru)

22 Aug
|
Net Connect
|
Bengaluru

22 Aug

Net Connect

Bengaluru

Key Responsibilities

- Administer and maintain a leading technology company and AI Factory environments, ensuring optimal uptime and performance.
- Manage compute nodes, GPU clusters, and InfiniBand networks, including hardware like a leading technology company DL380a, DL325, and Cray XD670.
- Perform configuration, patching, version upgrades, and firmware updates across hardware and software layers.
- Proactively monitor system health using tools like DCGM, NetQ, Grafana, and Exivity dashboards.
- Lead root cause analysis and corrective action plans to prevent recurring issues and maintain operational documentation.
- Oversee automation for provisioning, scaling, and patch management using Ansible and AWX.
- Operate a leading technology company Ezmeral Unified Analytics, Data Fabric, and AI Essentials platforms, supporting AI/ML workloads.
- Implement and maintain security protocols using Keycloak for authentication and role-based access control.

Overview

- a leading technology company is a forward-thinking IT services organization, dedicated to providing expert advice, integration, and acceleration of customer outcomes in the digital transformation landscape.
- We are seeking a Senior PCAI AI Factory Expert who will play a critical role in managing and optimizing our AI infrastructure platforms.
- This position calls for a deep understanding of AI, HPC, and GPU-accelerated environments, with the goal of ensuring operational stability and continuous improvement of our large-scale Private Cloud for AI (PCAI) and AI Factory deployments.
- The role is pivotal in overcoming IT complexities to match the speed of business opportunities with the right technology deployments.
- Join our cutting-edge team in Bangalore and help redefine the future by transforming insight into innovation.
- The Senior PCAI AI Factory Expert will be responsible for the administration and management of a leading technology company and AI Factory environments, ensuring optimal performance and uptime.




- The role involves overseeing complex compute nodes, GPU clusters, and network infrastructures.
- The expert will also lead operational monitoring, incident management, and lifecycle management of large-scale AI platforms.
- Additionally, the position requires collaboration with global teams to optimize automation workflows and drive continuous improvement in service response times.
- You will be a key player in ensuring the smooth operation of AI/ML workloads on containerized clusters and supporting advanced AI platform components.
- Your role will require a balance of technical expertise and leadership skills to conduct enablement sessions, manage operational dashboards, and ensure SLA compliance.
- This is an opportunity to work on next-generation AI infrastructure operations, contributing to service innovation and continuous improvement initiatives in AI infrastructure management.
- As part of a leading technology company, you will have the chance to work with cutting-edge technology and be part of a global team managing a leading technology company’s AI Factory and PCAI platforms, supporting large-scale AI workloads.
- We value strong analytical skills and the ability to lead operations improvement initiatives and mentor support engineers.

Requirements

- 8+ years of IT infrastructure administration experience, with 3+ years in AI/HPC or GPU-based environments.
- Proven experience in platform operations, monitoring, and lifecycle management of enterprise-grade AI environments.
- Strong knowledge of NVIDIA GPU stack, InfiniBand NDR, and Spectrum-X switches.
- Hands-on experience with virtualization and container platforms, including vSphere, RHEL, and Kubernetes.
- Expertise in automation tools such as Ansible, AWX, and a leading technology company Performance Cluster Manager.
- Familiarity with AI/ML tools and frameworks including TensorFlow, PyTorch, Spark, and Kubeflow.
- Preferred certifications include a leading technology company ASE / Master ASE, NVIDIA Certified Professional, and RHCE / Kubernetes Administrator.

📌 Senior AI Infra Engineer (Bengaluru)
🏢 Net Connect
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior ai infra engineer (bengaluru) / bengaluru