10 Aug
|
EnterOne
|
India
AI Infrastructure Engineer
Full-Time •Infrastructure • AI/ML Platform
ROLE OVERVIEW
We are looking for an AI Infrastructure Engineer to design and operate the physical and virtualized compute fabric that powers our AI workloads. You will own the full stack — from Cisco UCS rack and blade server provisioning through GPU-accelerated OpenShift deployments — ensuring that our AI engineers have the reliable, high-throughput infrastructure they need to build and serve large language models at scale, including at distributed edge locations.
KEY RESPONSIBILITIES
- Design, provision, and configure Cisco UCS rack servers and UCSX modular blade chassis (X-Series) hosting GPU-dense computing nodes, and deploy Cisco Unified Edge hardware for localized AI inferencing at edge locations using Cisco Intersight.
- Configure and maintain high-density Cisco Nexus switches for backend AI fabrics using Cisco Nexus Dashboard Fabric Controller (NDFC).
- Perform configuration and bare-metal installation of Red Hat OpenShift Container Platform (OCP) with full GPU hardware acceleration support.
- Deploy and manage the NVIDIA GPU Operator and NVIDIA Network Operator on OpenShift, ensuring operator health, version compatibility, and driver lifecycle management.
- Collaborate closely with AI engineers to deploy, schedule, and optimize OpenShift-hosted LLM inference engines for throughput, latency, and resource efficiency.
- Monitor infrastructure health across compute, network, and storage layers; drive root-cause analysis and resolution for incidents affecting AI workloads.
- Document configurations, runbooks, and architecture diagrams; contribute to infrastructure-as-code and automation efforts.
REQUIRED QUALIFICATIONS
- 5+ years of hands-on experience with Cisco UCS and Nexus platforms, including at least 1 year focused on configuration and deployment with Cisco Intersight and NDFC.
- Demonstrated experience provisioning and managing Cisco UCS X-Series or C-Series servers in production environments.
- Proficiency configuring Cisco Nexus switches, including VLANs, VPCs, ECMP, and fabric interconnects for high-bandwidth AI/storage networks.
- Hands-on experience with Red Hat OpenShift Container Platform (OCP) — installation, upgrade, and day-2 operations.
- Familiarity with NVIDIA GPU Operator and Network Operator deployment on Kubernetes/OpenShift.
- Working knowledge of Linux systems administration (RHEL/CoreOS preferred).
PREFERRED QUALIFICATIONS
- Experience deploying AI inference engines (vLLM, TGI, NVIDIA Triton) on GPU-accelerated container platforms.
- Familiarity with Cisco Intersight Workload Optimizer or Intersight Kubernetes Service (IKS).
- Understanding of RDMA, RoCEv2, or InfiniBand networking for GPU-to-GPU communication.
- Experience with infrastructure-as-code tools such as Terraform, Ansible, or ArgoCD.
- Cisco certifications (CCNA, CCNP Data Center, or CCIE Data Center) or Red Hat certifications (RHCSA, RHCE) are a plus.
- Exposure to edge computing architectures and distributed inference workloads.
WHAT WE OFFER
- Competitive salary and equity with a well-funded, quick-moving team.
- Access to cutting-edge GPU clusters, Cisco hardware, and the latest AI infrastructure tooling.
- A high-ownership role at the foundation of our AI platform — your work directly enables model deployment.
- Flexible remote-friendly environment with travel to data center and edge sites as needed.
📌 AI Infrastructure Engineer (India)
🏢 EnterOne
📍 India