04 Sep
|
Yotta Data Services Private
|
Navi Mumbai
04 Sep
Yotta Data Services Private
Navi Mumbai
– Hardware Engineer
Experience: 5+ Years
Role Overview:
We are looking for an experienced and hands-on Hardware Engineer to support the operations, maintenance, troubleshooting and lifecycle management of server hardware across our infrastructure. The role will involve working with both AI/GPU-based servers and conventional CPU-based enterprise servers in mission-critical data center environments.
The ideal candidate should have solid experience in server hardware break-fix, fault diagnosis, component replacement, spare-parts management, RMA processes and coordination with OEM/ODM support teams.
Key Responsibilities:
1. Server Hardware Operations & Maintenance
- Perform end-to-end troubleshooting, maintenance and repair of AI/GPU and conventional enterprise servers.
- Ensure high availability and reliability of server hardware deployed across the infrastructure.
- Monitor and identify hardware faults, degradation and recurring failure patterns.
- Perform hardware diagnosis and identify faulty components for replacement.
- Ensure servers are tested and operational after repair, replacement or maintenance activities.
- Work within defined SLAs to ensure timely resolution of hardware-related incidents.
2. Server Hardware Break-Fix
- Handle the complete hardware break-fix lifecycle, including:
- Fault detection and diagnosis
- Faulty component identification
- Spare allocation
- Onsite component replacement
- Server testing and service restoration
- Faulty part removal and tagging
- RMA coordination and closure
- Troubleshoot and replace server Field Replaceable Units (FRUs), including:
- GPUs
- CPUs
- Motherboards
- Memory
- NICs
- Power Supply Units (PSUs)
- Fans
- Storage and other server components
3. AI/GPU and Enterprise Server Support
- Provide hardware support for NVIDIA GPU-based servers and compute nodes.
- Work on high-density AI and HPC server environments.
- Support GPU server platforms, including HGX/DGX architecture and other GPU-based server platforms.
- Perform troubleshooting and replacement of GPU cards and associated server components.
- Support conventional enterprise and CPU-based servers, including:
- General-purpose compute servers
- Application and database servers
- Virtualization servers
- High-performance compute servers
4. Spare Parts & Inventory Management
- Maintain and manage server hardware spares required for break-fix activities.
- Ensure proper receipt, inspection, storage and issue of server components.
- Maintain accurate records of spare inventory and component movement.
- Monitor availability of critical server FRUs and escalate requirements for replenishment.
- Support inventory reconciliation of replaced, repaired and available spare components.
- Ensure critical AI/GPU server components are available as per operational requirements.
5. Faulty Part & RMA Management
- Identify, tag and maintain proper records of faulty or replaced components.
- Follow the defined process for segregation and storage of faulty hardware.
- Coordinate with OEMs/ODMs for raising and tracking RMA cases.
- Prepare faulty components for shipment to designated OEM/ODM service centers.
- Track repair and replacement status of faulty components.
- Ensure repaired or replacement components are received and appropriately updated in the inventory.
- Maintain accurate documentation to avoid unaccounted or misplaced components.
6. OEM/ODM & Vendor Coordination
- Coordinate with OEMs, ODMs and hardware service partners for technical support and issue resolution.
- Follow up on delayed parts, replacement requests and unresolved hardware issues.
- Support warranty and service-related activities.
- Escalate critical or recurring hardware failures to the relevant internal and external teams.
- Coordinate with central engineering teams and onsite support teams for timely issue resolution.
7.
Incident Management & Documentation
- Respond to server hardware incidents within defined response and resolution timelines.
- Maintain detailed records of hardware faults, repairs, replacements and RMA activities.
- Document troubleshooting steps and resolutions for recurring issues.
- Follow established operational processes, escalation procedures and SLAs.
- Participate in shift or standby support for critical 24x7 environments, as required.
Required Skills & Experience:
- 5+ years of experience in server hardware operations, data center hardware support or enterprise server infrastructure.
- Strong hands-on experience in server hardware troubleshooting and break-fix operations.
- Strong understanding of server hardware architecture and components.
- Experience in diagnosing and replacing server components such as GPUs, CPUs, motherboards, memory, NICs, PSUs, fans and other FRUs.
- Experience with spare-parts management and hardware inventory.
- Experience in faulty part handling and RMA management.
- Experience coordinating with OEMs, ODMs and hardware service partners.
- Understanding of hardware incident management and SLA-driven support environments.
- Experience working in 24x7 mission-critical data center or enterprise infrastructure environments.
- Ability to troubleshoot issues independently and coordinate with multiple technical teams.
Preferred Skills:
Experience with one or more of the following will be an added advantage:
- NVIDIA GPU-based servers and compute infrastructure
- NVIDIA H100, H200, B200, B300, GB200 or GB300 platforms
- NVIDIA HGX/DGX architecture
- AI/HPC server environments
- High-density compute infrastructure
- Supermicro
- ASUS
- Gigabyte
- Dell
- HPE
- GPUaaS, AI Cloud or large-scale AI data center environments
Educational Qualification:
Bachelor's degree or diploma in Computer Science, Electronics, Electrical Engineering, Information Technology, or a related technical discipline.
Relevant hardware, server or OEM certifications will be an added advantage.
📌 Hardware Engineer (Navi Mumbai)
🏢 Yotta Data Services Private
📍 Navi Mumbai