01 Oct
|
Consultbae India
|
India
01 Oct
Consultbae India
India
GPU & ML Infrastructure Engineer
Experience: 10+ Years
Location: Remote
About the Role
We are looking for an experienced GPU & ML Infrastructure Engineer to build, test, monitor, and maintain reliable infrastructure for GPU-based benchmarking, stress testing, telemetry collection, and machine learning workloads.
The role involves working with NVIDIA data-center GPUs and embedded/edge platforms, deploying ML workloads, collecting hardware telemetry, automating testing processes, and ensuring that the resulting data is accurate, consistent, and reproducible.
The engineer will own the end-to-end testing and data-generation process, from preparing current GPU platforms and deploying workloads to collecting telemetry, validating data, and documenting results.
Key Responsibilities
- Port and adapt existing GPU test procedures to new NVIDIA data-center and edge/embedded platforms.
- Analyze differences in GPU drivers, power/thermal limits, hardware sensors, and telemetry sources across platforms.
- Build and maintain GPU benchmarking, stress-testing, and workload suites.
- Deploy and execute LLM inference and training workloads, including Llama-family models.
- Work with both quantized and full-precision ML models.
- Develop and execute synthetic workloads such as:
- GEMM
- Convolution
- Compute workloads
- Memory-bandwidth workloads
- Multi-GPU workloads
- Vision models for edge platforms
- Ensure GPU workloads are reproducible and consistent across repeated runs.
- Collect and validate hardware telemetry using:
- NVML
- DCGM
- BMC
- IPMI
- Redfish
- Other external/lab-grade measurement equipment
- Ensure consistent sampling rates, timestamps, field names, and clock alignment across telemetry sources.
- Identify and troubleshoot missing, irregular, inconsistent, or inaccurate sensor data.
- Automate infrastructure deployment, workload execution, logging, data collection, and cleanup.
- Build automated data-quality checks for time-series and hardware telemetry data.
- Detect missing samples, clock mismatches, workload/telemetry misalignment, and invalid sensor data.
- Maintain consistent and well-documented datasets.
- Maintain complete run metadata, including:
- Hardware model
- Driver version
- Firmware version
- Procedure version
- Execution schedule
- Document platform configurations, limitations, and hardware behavior.
Required Skills & Experience
- 10+ years of relevant experience preferred; exceptional candidates with 8+ years may be considered.
- Hands-on experience deploying LLM inference and training workloads on GPUs.
- Experience with quantized and full-precision ML models.
- Strong hands-on experience with NVIDIA GPU environments and infrastructure.
- Strong Linux system administration and troubleshooting skills.
- Good understanding of:
- GPU driver stacks
- Linux processes
- Process orchestration
- Scheduling
- Timing behavior
- Hands-on experience collecting and analyzing hardware telemetry programmatically.
- Experience with telemetry technologies such as NVML, DCGM, BMC, IPMI, and/or Redfish.
- Ability to troubleshoot sensor and sampling issues in hardware time-series data.
- Strong understanding of workload reproducibility and non-determinism.
- Experience with automation and scripting for infrastructure deployment and workload execution.
Good to Have
- Experience developing data-collection pipelines for:
- Hardware testing
- Hardware qualification
- Systems research
- Performance benchmarking
- Experience with GPU benchmarking and stress-testing tools.
- Understanding of benchmark methodology, including:
- Warm-up
- Steady-state execution
- Run-to-run variance
- Performance consistency
- Understanding of GPU power and thermal management.
- Familiarity with GPU clock-throttling reasons and related telemetry.
- Experience with multi-GPU scaling.
- Knowledge of NCCL.
- Knowledge of tensor parallelism and pipeline parallelism.
- Experience developing automated data-quality validation for time-series or sensor data.
Key Technical Skills
GPU: NVIDIA Data Center GPUs, Embedded/Edge GPUs
ML Workloads: LLM Inference, LLM Training, Llama, Vision Models
GPU Telemetry: NVML, DCGM
Hardware Management: BMC, IPMI, Redfish
Systems: Linux, GPU Drivers, Process Orchestration
Benchmarking: GEMM, Convolution, Compute, Memory Bandwidth, Stress Testing
Multi-GPU: NCCL, Tensor Parallelism, Pipeline Parallelism
Data: Time-Series Telemetry, Data Validation, Dataset Generation
Automation: Scripting, Infrastructure Deployment, Workload Automation The candidate should be comfortable owning the complete workflow—from preparing a new GPU platform and deploying workloads to collecting reliable telemetry and producing reproducible, well-documented datasets.
📌 GPU & ML Infrastructure Engineer (India)
🏢 Consultbae India
📍 India