GPU & ML Infrastructure Engineer (India)

GPU & ML Infrastructure Engineer (India)

01 Oct
|
Consultbae India
|
India

01 Oct

Consultbae India

India

GPU & ML Infrastructure Engineer

Experience: 10+ Years

Location: Remote

About the Role

We are looking for an experienced GPU & ML Infrastructure Engineer to build, test, monitor, and maintain reliable infrastructure for GPU-based benchmarking, stress testing, telemetry collection, and machine learning workloads.

The role involves working with NVIDIA data-center GPUs and embedded/edge platforms, deploying ML workloads, collecting hardware telemetry, automating testing processes, and ensuring that the resulting data is accurate, consistent, and reproducible.

The engineer will own the end-to-end testing and data-generation process, from preparing current GPU platforms and deploying workloads to collecting telemetry, validating data, and documenting results.

Key Responsibilities

- Port and adapt existing GPU test procedures to new NVIDIA data-center and edge/embedded platforms.
- Analyze differences in GPU drivers, power/thermal limits, hardware sensors, and telemetry sources across platforms.
- Build and maintain GPU benchmarking, stress-testing, and workload suites.
- Deploy and execute LLM inference and training workloads, including Llama-family models.
- Work with both quantized and full-precision ML models.
- Develop and execute synthetic workloads such as:
- GEMM
- Convolution
- Compute workloads
- Memory-bandwidth workloads
- Multi-GPU workloads
- Vision models for edge platforms

- Ensure GPU workloads are reproducible and consistent across repeated runs.
- Collect and validate hardware telemetry using:

- NVML
- DCGM
- BMC
- IPMI
- Redfish
- Other external/lab-grade measurement equipment

- Ensure consistent sampling rates, timestamps, field names, and clock alignment across telemetry sources.




- Identify and troubleshoot missing, irregular, inconsistent, or inaccurate sensor data.
- Automate infrastructure deployment, workload execution, logging, data collection, and cleanup.
- Build automated data-quality checks for time-series and hardware telemetry data.
- Detect missing samples, clock mismatches, workload/telemetry misalignment, and invalid sensor data.
- Maintain consistent and well-documented datasets.
- Maintain complete run metadata, including:

- Hardware model
- Driver version
- Firmware version
- Procedure version
- Execution schedule

- Document platform configurations, limitations, and hardware behavior.

Required Skills & Experience
- 10+ years of relevant experience preferred; exceptional candidates with 8+ years may be considered.
- Hands-on experience deploying LLM inference and training workloads on GPUs.
- Experience with quantized and full-precision ML models.
- Strong hands-on experience with NVIDIA GPU environments and infrastructure.
- Strong Linux system administration and troubleshooting skills.
- Good understanding of:
- GPU driver stacks
- Linux processes
- Process orchestration
- Scheduling
- Timing behavior

- Hands-on experience collecting and analyzing hardware telemetry programmatically.
- Experience with telemetry technologies such as NVML, DCGM, BMC, IPMI, and/or Redfish.




- Ability to troubleshoot sensor and sampling issues in hardware time-series data.
- Strong understanding of workload reproducibility and non-determinism.
- Experience with automation and scripting for infrastructure deployment and workload execution.

Good to Have
- Experience developing data-collection pipelines for:
- Hardware testing
- Hardware qualification
- Systems research
- Performance benchmarking

- Experience with GPU benchmarking and stress-testing tools.
- Understanding of benchmark methodology, including:

- Warm-up
- Steady-state execution
- Run-to-run variance
- Performance consistency

- Understanding of GPU power and thermal management.
- Familiarity with GPU clock-throttling reasons and related telemetry.
- Experience with multi-GPU scaling.
- Knowledge of NCCL.
- Knowledge of tensor parallelism and pipeline parallelism.
- Experience developing automated data-quality validation for time-series or sensor data.

Key Technical Skills

GPU: NVIDIA Data Center GPUs, Embedded/Edge GPUs

ML Workloads: LLM Inference, LLM Training, Llama, Vision Models

GPU Telemetry: NVML, DCGM

Hardware Management: BMC, IPMI, Redfish

Systems: Linux, GPU Drivers, Process Orchestration

Benchmarking: GEMM, Convolution, Compute, Memory Bandwidth, Stress Testing

Multi-GPU: NCCL, Tensor Parallelism, Pipeline Parallelism

Data: Time-Series Telemetry, Data Validation, Dataset Generation

Automation: Scripting, Infrastructure Deployment, Workload Automation The candidate should be comfortable owning the complete workflow—from preparing a new GPU platform and deploying workloads to collecting reliable telemetry and producing reproducible, well-documented datasets.

📌 GPU & ML Infrastructure Engineer (India)
🏢 Consultbae India
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: gpu & ml infrastructure engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: gpu & ml infrastructure engineer (india) / india