GPU & ML Infrastructure Engineer (Exp 10+ yrs) (Pune)

GPU & ML Infrastructure Engineer (Exp 10+ yrs) (Pune)

27 Sep
|
ZecData Technology
|
Pune

27 Sep

ZecData Technology

Pune

Key Responsibilities

Port and adapt existing GPU test procedures to new data-center GPUs and edge/embedded platforms.

Analyze differences in GPU drivers, power/thermal limits, hardware sensors, and available telemetry sources across platforms.

Build and maintain controlled GPU benchmark and workload suites.

Deploy and execute LLM inference and training workloads, including Llama-family models in both quantized and full-precision configurations.

Develop and maintain synthetic workloads such as:

GEMM

Convolution

Compute workloads

Memory-bandwidth workloads

Multi-GPU workloads

Vision models for edge platforms

Ensure GPU workloads are reproducible and deterministic across repeated runs.

Collect and validate hardware telemetry from:

NVIDIA Management Library (NVML)

Data Center GPU Manager (DCGM)

BMC

IPMI

Redfish

Other external/lab-grade measurement equipment

Ensure consistent sampling rates, timestamps, field naming, and clock alignment across telemetry sources.

Diagnose missing, irregular, inconsistent, or inaccurate sensor data and identify root causes.

Automate infrastructure deployment, workload execution, logging, data collection, and cleanup.

Build automated data-quality validation checks for time-series and hardware telemetry data.

Detect missing samples, clock mismatches, telemetry/workload misalignment, and sensors returning invalid or no data.

Ensure collected datasets are stored in a consistent and documented structure.

Maintain complete run metadata, including hardware model, driver version, firmware version, procedure version, and execution schedule.

Document configuration, limitations, and behavior of each supported hardware platform.

Required Skills & Experience

10+ years of overall relevant experience preferred; exceptional candidates with 8+ years may be considered.

Hands-on experience deploying LLM inference and training workloads on GPUs.





Experience working with quantized and full-precision ML models.

Solid experience with NVIDIA GPU environments and GPU infrastructure.

Strong Linux systems administration and troubleshooting skills.

Good understanding of

GPU driver stacks

Linux processes

Process orchestration

Scheduling

Timing behavior

Hands-on experience collecting and analyzing hardware telemetry programmatically.

Experience with telemetry technologies such as:

NVML

DCGM

BMC

IPMI

Redfish

Ability to independently diagnose sensor and sampling issues in hardware time-series data.

Strong understanding of workload reproducibility and controlling sources of non-determinism.

Experience with automation and scripting for infrastructure deployment and workload execution.

Good to Have

Experience developing data-collection pipelines for:

Hardware testing

Hardware qualification

Systems research

Performance benchmarking

Experience with GPU benchmark and stress-testing tools.

Understanding of benchmark methodology, including:

Warm-up

Steady-state execution

Run-to-run variance

Performance consistency

Understanding of GPU power and thermal management.

Familiarity with GPU clock-throttling reasons and related telemetry.

Experience with multi-GPU scaling.

Knowledge of NCCL.

Experience with tensor parallelism and pipeline parallelism.

Experience developing automated data-quality validation for time-series or sensor data.

Key Technical Areas

GPU: NVIDIA Data Center GPUs, Embedded/Edge GPUs

ML Workloads: LLM Inference, LLM Training, Llama-family Models, Vision Models

GPU APIs/Telemetry: NVML, DCGM

Hardware Management: BMC, IPMI, Redfish

Systems: Linux, GPU Drivers, Process Orchestration

Benchmarking: GEMM, Convolution, Compute, Memory Bandwidth, Stress Testing

Multi-GPU: NCCL, Tensor Parallelism, Pipeline Parallelism

Data: Time-Series Telemetry, Data Validation, Dataset Generation

Pay: ₹384,026.28 - ₹1,538,362.00 per year

Work Location: In person

📌 GPU & ML Infrastructure Engineer (Exp 10+ yrs) (Pune)
🏢 ZecData Technology
📍 Pune

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: gpu & ml infrastructure engineer (exp 10+ yrs) (pune) / pune

Subscribe to this job alert:

Get the latest job offers by email for: gpu & ml infrastructure engineer (exp 10+ yrs) (pune) / pune