27 Sep
|
ZecData Technology
|
Pune
27 Sep
ZecData Technology
Pune
Key Responsibilities
Port and adapt existing GPU test procedures to new data-center GPUs and edge/embedded platforms.
Analyze differences in GPU drivers, power/thermal limits, hardware sensors, and available telemetry sources across platforms.
Build and maintain controlled GPU benchmark and workload suites.
Deploy and execute LLM inference and training workloads, including Llama-family models in both quantized and full-precision configurations.
Develop and maintain synthetic workloads such as:
GEMM
Convolution
Compute workloads
Memory-bandwidth workloads
Multi-GPU workloads
Vision models for edge platforms
Ensure GPU workloads are reproducible and deterministic across repeated runs.
Collect and validate hardware telemetry from:
NVIDIA Management Library (NVML)
Data Center GPU Manager (DCGM)
BMC
IPMI
Redfish
Other external/lab-grade measurement equipment
Ensure consistent sampling rates, timestamps, field naming, and clock alignment across telemetry sources.
Diagnose missing, irregular, inconsistent, or inaccurate sensor data and identify root causes.
Automate infrastructure deployment, workload execution, logging, data collection, and cleanup.
Build automated data-quality validation checks for time-series and hardware telemetry data.
Detect missing samples, clock mismatches, telemetry/workload misalignment, and sensors returning invalid or no data.
Ensure collected datasets are stored in a consistent and documented structure.
Maintain complete run metadata, including hardware model, driver version, firmware version, procedure version, and execution schedule.
Document configuration, limitations, and behavior of each supported hardware platform.
Required Skills & Experience
10+ years of overall relevant experience preferred; exceptional candidates with 8+ years may be considered.
Hands-on experience deploying LLM inference and training workloads on GPUs.
Experience working with quantized and full-precision ML models.
Solid experience with NVIDIA GPU environments and GPU infrastructure.
Strong Linux systems administration and troubleshooting skills.
Good understanding of
GPU driver stacks
Linux processes
Process orchestration
Scheduling
Timing behavior
Hands-on experience collecting and analyzing hardware telemetry programmatically.
Experience with telemetry technologies such as:
NVML
DCGM
BMC
IPMI
Redfish
Ability to independently diagnose sensor and sampling issues in hardware time-series data.
Strong understanding of workload reproducibility and controlling sources of non-determinism.
Experience with automation and scripting for infrastructure deployment and workload execution.
Good to Have
Experience developing data-collection pipelines for:
Hardware testing
Hardware qualification
Systems research
Performance benchmarking
Experience with GPU benchmark and stress-testing tools.
Understanding of benchmark methodology, including:
Warm-up
Steady-state execution
Run-to-run variance
Performance consistency
Understanding of GPU power and thermal management.
Familiarity with GPU clock-throttling reasons and related telemetry.
Experience with multi-GPU scaling.
Knowledge of NCCL.
Experience with tensor parallelism and pipeline parallelism.
Experience developing automated data-quality validation for time-series or sensor data.
Key Technical Areas
GPU: NVIDIA Data Center GPUs, Embedded/Edge GPUs
ML Workloads: LLM Inference, LLM Training, Llama-family Models, Vision Models
GPU APIs/Telemetry: NVML, DCGM
Hardware Management: BMC, IPMI, Redfish
Systems: Linux, GPU Drivers, Process Orchestration
Benchmarking: GEMM, Convolution, Compute, Memory Bandwidth, Stress Testing
Multi-GPU: NCCL, Tensor Parallelism, Pipeline Parallelism
Data: Time-Series Telemetry, Data Validation, Dataset Generation
Pay: ₹384,026.28 - ₹1,538,362.00 per year
Work Location: In person
📌 GPU & ML Infrastructure Engineer (Exp 10+ yrs) (Pune)
🏢 ZecData Technology
📍 Pune