Staff/Principal Engineer - AI/ML & System-Level Validation (Hyderabad)

Staff/Principal Engineer - AI/ML & System-Level Validation (Hyderabad)

07 Aug
|
Advanced Micro Devices (AMD)
|
Hyderabad

07 Aug

Advanced Micro Devices (AMD)

Hyderabad

Principal Member of Technical Staff ROCm Validation & System-Level Quality

About the Role

We are seeking a Principal Member of Technical Staff (PMTS) to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define

how AMD proves ROCm is ready to ship

from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

What You Will Do

- Own the end-to-end validation architecture for ROCm unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers across multiple GPU generations and server platforms.
- Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
- Lead system-level testing for server nodes multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.
- Drive compute workload validation and characterization LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks establishing reproducible methodology, baselines, and regression tracking.
- Architect the test infrastructure distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration,



result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
- Champion modern, agile quality engineering shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
- Set the bar for GitHub-based quality workflows PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/ repositories and partner upstream projects.
- Lead complex escalation debug partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
- Influence the roadmap work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms

before

tape-in milestones and silicon arrival.
- Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
- Represent ROCm validation externally strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.

Minimum Qualifications

- 12+ years of professional software engineering experience with a strong validation, SDET, or quality-engineering focus, including 5+ years in a senior IC role (Staff/Principal/PMTS or equivalent)



leading validation of complex systems software.
- BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).
- Expert-level Python for test automation and infrastructure; robust C++ for debugging, and extending production code paths under test.
- Deep, demonstrable validation experience in at least two of the following domains:

- GPU compute software stacks (ROCm, CUDA, oneAPI, SYCL)
- Deep-learning frameworks and inference engines (PyTorch, TensorFlow, JAX, Triton, vLLM)
- HPC / parallel runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)
- Linux kernel, GPU drivers, or accelerator firmware
- Distributed systems and large-scale cluster software
- System-level validation for server-class compute nodes multi-GPU, multi-node, fabric-attached environments including stress/stability, soak, fault-injection, and RAS testing.

- Proven, hands-on experience working efficiently in an agentic AI engineering workplace daily, production use of LLM-based coding agents (e.g., Cursor, Claude Code, Copilot Workspace, Codex-class agents) and orchestration frameworks for real engineering work, with demonstrable productivity, quality, or coverage gains attributable to those workflows. Comfort designing prompts, tool/MCP integrations, evaluation harnesses, and guardrails for autonomous and semi-autonomous agents.
- Hands-on experience defining and shipping release qualification programs for software consumed by hyperscalers, OEMs, or other Tier-1 customers.
- Mastery of GitHub at scale for quality engineering PR gating, GitHub Actions, self-hosted runners, required

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

📌 Staff/Principal Engineer - AI/ML & System-Level Validation (Hyderabad)
🏢 Advanced Micro Devices (AMD)
📍 Hyderabad

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: staff/principal engineer - ai/ml & system-level validation (hyderabad) / hyderabad

Subscribe to this job alert:

Get the latest job offers by email for: staff/principal engineer - ai/ml & system-level validation (hyderabad) / hyderabad