07 Aug
|
Advanced Micro Devices (AMD)
|
Bengaluru
07 Aug
Advanced Micro Devices (AMD)
Bengaluru
Lead DCGPU Performance Engineer
THE ROLE:
AMD is looking for an outstanding technical contributor to drive performance measurement and characterization of Data Center GPU (DCGPU) systems for AI workloads.
This role focuses on producing accurate, repeatable, and trustworthy performance data across a wide range of AI workloads, platforms, and configurations. The engineer will establish robust measurement methodologies and ensure consistency across environments, enabling reliable performance insights for engineering, product, and business decisions.
THE PERSON:
As a highly detail-oriented and data-driven DCGPU Performance Engineer, you will specialize in performance measurement at system and workload levels, ensuring that results are reproducible, comparable, and representative of real-world behavior.
You will define and enforce best practices for performance measurement across AI workloads, including training and inference, while accounting for system variability, configuration differences, and evolving software stacks. You are expected to become an expert user of internal performance tools and workflows, enabling efficient data collection and high-quality reporting.
The ideal candidate combines deep technical expertise with a strong sense of rigor and discipline in experimentation, validation, and reporting. You are expected to question results, validate assumptions, and continuously improve measurement infrastructure through close collaboration with tools teams.
KEY RESPONSIBILITIES:
Performance Measurement & Characterization
- Measure performance of DCGPU systems across AI workloads (training, inference, microbenchmarks)
- Ensure accurate capture of key metrics such as throughput, latency, efficiency, and scaling behavior
- Validate performance across different system configurations, software stacks, and runtime environments
Reproducibility & Methodology
- Define and enforce best practices for reproducible performance measurement
- Ensure experiments are repeatable across systems, teams, and time
- Establish controls for variables such as software versions, system configuration, and workload parameters
- Develop standardized methodologies for fair and consistent comparisons
Benchmarking & Workload Execution
- Execute AI workloads (LLMs, training, inference) with well-defined configurations
- Ensure consistency in workload setup, execution, and reporting across runs
- Maintain benchmark definitions and configuration baselines
- Support internal and competitive benchmarking efforts
Data Accuracy & Validation
- Cross-check results for anomalies, inconsistencies, and measurement errors
- Validate data using multiple methods (profiling tools, logs, counters, independent runs)
- Identify sources of measurement noise and variability and mitigate them
- Ensure published results are reliable, defensible, and aligned with methodology
Tooling & Measurement Infrastructure
- Develop and enhance tools for performance measurement, logging, and reporting
- Build automation for workload execution, result collection, and validation
- Enable standardized output formats for consistent analysis and reporting
- Improve measurement workflows to increase efficiency and reliability
Tools Expertise & Feedback Loop
- Become an expert user of internal performance tools to collect metrics, analyze results, and generate reports
- Work closely with tools and infrastructure teams to enable efficient and scalable measurement workflows
- Provide actionable feedback to tools teams to improve automation, usability, performance, and coverage
- Help drive adoption of standardized tools and workflows across the organization
Cross-Functional Collaboration
- Work with performance engineers, software teams, and system teams to align on measurement practices
- Provide trusted performance data to architecture, product, and business stakeholders
- Support performance deep dives and root-cause investigations when discrepancies arise
Reporting & Insights
- Generate clear, structured performance reports with documented methodology
- Ensure performance results are communicated with appropriate context, assumptions, and limitations
- Enable decision-making through reliable and high-quality data
PREFERRED EXPERIENCE:
- 8 12+ years of experience in performance measurement, benchmarking, or system characterization
- Strong understanding of performance measurement of complex SoCs (GPUs is a plus) and AI workloads (training and inference)
- Experience with performance benchmarking methodologies and reproducibility practices
- Hands-on experience with profiling and measurement tools (rocprof, Nsight, perf, etc.)
- Experience with large-scale AI workloads (LLMs, distributed training, inference serving)
- Familiarity with system-level performance variability and benchmarking challenges
- Programming experience in Python, C/C++, or scripting for automation
- Robust analytical mindset with attention to detail and data validation
- Experience building or maintaining benchmarking frameworks or infrastructure
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Lead DCGPU Performance Engineer (Bengaluru)
🏢 Advanced Micro Devices (AMD)
📍 Bengaluru