We are seeking an experienced AI Software Stack Test Lead to define, drive, and execute validation strategy for a next-generation AI computational storage platform. This role owns end-to-end verification of the AI software stack, including neural network operators, compiler-generated workloads, runtime execution, memory management, scheduling systems, and model-level performance.
The ideal candidate combines solid software validation expertise with deep knowledge of AI/ML execution pipelines and will lead validation efforts across compiler, runtime, kernel, and hardware teams to ensure correctness, scalability, reliability, and performance.
Key Responsibilities
Validation Strategy & Technical Leadership
- Own and define the end-to-end validation strategy for the AI software stack.
- Establish validation methodologies, coverage metrics, quality gates, and release-readiness criteria.
- Lead validation efforts across operator libraries, runtime systems, execution frameworks, and AI workloads.
- Mentor engineers and drive best practices for automation, performance validation, and debugging.
Functional Validation
- Validate neural network operators and compute kernels
- Validate graph-level execution and end-to-end model inference workflows.
- Verify numerical correctness against reference frameworks such as PyTorch or TensorFlow.
- Ensure behavioral consistency across software releases and hardware revisions.
Runtime & Execution Validation
- Validate memory allocation and movement across host and device environments.
- Execute stress, concurrency, and multi-device validation scenarios.