11 Sep
|
Uber AI Solutions
|
India
11 Sep
Uber AI Solutions
India
Freelance Kernel Optimization / Performance Engineering Expert
Location: Remote - INDIA
This role focuses on improving the performance, portability, and correctness of machine learning workloads on TPUs. The engineer will work across high-level ML frameworks, compiler/runtime behavior, distributed execution, and low-level custom kernels to identify bottlenecks and deliver productive implementations.
The ideal candidate is comfortable moving between model-level operations and hardware-aware execution. They can inspect GPU-oriented implementations, reason about their exact semantics, translate them into TPU-compatible forms, and validate both numerical correctness and performance.
Minimum qualifications:
- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
- Strong programming experience in Python and at least one systems language (C++, CUDA, or Pallas).
- Experience developing and optimizing ML accelerator code using JAX/Pallas, PyTorch-XLA, or OpenXLA.
- Experience with distributed machine learning concepts.
Preferred qualifications:
- Master's degree or PhD in Computer Science, Data Science, or a related technical field.
- Knowledge of TPU architecture and performance characteristics.
- Experience converting CUDA or Triton implementations to standard PyTorch or JAX APIs for TPU execution while preserving numerical and gradient semantics.
- Experience with distributed machine learning concepts such as data, tensor, sequence, and expert parallelism; sharding; and collectives including all-reduce, all-gather, reduce-scatter, and all-to-all.
- Kernel Extraction, Profiling, & Optimization: Experience extracting, profiling, and optimizing kernels from frontier open-weight models (e.g., DeepSeek V3/V4, Qwen 2.5/3.8, Kimi K2.5)
- Experience debugging numerical correctness issues involving bf16/fp32 accumulation, quantization scales, softmax/log-sum-exp stability, masking, and backward-pass behavior.
- Experience with performance-oriented open-source ML systems such as vLLM, FlashAttention, JAX/Flax, PyTorch/XLA, or similar frameworks and kernel libraries.
- Experience building reproducible microbenchmarks, performance regression tests, or profiling and diagnostic tooling for accelerator workloads.
Responsibilities
- Develop and optimize high-performance custom kernels and performance-critical ML operations for TPU execution.
- Translate CUDA and Triton kernel behavior into standard PyTorch or JAX implementations that can compile efficiently through OpenXLA/XLA.
- Profile training and inference workloads to identify compute, memory, communication, compilation, and data-movement bottlenecks.
- Debug correctness and performance issues across ML frameworks, compiler/runtime layers, distributed execution, and accelerator kernels.
- Adversarial Agent Benchmarking & Task Formulation: Design authentic, production-derived engineering challenges derived from modern open-weight architectures (e.g., DeepSeek, Qwen, GLM) to stress-test and stump baseline frontier models (Gemini 3.7 Flash, Claude Code), establishing ground-truth solutions and rigorous evaluation rubrics for autonomous agents.
- Design TPU-friendly tiling, sharding, reduction, data-movement, and mixed-precision strategies for common and custom ML operations.
- Validate numerical equivalence, gradient correctness, masking/indexing behavior, and edge cases between source GPU implementations and TPU-targeted implementations.
- Build and maintain benchmarks and diagnostic workflows that measure latency, throughput, utilization, memory usage, communication cost, and performance regressions.
- Analyze transformer and LLM workloads, including attention, MoE, KV-cache, quantized matmuls, and distributed training or serving patterns, to identify optimization opportunities.
- Collaborate with ML systems, compiler, infrastructure, and model engineers to translate workload requirements into efficient accelerator implementations.
📌 System Engineer (India)
🏢 Uber AI Solutions
📍 India