System Engineer (India)

System Engineer (India)

11 Sep
|
Uber AI Solutions
|
India

11 Sep

Uber AI Solutions

India

Freelance Kernel Optimization / Performance Engineering Expert

Location: Remote - INDIA

This role focuses on improving the performance, portability, and correctness of machine learning workloads on TPUs. The engineer will work across high-level ML frameworks, compiler/runtime behavior, distributed execution, and low-level custom kernels to identify bottlenecks and deliver productive implementations.

The ideal candidate is comfortable moving between model-level operations and hardware-aware execution. They can inspect GPU-oriented implementations, reason about their exact semantics, translate them into TPU-compatible forms, and validate both numerical correctness and performance.

Minimum qualifications:

- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
- Strong programming experience in Python and at least one systems language (C++, CUDA, or Pallas).
- Experience developing and optimizing ML accelerator code using JAX/Pallas, PyTorch-XLA, or OpenXLA.
- Experience with distributed machine learning concepts.

Preferred qualifications:

- Master's degree or PhD in Computer Science, Data Science, or a related technical field.
- Knowledge of TPU architecture and performance characteristics.
- Experience converting CUDA or Triton implementations to standard PyTorch or JAX APIs for TPU execution while preserving numerical and gradient semantics.
- Experience with distributed machine learning concepts such as data, tensor, sequence, and expert parallelism; sharding; and collectives including all-reduce, all-gather, reduce-scatter, and all-to-all.
- Kernel Extraction, Profiling, & Optimization: Experience extracting, profiling, and optimizing kernels from frontier open-weight models (e.g., DeepSeek V3/V4, Qwen 2.5/3.8, Kimi K2.5)
- Experience debugging numerical correctness issues involving bf16/fp32 accumulation, quantization scales, softmax/log-sum-exp stability, masking, and backward-pass behavior.




- Experience with performance-oriented open-source ML systems such as vLLM, FlashAttention, JAX/Flax, PyTorch/XLA, or similar frameworks and kernel libraries.
- Experience building reproducible microbenchmarks, performance regression tests, or profiling and diagnostic tooling for accelerator workloads.

Responsibilities

- Develop and optimize high-performance custom kernels and performance-critical ML operations for TPU execution.
- Translate CUDA and Triton kernel behavior into standard PyTorch or JAX implementations that can compile efficiently through OpenXLA/XLA.
- Profile training and inference workloads to identify compute, memory, communication, compilation, and data-movement bottlenecks.
- Debug correctness and performance issues across ML frameworks, compiler/runtime layers, distributed execution, and accelerator kernels.
- Adversarial Agent Benchmarking & Task Formulation: Design authentic, production-derived engineering challenges derived from modern open-weight architectures (e.g., DeepSeek, Qwen, GLM) to stress-test and stump baseline frontier models (Gemini 3.7 Flash, Claude Code), establishing ground-truth solutions and rigorous evaluation rubrics for autonomous agents.
- Design TPU-friendly tiling, sharding, reduction, data-movement, and mixed-precision strategies for common and custom ML operations.
- Validate numerical equivalence, gradient correctness, masking/indexing behavior, and edge cases between source GPU implementations and TPU-targeted implementations.
- Build and maintain benchmarks and diagnostic workflows that measure latency, throughput, utilization, memory usage, communication cost, and performance regressions.
- Analyze transformer and LLM workloads, including attention, MoE, KV-cache, quantized matmuls, and distributed training or serving patterns, to identify optimization opportunities.
- Collaborate with ML systems, compiler, infrastructure, and model engineers to translate workload requirements into efficient accelerator implementations.

📌 System Engineer (India)
🏢 Uber AI Solutions
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: system engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: system engineer (india) / india