06 Aug
|
Deccan AI Experts
|
India
06 Aug
Deccan AI Experts
India
About Us
Deccan AI Experts is a pioneering AI company founded by IIT Bombay and IIM Ahmedabad alumni, with a strong founding team from IITs, NITs, and BITS. We specialize in high-quality human-curated data, AI-first operations, and advanced AI evaluation systems. Our global network of technology experts helps train and evaluate next-generation AI models through expert engineering judgment and domain-specific expertise.
About the Role
We are seeking a Kernel Optimization Engineer (Freelancer) to support advanced AI evaluation initiatives focused on GPU kernel optimization, high-performance computing (HPC), parallel programming, compiler optimizations, AI acceleration, and AI-generated technical content evaluation.
In this role, you will evaluate AI-generated GPU kernels, CUDA implementations, parallel algorithms, performance optimization strategies, compiler recommendations, low-level systems code, and technical documentation. Your expertise will help improve AI systems designed for high-performance machine learning, scientific computing, GPU programming, and systems optimization.
This position is ideal for professionals with experience in GPU programming, kernel optimization, compiler engineering, HPC, embedded systems, AI acceleration, or performance engineering.
Responsibilities
- Review AI-generated CUDA kernels, OpenCL code, GPU optimization strategies, compiler optimizations, benchmarking reports, and performance analyses.
- Evaluate optimization techniques involving thread scheduling, memory coalescing, shared memory utilization, vectorization, SIMD/SIMT execution, cache optimization, kernel fusion, and asynchronous execution.
- Verify AI-generated recommendations for latency reduction, throughput optimization, memory bandwidth utilization, profiling, and hardware acceleration.
- Assess AI-generated code for correctness, maintainability, portability, and production readiness.
- Identify synchronization issues, race conditions, memory bottlenecks,
inefficient execution patterns, compiler optimization opportunities, and scalability limitations.
- Provide structured feedback to improve AI performance in kernel optimization, GPU programming, and high-performance systems engineering.
- Review peer-developed deliverables to maintain quality and consistency standards.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Software Engineering, Applied Mathematics, or a related field is required.
- Master's degree or Ph.D. in Computer Science, High-Performance Computing, Parallel Computing, Computer Architecture, or a related discipline is preferred.
- 2+ years of hands-on experience in GPU programming, kernel optimization, performance engineering, HPC, compiler engineering, or systems programming.:
- Developing and optimizing GPU kernels for production workloads.
- Profiling applications and identifying performance bottlenecks.
- Optimizing memory access patterns and computational efficiency.
- Reviewing AI-generated low-level systems code and optimization recommendations.
- Proficiency in C, C++, CUDA , and familiarity with Python for benchmarking and automation.
- Experience with technologies such as NVIDIA CUDA Toolkit, Nsight Systems, Nsight Compute, OpenMP, MPI, OpenCL, SYCL, ROCm, Triton, LLVM, GCC, Clang , or equivalent performance engineering tools.
- Familiarity with AI frameworks such as PyTorch, TensorFlow, JAX , and optimized inference libraries including TensorRT, cuDNN, NCCL , or equivalent is preferred.
- Excellent analytical thinking, systems programming,
performance tuning, and written English communication skills.
- Strong attention to detail and ability to evaluate production-grade optimized code.
- Ability to work independently in a remote, fast-paced environment.
Preferred Qualifications
- Experience working with semiconductor companies, AI infrastructure providers, cloud computing platforms, research laboratories, supercomputing centers, GPU vendors, or high-performance computing organizations.
- Expertise in one or more areas such as LLM inference optimization, distributed GPU training, compiler backends, AI accelerator development, scientific computing, numerical optimization, or embedded GPU systems.
- Experience evaluating AI-generated CUDA code, performance reports, compiler optimization strategies, benchmarking methodologies, or technical documentation.
- Familiarity with Generative AI, prompt engineering, RLHF (Reinforcement Learning from Human Feedback), AI benchmarking, AI compiler optimization, or AI systems evaluation is highly desirable.
- Contributions to open-source compiler projects, CUDA libraries, HPC frameworks, technical blogs, research publications, or conference presentations are a plus.
- Professional certifications or advanced training in CUDA programming, HPC, GPU optimization, or cloud computing are advantageous.
Why Join Us
- Competitive hourly pay: ₹2,000/hour
- Fully remote with adaptable working hours.
- Opportunity to contribute to cutting-edge AI initiatives in GPU optimization, high-performance computing, and Generative AI.
- Exposure to advanced AI systems focused on compiler optimization, kernel performance, distributed computing, and AI infrastructure.
- Flexible project-based opportunities with global teams.
- Work on next-generation AI solutions supporting machine learning platforms, semiconductor companies, cloud providers, and high-performance computing environments.
📌 Kernel Optimization Engineer (Freelancer) (India)
🏢 Deccan AI Experts
📍 India