06 Aug
|
Weekday AI (YC W21)
|
Bengaluru
06 Aug
Weekday AI (YC W21)
Bengaluru
???? ???? ?? ??? ??? ?? ??? ???????'? ???????
?????? ?????: ?? ??????? - ?? ??????? (?? ??? ??-?? ???)
Experience: 5+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are seeking a highly skilled Inference Engineer with strong expertise in Large Language Models (LLMs) and vLLM to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.
As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the opportunity to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.
Requirements Key Responsibilities
- Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments
- Build and optimize high-performance model serving pipelines using vLLM and other modern inference frameworks
- Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters
- Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability
- Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference
- Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services
- Develop APIs, microservices, and deployment workflows for AI-powered applications
- Implement monitoring, logging, benchmarking, and performance profiling for inference workloads
- Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems
- Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices
What Makes You a Great Fit
- 5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering
- Strong hands-on expertise with Large Language Models (LLMs) and vLLM for production-scale inference
- Experience deploying and optimizing transformer-based models using modern inference frameworks
- Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques
- Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems
- Knowledge of containerization and orchestration technologies including Docker and Kubernetes
- Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment
- Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures
- Experience implementing monitoring, benchmarking, and observability for AI inference workloads
- Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems
- Robust collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams
- Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications
📌 Inference Engineer (Bengaluru)
🏢 Weekday AI (YC W21)
📍 Bengaluru