Inference Engineer (Bengaluru)

Inference Engineer (Bengaluru)

04 Aug
|
Weekday AI
|
Bengaluru

04 Aug

Weekday AI

Bengaluru

???? ???? ?? ??? ??? ?? ??? ???????'? ???????

?????? ?????: ?? ??????? - ?? ??????? (?? ??? ??-?? ???)

Experience: 5+ yrs

Location: Bengaluru, Karnataka, India

Job Type: Full-time

We are seeking a highly skilled Inference Engineer with strong expertise in Large Language Models (LLMs) and vLLM to build, optimize, and scale high-performance AI inference platforms. This role is ideal for professionals who are passionate about deploying production-grade generative AI systems, optimizing model serving, and improving inference efficiency for large-scale applications.

As an Inference Engineer, you will be responsible for designing, implementing, and maintaining scalable LLM inference infrastructure that delivers low-latency, high-throughput AI experiences. You will work closely with AI Researchers, Machine Learning Engineers, Platform Engineers, and Product teams to deploy state-of-the-art language models, optimize serving pipelines, and improve resource utilization across GPU-based environments. This role offers the prospect to work with cutting-edge AI technologies while solving complex performance, scalability, and infrastructure challenges.

Key Responsibilities

Design, deploy, and maintain scalable inference infrastructure for Large Language Models (LLMs) in production environments.

Build and optimize high-performance model serving pipelines using vLLM and other modern inference frameworks.

Improve inference latency, throughput, memory utilization, and overall system performance across GPU clusters.

Deploy, monitor, and manage open-source and proprietary LLMs while ensuring reliability and scalability.

Optimize GPU resource allocation, batching strategies, caching mechanisms, and model parallelism for efficient inference.





Collaborate with Machine Learning Engineers to transition trained models into production-ready inference services.

Develop APIs, microservices, and deployment workflows for AI-powered applications.

Implement monitoring, logging, benchmarking, and performance profiling for inference workloads.

Troubleshoot production issues related to model serving, infrastructure, and distributed inference systems.

Contribute to the continuous improvement of AI platform architecture, automation, and deployment best practices.

What Makes You a Great Fit

5+ years of experience in Machine Learning Infrastructure, AI Platform Engineering, MLOps, or Inference Engineering.

Strong hands-on expertise with Large Language Models (LLMs) and vLLM for production-scale inference.

Experience deploying and optimizing transformer-based models using modern inference frameworks.

Solid understanding of GPU computing, CUDA, distributed inference, model quantization, and performance optimization techniques.

Experience with Python and deep learning frameworks such as PyTorch, Hugging Face Transformers, or similar ecosystems.

Knowledge of containerization and orchestration technologies including Docker and Kubernetes.

Familiarity with cloud platforms such as AWS, Azure, or GCP for AI infrastructure deployment.

Strong understanding of REST APIs, microservices, distributed systems, and scalable backend architectures.

Experience implementing monitoring, benchmarking, and observability for AI inference workloads.

Excellent analytical, debugging, and problem-solving skills with the ability to optimize complex AI systems.

Strong collaboration and communication skills, with the ability to work effectively across AI, platform, and product engineering teams.

Passion for advancing generative AI infrastructure and delivering reliable, high-performance inference solutions for real-world applications.

📌 Inference Engineer (Bengaluru)
🏢 Weekday AI
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: inference engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: inference engineer (bengaluru) / bengaluru