Machine Learning Engineer — GPU / LLM Inference (Bengaluru)

Machine Learning Engineer — GPU / LLM Inference (Bengaluru)

04 Aug
|
Taghire
|
Bengaluru

04 Aug

Taghire

Bengaluru

The opportunity

Taghire is partnering with a Series A-funded GenAI inference infrastructure startup — backed by marquee investors and building one of the fastest platforms in the world to deploy, scale, and monitor GenAI models (LLMs, speech, vision, diffusion) across cloud and on-prem, engineered for strict SLAs, enterprise-grade security, and full observability.

This is not an "API-on-top-of-a-model" job. You'll work deep in the ML stack — GPUs, inference engines, and distributed systems — helping build core infrastructure that has real, measurable impact for enterprise customers around the world.

What you'll do

- Design and build the core architecture of a next-generation inference/MLOps platform running diverse GPU-accelerated workloads at scale .
- Standardize heterogeneous ML workloads — LLM / VLM / ASR / diffusion pipelines — and build clean orchestration abstractions for them.
- Build internal systems for continuous deployment of services, modules, and model pipelines across multi-cloud and hybrid environments.
- Create frameworks for high reliability, observability, and fault tolerance across mission-critical inference, training, and data pipelines.
- Partner with Applied ML and Core ML teams to push reliability, latency, and cost-efficiency .
- Build tooling to benchmark, evaluate, and deploy models quickly and consistently.
- Ship production-grade code and infrastructure with strong engineering fundamentals and test-driven development.
- Troubleshoot complex systems: performance bottlenecks, GPU behavior,



and distributed workloads.

What we're looking for

- Deep expertise in system design, distributed systems, and GPU-based ML workloads .
- Solid software-engineering fundamentals — data structures, APIs, testing, debugging.
- Experience with infrastructure-as-code (Terraform, Ansible) and cloud platforms ( AWS / GCP / Azure ).
- Solid ML fundamentals — model architectures (Transformers, CNNs) and inference behavior.
- Ability to build and reason about multi-step pipelines (ETL → model → evaluation → deploy).
- Strong systems knowledge: Linux internals, networking, performance tuning, GPU memory behavior .
- Comfortable owning large, ambiguous problems independently — and communicating design decisions and trade-offs clearly.

Bonus points

- Hands-on with modern inference stacks — TensorRT, Triton, vLLM / TGI, SGLang .
- Exposure to quantization, model optimization, or CUDA .
- Experience with Llama / Mistral, Whisper, or Stable Diffusion pipelines.
- Familiarity with CI/CD, Docker, GitHub workflows, and IaC-driven deployments.
- Experience designing high-availability or fault-tolerant production systems.

Why this role

- Work on deep ML and GenAI infrastructure problems — not a wrapper, a true infrastructure company.
- Solve real problems for enterprise clients where your work has direct, visible impact.
- Join a team of engineers from some of the most innovative companies in the world, and stay on the frontier of the AI transformation.

📌 Machine Learning Engineer — GPU / LLM Inference (Bengaluru)
🏢 Taghire
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: machine learning engineer — gpu / llm inference (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: machine learning engineer — gpu / llm inference (bengaluru) / bengaluru