04 Aug
|
Taghire
|
Bengaluru
The opportunity
Taghire is partnering with a Series A-funded GenAI inference infrastructure startup — backed by marquee investors and building one of the fastest platforms in the world to deploy, scale, and monitor GenAI models (LLMs, speech, vision, diffusion) across cloud and on-prem, engineered for strict SLAs, enterprise-grade security, and full observability.
This is not an "API-on-top-of-a-model" job. You'll work deep in the ML stack — GPUs, inference engines, and distributed systems — helping build core infrastructure that has real, measurable impact for enterprise customers around the world.
What you'll do
- Design and build the core architecture of a next-generation inference/MLOps platform running diverse GPU-accelerated workloads at scale .
- Standardize heterogeneous ML workloads — LLM / VLM / ASR / diffusion pipelines — and build clean orchestration abstractions for them.
- Build internal systems for continuous deployment of services, modules, and model pipelines across multi-cloud and hybrid environments.
- Create frameworks for high reliability, observability, and fault tolerance across mission-critical inference, training, and data pipelines.
- Partner with Applied ML and Core ML teams to push reliability, latency, and cost-efficiency .
- Build tooling to benchmark, evaluate, and deploy models quickly and consistently.
- Ship production-grade code and infrastructure with strong engineering fundamentals and test-driven development.
- Troubleshoot complex systems: performance bottlenecks, GPU behavior,
and distributed workloads.
What we're looking for
- Deep expertise in system design, distributed systems, and GPU-based ML workloads .
- Solid software-engineering fundamentals — data structures, APIs, testing, debugging.
- Experience with infrastructure-as-code (Terraform, Ansible) and cloud platforms ( AWS / GCP / Azure ).
- Solid ML fundamentals — model architectures (Transformers, CNNs) and inference behavior.
- Ability to build and reason about multi-step pipelines (ETL → model → evaluation → deploy).
- Strong systems knowledge: Linux internals, networking, performance tuning, GPU memory behavior .
- Comfortable owning large, ambiguous problems independently — and communicating design decisions and trade-offs clearly.
Bonus points
- Hands-on with modern inference stacks — TensorRT, Triton, vLLM / TGI, SGLang .
- Exposure to quantization, model optimization, or CUDA .
- Experience with Llama / Mistral, Whisper, or Stable Diffusion pipelines.
- Familiarity with CI/CD, Docker, GitHub workflows, and IaC-driven deployments.
- Experience designing high-availability or fault-tolerant production systems.
Why this role
- Work on deep ML and GenAI infrastructure problems — not a wrapper, a true infrastructure company.
- Solve real problems for enterprise clients where your work has direct, visible impact.
- Join a team of engineers from some of the most innovative companies in the world, and stay on the frontier of the AI transformation.
📌 Machine Learning Engineer — GPU / LLM Inference (Bengaluru)
🏢 Taghire
📍 Bengaluru