Infrastructure Engineer (India)

Infrastructure Engineer (India)

23 Sep
|
Emergys
|
India

23 Sep

Emergys

India

Experience 4-8 Years

Location: Pune

Role Overview

This role owns the inference path in depth. You will configure, deploy, tune and operate the model serving layer: how bundles and model profiles are composed, how replicas and service tiers are laid out across accelerator nodes, how throughput and latency behave under real concurrency, and how LLM application components built on the platform hold up in production. This is a hands-on engineering role rather than a design authority. You run the benchmark, read the profile and change the configuration yourself, and you supply sizing and performance inputs into the architecture function rather than owning solution architecture. On call for the serving and inference path is part of the role.

Key Responsibilities

- Own configuration of the model serving layer: model bundle composition, model profiles, deployment groups, replica counts and service-tier mapping.
- Deploy and manage model deployments through the platform’s Kubernetes custom resources and migrate configurations forward as the resource model changes between platform versions.
- Decide bundle strategy: which models are co-resident, which batch size and sequence-length configurations are worth the memory footprint, and where speculative decoding earns its cost.
- Manage checkpoint and artifact delivery, model versioning and rollback, and validate every bundle change before it reaches customer traffic.
- Enforce node-level memory and resource limits so that bundle configuration stays inside declared node capacity.
- Own inference performance targets: time to first token, inter-token latency, sustained tokens per second and latency percentiles under concurrency.
- Design and run benchmark campaigns with a defined methodology and turn the results into reusable sizing guidance rather than a one-off report.
- Profile and tune the serving path, including batching behavior, queue depth, request routing, cache utilization and concurrency limits.
- Diagnose whether a latency regression originates in the application, gateway, router, serving runtime, accelerator or network path, and prove it.
- Quantify the throughput and quality cost of quantization and configuration changes before recommending them.
- Build production LLM application components on the platform: retrieval-augmented generation, agentic workflows, tool and function calling, and structured extraction.
- Implement retrieval pipelines with a defined chunking strategy, embedding model selection, hybrid search and reranking, and measured recall.
- Build evaluation suites covering accuracy, regression across model versions,



and cost and latency alongside quality, and gate configuration changes.
- Implement guardrails and safety controls, including prompt-injection defense, input and output filtering and abuse detection.
- Build the pipelines that move model and bundle changes from validation into production with defined promotion, canary and rollback criteria, and own model-layer observability covering per-request token accounting, latency breakdown, queue metrics, worker state and error taxonomy.
- Run day-two model operations: bundle changes, configuration updates, and capacity rebalancing.
- Supply sizing inputs, routing semantics and quota behavior to the architecture, infrastructure and gateway teams, and support field and forward-deployed engineers on customer benchmarking, tuning and inference-path escalations.
- Work alongside agentic AI tooling for benchmark harnesses, evaluation suites and code review, and automate repetitive operational work.

Tools and Technologies

Category

Tools and technologies

Model serving and runtimes

vLLM, NVIDIA Triton and TensorRT-LLM, Text Generation Inference, Ray Serve, KServe, accelerator-specific serving runtimes, continuous batching, paged attention and KV cache management, speculative decoding, quantization including FP8, INT8, AWQ and GPTQ

Platform model deployment

Kubernetes custom resources for model bundles, profiles and deployments, replica groups and quality-of-service tiers, checkpoint and artifact management, node memory limit enforcement, Helm-managed model charts

LLM application frameworks

LangChain, LlamaIndex, LangGraph, DSPy, Semantic Kernel, Model Context Protocol, OpenAI-compatible client SDKs

Retrieval and vector search

pgvector, Milvus, Qdrant, Weaviate, OpenSearch k-NN, embedding model selection, cross-encoder rerankers, hybrid and BM25 search, chunking strategy

Evaluation and guardrails

Ragas, DeepEval, promptfoo, LLM-as-judge design, golden datasets and regression suites, NeMo Guardrails, Llama Guard, prompt-injection and jailbreak testing

Performance engineering

Time to first token, inter-token latency and tokens-per-second profiling, concurrency sweeps, k6, Locust, vegeta, py-spy and flame graphs, tokenizer analysis,



latency percentile reporting

MLOps and LLMOps

MLflow, Weights and Biases, Kubeflow, Argo Workflows, model registries, DVC, experiment tracking, canary and progressive rollout

Languages, libraries and data

Python with asyncio and FastAPI, PyTorch, Hugging Face Transformers, Go, Bash, SQL, PostgreSQL, Redis queues, object storage, Parquet, batch and streaming data paths

Containers and orchestration

Docker, Kubernetes, RKE2, Helm, operators, node affinity and accelerator scheduling

Observability

Prometheus and router-level inference metrics, Grafana, OpenTelemetry tracing, OpenSearch and Fluent Bit for per-request execution logs, structured usage telemetry

AI-assisted engineering

Agentic AI assistants used for benchmark tooling, evaluation harness generation and code review

Minimum Requirements

- Strong hands-on ML systems engineering, MLOps or LLMOps experience with production accountability, not experimentation only.
- Demonstrated experience deploying and operating LLM inference at scale, including capacity, latency and cost management.
- Deep understanding of LLM serving internals: batching, KV cache, context handling, tokenisation, sampling and decoding strategies.
- Practical performance-engineering ability: you can design a benchmark, run it, interpret percentile behavior and act on the result.
- Robust grasp of retrieval-augmented generation architecture and its failure modes, plus evaluation methodology that catches regressions.
- Understanding LLM application security: prompt injection, output handling, guardrails and abuse patterns.
- Strong Kubernetes, containers, Helm and CI/CD fundamentals, and the ability to work within a cluster you do not own.
- Experience with at least one major public cloud and with on-premises, hybrid or data-center deployment models.
- Strong Python engineering ability.
- Rigorous, measurement-led problem solving, willingness to stay hands-on at the configuration and code level, and availability for serving-path on-call.

Preferred Requirements

- Experience with non-GPU AI accelerators and their runtime, compilation, and scheduling models.
- Exposure to AI hardware, semiconductor platforms or simulation environments.
- Fine-tuning, LoRA and adapter workflows, and the judgement to know when not to use them.
- Experience with B2B enterprise products supporting both datacenter and cloud deployment.
- Contributions to open-source serving, evaluation or orchestration projects, or direct experience supporting customer benchmarking and performance escalations.

📌 Infrastructure Engineer (India)
🏢 Emergys
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: infrastructure engineer (india) / india