Senior Infrastructure Mlops Engineer (India)

Senior Infrastructure Mlops Engineer (India)

27 Sep
|
Prelect Networks
|
India

27 Sep

Prelect Networks

India

Senior Infrastructure MLOps Engineer

Remote Role

Responsibilities

● Manage the core infrastructure with a focus on supporting ML and LLM workloads. This infrastructure is relied on by a large subscriber base product as well as internal stakeholders across the company.

● Design, build, and maintain cloud infrastructure components (AWS and Kubernetes)

using a combination of in-house technology and open source software.

● Build and scale model serving and inference infrastructure for both real-time recommendation models and LLM workloads including model servers (e.G. NVIDIA

Triton), LLM inference engines (e.G. vLLM, TGI, TensorRT-LLM), autoscaling, and latency/throughput optimization.

● Manage GPU and accelerator infrastructure on Kubernetes scheduling, capacity and quota management, and cost optimization across on-demand and spot capacity.

● Design, build, and maintain monitoring and observability systems spanning both system health and model quality (drift, evaluation, and online performance) that enable engineers and the support team to gain insight and to discover and debug issues.

● Build and operate model evaluation infrastructure: offline and online eval pipelines,

golden/regression datasets, LLM-as-judge harnesses, A/B testing and experimentation,

and feedback/ground-truth loops

● Design, build, and maintain CI/CD systems enabling engineers to create pipelines to test and deploy their code.

● Establish tools, methods and best practices for other engineers interfacing with the ML infrastructure. Ensure reliability, security, and scalability of the platform. Promote

Infrastructure as Code.

● Work closely with other engineers to deploy and instrument software systems.

● Drive evaluation, selection,



and integration of third-party vendor systems and work closely with vendors to configure and manage them.

Qualifications

● 8+ years of software development and infrastructure management experience.

● 4+ years of experience developing and deploying ML systems at scale across training and inference

● Experience working with distributed systems and an understanding of microservices architecture principles.

● Experience with Linux and containerized (i.E. Docker) environments.

● Experience managing cloud computing environments (AWS or GCP) and configuring cloud services e.G. CloudWatch, Route 53, RDS, ElastiCache, SQS, ALB/NLB/ELB, VPC networking, IAM security.

● Experience with container orchestration platforms (i.E. Kubernetes), and familiarity with running GPU/accelerated workloads on them (e.G. NVIDIA device plugin, MIG,

timeslicing).

● Robust understanding of networking and internet application protocols including, but not limited to TCP/IP, DNS, and HTTP.

● Strong understanding of network and application security principles and best practices.

● Familiarity or hands-on experience with configuration management systems and

Infrastructure as Code (e.G. Terraform, CloudFormation).

● Familiarity or hands-on experience with Monitoring/Observability systems (e.G.

Prometheus, Grafana, TICK/InfluxDB, Fluentd, ELK, Datadog).

● Familiarity or hands-on experience with CI/CD automation systems e.G. Jenkins, Gitlab.





● Experience with relational and non-relational databases and familiarity with modern data warehousing and querying.

● Proficient in writing, testing, and profiling software in Golang, JavaScript/TypeScript, C++,

Ruby, Python or similar programming languages.

● Experience and aptitude for collaborating and communicating with internal and external stakeholders in both business and technical roles.

A strong candidate may also have one or more of these:

● Hands-on experience optimizing inference dynamic/continuous batching, quantization,

KV-cache management, or GPU memory tuning for low-latency or highthroughput serving.

● Experience with ML platform tooling: model registry and experiment tracking (e.G.

MLflow, Weights & Biases), feature stores (e.G. Feast, Tecton), and pipeline/workflow orchestration (e.G. Airflow, Dagster, Kubeflow, Ray, Metaflow).

● Experience building model evaluation, monitoring, or experimentation systems including drift detection, LLM-as-judge, or A/B testing for models.

● Experience with LLMOps: inference gateways/routers (e.G. LiteLLM), hosted model APIs

(Bedrock, OpenAI, Anthropic), RAG pipelines, vector databases (e.G. pgvector, Pinecone,

Weaviate), and prompt/version management.

● Experience building or operating agent harnesses or tool-using LLM applications in production.

● Experience with data and model versioning (e.G. DVC, LakeFS) and data/labeling pipelines.

● Experience with cost management / FinOps for GPU and inference workloads We are language agnostic, but most of our backend code is written in Golang, Ruby and

TypeScript, with some C++ and Python. Our services run on Kubernetes, and we practice continuous deployment across all of our systems.

📌 Senior Infrastructure Mlops Engineer (India)
🏢 Prelect Networks
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior infrastructure mlops engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: senior infrastructure mlops engineer (india) / india