Sr. Backend Software Engineer (Infrastructure MLOps) (Lucknow)

Sr. Backend Software Engineer (Infrastructure MLOps) (Lucknow)

24 Sep
|
Sdlc Technologies
|
Lucknow

24 Sep

Sdlc Technologies

Lucknow

Senior Backend Software Engineer (Infrastructure MLOps)

The infrastructure team builds the foundation for a 24/7 live and on-demand streaming product serving video to our growing subscriber base. It provides a platform that enables groups across the company such as video delivery, product, support, and analytics to do their best work easily and smoothly. Our systems are highly distributed, horizontally scalable and run in the cloud on Kubernetes.

We continuously test and deploy our code and closely monitor the health and performance of our production workplace. We are seeking an enthusiastic Infrastructure MLOps Software Engineer to help us build and improve the ML infrastructure. We are building out ML and LLM infrastructure and this role will be pivotal in setting direction, adopting best practices, and shipping infrastructure that will be used across the company.

As a sample, the ML infrastructure supports teams like Discovery in building content recommendation models that will help our users find more favorite shows from the large catalog. The engineering team is designing and developing our own agent harnesses to support coding and empowering everyone at the company to get answers across our code bases, analytics, logs, and communication tools. The video player team is using LLMs to classify ads and label our video catalog so that users have a better experience.

Youll build the serving, evaluation and platform infrastructure that these workloads depend on.

Responsibilities:

Manage the core infrastructure with a focus on supporting ML and LLM workloads. This infrastructure is relied on by a large subscriber base product as well as internal stakeholders across the company.

Design, build, and maintain cloud infrastructure components (AWS and Kubernetes) using a combination of in-house technology and open source software.

Build and scale model serving and inference infrastructure for both real-time recommendation models and LLM workloads including model servers (e.g. NVIDIA Triton), LLM inference engines (e.g. vLLM, TGI, TensorRT-LLM), autoscaling, and latency/throughput optimization.

Manage GPU and accelerator infrastructure on Kubernetes scheduling, capacity and quota management,



and cost optimization across on-demand and spot capacity.

Design, build, and maintain monitoring and observability systems spanning both system health and model quality (drift, evaluation, and online performance) that enable engineers and the support team to gain insight and to discover and debug issues.

Build and operate model evaluation infrastructure: offline and online eval pipelines, golden/regression datasets, LLM-as-judge harnesses, A/B testing and experimentation, and feedback/ground-truth loops

Design, build, and maintain CI/CD systems enabling engineers to create pipelines to test and deploy their code.

Establish tools, methods and best practices for other engineers interfacing with the ML infrastructure. Ensure reliability, security, and scalability of the platform. Promote Infrastructure as Code.

Work closely with other engineers to deploy and instrument software systems.

Drive evaluation, selection, and integration of third-party vendor systems and work closely with vendors to configure and manage them.

Qualifications:

8+ years of software development and infrastructure management experience.

4+ years of experience developing and deploying ML systems at scale across training and inference

Experience working with distributed systems and an understanding of microservices architecture principles.

Experience with Linux and containerized (i.e. Docker) environments.

Experience managing cloud computing environments (AWS or GCP) and configuring cloud services e.g. CloudWatch, Route 53, RDS, ElastiCache, SQS, ALB/NLB/ELB, VPC networking, IAM security.

Experience with container orchestration platforms (i.e. Kubernetes), and familiarity with running GPU/accelerated workloads on them (e.g. NVIDIA device plugin, MIG, timeslicing).

Strong understanding of networking and internet application protocols including, but not limited to TCP/IP, DNS,



and HTTP.

Strong understanding of network and application security principles and best practices.

Familiarity or hands-on experience with configuration management systems and Infrastructure as Code (e.g. Terraform, CloudFormation).

Familiarity or hands-on experience with Monitoring/Observability systems (e.g. Prometheus, Grafana, TICK/InfluxDB, Fluentd, ELK, Datadog).

Familiarity or hands-on experience with CI/CD automation systems e.g. Jenkins, Gitlab.

Experience with relational and non-relational databases and familiarity with modern data warehousing and querying.

Proficient in writing, testing, and profiling software in Golang, JavaScript/TypeScript, C++, Ruby, Python or similar programming languages.

Experience and aptitude for collaborating and communicating with internal and external stakeholders in both business and technical roles.

A strong candidate may also have one or more of these:

Hands-on experience optimizing inference dynamic/continuous batching, quantization, KV-cache management, or GPU memory tuning for low-latency or highthroughput serving.

Experience with ML platform tooling: model registry and experiment tracking (e.g. MLflow, Weights & Biases), feature stores (e.g. Feast, Tecton), and pipeline/workflow orchestration (e.g. Airflow, Dagster, Kubeflow, Ray, Metaflow).

Experience building model evaluation, monitoring, or experimentation systems including drift detection, LLM-as-judge, or A/B testing for models.

Experience with LLMOps: inference gateways/routers (e.g. LiteLLM), hosted model APIs (Bedrock, OpenAI, Anthropic), RAG pipelines, vector databases (e.g. pgvector, Pinecone, Weaviate), and prompt/version management.

Experience building or operating agent harnesses or tool-using LLM applications in production.

Experience with data and model versioning (e.g. DVC, LakeFS) and data/labeling pipelines.

Experience with cost management / FinOps for GPU and inference workloads We are language agnostic, but most of our backend code is written in Golang, Ruby and TypeScript, with some C++ and Python. Our services run on Kubernetes, and we practice continuous deployment across all of our systems.

📌 Sr. Backend Software Engineer (Infrastructure MLOps) (Lucknow)
🏢 Sdlc Technologies
📍 Lucknow

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr. backend software engineer (infrastructure mlops) (lucknow) / lucknow