Key Responsibilities
- Build and optimize ML model serving infrastructure with a focus on low-latency, cost-efficient inference.
- Architect scalable inference pipelines balancing latency, throughput, and infrastructure cost.
- Develop monitoring, logging, and observability solutions for production ML systems.
- Collaborate with ML Engineers to establish best practices for model deployment and inference optimization.
- Design and implement enterprise-scale, cost-optimized ML infrastructure.
- Work closely with MLEs, QA Engineers, DevOps Engineers, and cross-functional teams.
- Evaluate, benchmark, and adopt new ML infrastructure technologies and tools.
- Contribute to architecture and design decisions for distributed ML systems.
Mandatory Skills & Experience
- 5+ years of software engineering experience with
Python
.
- Robust hands-on experience with
PyTorch
.
- Experience optimizing ML models using
AWS Neuron, ONNX, and TensorRT
.
- Hands-on experience with
AWS SageMaker
,
Inferentia
, and
Trainium
.
- Experience building and operating
AWS serverless architectures
.
- Strong understanding of
event-driven architectures
using
SQS
,
SNS
, and serverless caching.
- Experience with
Docker
and container orchestration.
- Strong knowledge of
RESTful API
design and development.
- Experience writing secure, high-quality, production-grade code and using static code analysis tools.
- Strong understanding of algorithms, data structures, problem-solving, and complexity analysis.
- Excellent verbal and written communication skills.
📌 MLOps Engineer (India)
🏢 Talentoj
📍 India
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.