LLM Ops Engineer (Infrastructure & Performance) (Bengaluru)

LLM Ops Engineer (Infrastructure & Performance) (Bengaluru)

02 Aug
|
Krazy Bee Services
|
Bengaluru

02 Aug

Krazy Bee Services

Bengaluru

LLM Ops Engineer (Infrastructure & Performance)

Role Overview

We are looking for a hands-on Junior AI MLOps Engineer to bridge the gap between model development and production-grade deployment. In this role, you will be responsible for the health of our AI infrastructure, focusing on the deployment, monitoring, and financial optimization of machine learning models. A key part of your responsibility will be evaluating infrastructure choices (e.g., Lambda vs.

EC2 vs. Bedrock) to provide leadership with meaningful cost-to-performance metrics that guide our scaling strategy.

Key Responsibilities

- Infrastructure Evaluation: Conduct comparative analysis of different AWS hosting environments to determine the most cost-effective deployment strategy for LLMs and specialized ML models.
- Cost & Performance Tracking: Develop dashboards and reporting frameworks to monitor "cost-per-inference" and "latency-vs-throughput" metrics across production batches.
- Pipeline Automation: Build and maintain CI/CD pipelines for automated model testing, containerization (Docker), and deployment to AWS environments (EKS, ECS, or Lambda).
- Resource Optimization: Hands-on management of GPU/CPU utilization to ensure high availability while minimizing cloud spend.
- Monitoring & Observability: Implement logging and alerting for model drift, API performance, and infrastructure health to ensure model reliability.
- Collaborative MLOps: Work closely with the Data Science team to move models from experimental notebooks to scalable, version-controlled microservices.

Required Skills & Qualifications

- Experience: 12 years of hands-on experience in MLOps, DevOps,



or Backend Engineering with a solid focus on ML infrastructure.
- Technical Stack: Proficiency in Python and experience with containerization tools like Docker and orchestration via Kubernetes.
- AWS Proficiency: Practical experience with AWS services such as Lambda, EC2, S3, and SageMaker. Knowledge of AWS Bedrock is a plus.
- Analytical Mindset: Ability to synthesize complex technical data into clear performance-to-cost reports for stakeholders.
- Foundational ML Knowledge: Understanding of the ML lifecycle, including model versioning, feature stores, and API integration (FastAPI/Flask).

Education: Bachelors degree in Computer Science, Information Technology, or a related technical field.

Disclaimer:

This is intended to outline the general nature and key responsibilities of the position. It is not intended to be an exhaustive list of all duties, responsibilities, or qualifications associated with the role. The responsibilities and qualifications described may be subject to change, and other duties may be assigned as needed.

Employment is at-will, meaning the employee or the employer may terminate the employment relationship at any time, with or without cause, and with or without notice.

Data Utilization Disclaimer:

By applying for this position, you acknowledge and agree that any personal data you provide may be used for recruitment and employment purposes. The data collected will be stored and processed in accordance with our privacy policy and applicable data protection laws. Your information will only be shared with relevant internal stakeholders and will not be disclosed to third parties without your consent, unless required by law.

📌 LLM Ops Engineer (Infrastructure & Performance) (Bengaluru)
🏢 Krazy Bee Services
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: llm ops engineer (infrastructure & performance) (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: llm ops engineer (infrastructure & performance) (bengaluru) / bengaluru