18 Aug
|
Nasugroup.com
|
Bengaluru
18 Aug
Nasugroup.com
Bengaluru
This role is responsible for designing, deploying, and operating scalable AI/ML and Generative AI platforms on AWS. The engineer will enable robust MLOps, LLMOps, and DevOps practices,
supporting production-grade AI solutions (including LLMs and agent-based systems) across lending,
risk, and customer journeys.
Key Responsibilities
- AI/ML & GenAI Operations (AI Ops)
Operate and support production ML and GenAI workloads on AWS
Monitor:
o Model performance, drift, and reliability o LLM outputs (hallucination, latency, response quality)
Ensure high availability of AI services powering lending workflows
- MLOps & LLMOps (AWS Native)
Build and manage end-to-end ML pipelines using:
o Amazon SageMaker (training, deployment, pipelines)
o SageMaker Model Registry for versioning
Implement:
o CI/CD pipelines for ML and GenAI workloads o Experiment tracking and reproducibility
Enable LLMOps practices:
o Prompt lifecycle management o Evaluation frameworks for GenAI outputs
- DevOps & Cloud Engineering (AWS)
Design and manage infrastructure using:
o Terraform / AWS CloudFormation
Implement CI/CD pipelines using:
o AWS CodePipeline, CodeBuild, GitHub Actions
Orchestrate workloads with:
o EKS (Kubernetes) and Docker
Ensure scalability, resilience, and cost optimisation
- GenAI & Agentic AI Enablement
Deploy and manage:
o Amazon Bedrock (LLMs, foundation models)
o RAG pipelines using vector databases (OpenSearch, Pinecone, etc.)
Enable runtime support for:
o AI agents and multi-agent workflows
Integrate AI systems with:
o APIs, event-driven services, and enterprise platforms
- Data & Pipeline Integration
Build pipelines using:
o AWS Glue, Lambda, Step Functions
Manage data storage and access via:
o S3, Redshift, DynamoDB
Enable real-time and batch AI workflows
- Monitoring, Observability & Reliability
Implement monitoring using:
o CloudWatch, Prometheus, Grafana
Track:
o Model metrics, pipeline performance, system health
Define SLAs/SLOs and manage incident response
- Security, Risk & Compliance
Ensure secure AI deployments using:
o IAM, KMS, Secrets Manager
Implement data governance and privacy controls
Enforce Responsible AI and model governance standards
- Collaboration & Enablement
Work with:
o AI Architects, Data Scientists, Platform Engineers
Enable teams with:
o Reusable MLOps templates and frameworks o Self-service AI deployment capabilities
Key Skills and Experience
AWS AI/ML & Cloud Stack
Robust experience with:
o Amazon SageMaker (end-to-end ML lifecycle)
o Amazon Bedrock (GenAI / LLMs)
Familiarity with:
o OpenSearch, S3, Lambda, API Gateway
MLOps / LLMOps
Experience implementing:
o ML pipelines, model registries, CI/CD for ML
Knowledge of:
o Prompt engineering workflows o GenAI evaluation techniques
DevOps & Platform Engineering
Hands-on experience with:
o Docker, Kubernetes (EKS)
o Terraform / CloudFormation
CI/CD:
o CodePipeline, Jenkins, GitHub Actions
Programming
Python (primary), Bash scripting
Experience building APIs (FastAPI preferred)
Monitoring & Reliability
Experience with:
o CloudWatch, ELK stack, Prometheus
Understanding of:
o AI system observability and logging
- .
📌 AI/ML [Sagemaker with AWS] - 5+yrs (Bengaluru)
🏢 Nasugroup.com
📍 Bengaluru