06 Aug
|
Knowledge Foundry
|
Chennai
06 Aug
Knowledge Foundry
Chennai
About the Role:
We are looking for an experienced AI Infrastructure & MLOps Engineer for our US InsurTech client to build and operate the production platform behind its Agentic AI ecosystem.
This is a hands-on engineering role focused on designing and managing secure, scalable, reliable, and observable infrastructure for production AI workloads. You will work across AWS cloud infrastructure, container orchestration, Infrastructure as Code, CI/CD, observability, security, and MLOps/LLMOps to enable reliable deployment and operation of enterprise AI systems.
If you have strong hands-on experience in AWS, Kubernetes/EKS, Docker, Terraform, CI/CD, and production AI/ML platform operations, we would like to hear from you.
Roles & Responsibilities:
- Build and manage AWS infrastructure for AI and machine learning workloads.
- Design, deploy, and operate containerized AI services using Docker and Kubernetes / Amazon EKS.
- Build and maintain Infrastructure as Code (IaC) using Terraform and/or AWS CloudFormation.
- Design and maintain CI/CD pipelines using tools such as GitHub Actions.
- Support the deployment and operation of AI/ML and LLM-based services in production environments.
- Implement and maintain MLOps and LLMOps practices for model, prompt, and AI application lifecycle management.
- Establish production-grade monitoring, logging, tracing, and observability for AI services and infrastructure.
- Implement platform monitoring using technologies such as Amazon CloudWatch, Prometheus, and Grafana.
- Ensure platform security through appropriate IAM, networking, secrets management, and least-privilege access controls.
- Build secure and scalable AWS environments using services such as Amazon EKS/ECS, EC2, Lambda, S3, VPC, IAM, and API Gateway.
- Support production reliability, scalability, availability, performance, and cost optimization of AI workloads.
- Collaborate with AI Platform Engineers and development teams to enable reliable production deployment of AI applications.
Required Skills & Experience:
- 5 - 8 years of relevant experience in cloud infrastructure, DevOps, platform engineering, MLOps, or related areas.
- Strong hands-on experience with AWS cloud infrastructure and services.
- Strong experience with Docker and containerized application deployment.
- Hands-on experience with Kubernetes, preferably Amazon EKS.
- Solid experience with Infrastructure as Code, particularly Terraform.
- Experience building and maintaining CI/CD pipelines, including GitHub Actions.
- Hands-on experience with AWS services relevant to production platforms, including:
- Amazon EKS / ECS
- Amazon EC2
- AWS Lambda
- Amazon S3
- AWS IAM
- Amazon VPC
- Amazon API Gateway
- Experience with monitoring and observability tools such as:
- Amazon CloudWatch
- Prometheus
- Grafana
- Strong understanding of:
- Cloud security
- IAM and least-privilege access
- Networking
- Secrets management
- Production reliability and scalability
- Experience with MLOps and/or LLMOps practices for production AI/ML systems.
- Experience supporting the deployment and operation of AI, ML, or LLM-based workloads in production.
AI / MLOps Platform Experience
The ideal candidate should have hands-on experience in areas such as:
- Deploying AI/ML services to production cloud environments.
- Containerizing AI applications using Docker.
- Deploying and operating workloads on Kubernetes / Amazon EKS.
- Automating infrastructure provisioning using Terraform or CloudFormation.
- Building automated CI/CD pipelines for application and infrastructure deployment.
- Implementing monitoring, logging, and observability for production AI services.
- Managing model, application, or AI platform lifecycle processes.
- Supporting scalable and reliable production environments for AI/LLM workloads.
- Troubleshooting infrastructure, deployment, performance, and reliability issues in production.
What We Are Looking For:
We are particularly interested in engineers who have actually built, deployed, and operated production infrastructure for AI/ML or LLM-based applications.
Interested candidates who can join within 30 days or less are encouraged to apply.
📌 AI Infrastructure & MLOps Engineer (AWS Platform) (Chennai)
🏢 Knowledge Foundry
📍 Chennai