09 Oct
|
Cornerstone OnDemand
|
India
09 Oct
Cornerstone OnDemand
India
We're looking for a
Senior Dev Ops Engineer
This role is Office Based, Hyderabad Office
Senior Dev Ops / Cloud Platform Engineer – ML & AI Infrastructure
Job Summary
We are looking for a Senior Dev Ops / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services.
The ideal candidate will have hands-on experience with AWS EKS, Sage Maker, Bedrock, Docker, Kubernetes, Terraform, Helm, Git Hub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments. This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost-efficient platforms for ML services across development and production environments.
In this role you will...
Key Responsibilities
AWS & Kubernetes Infrastructure
- Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.
- Manage Amazon EKS clusters, including cluster provisioning, upgrades, scaling, networking, and troubleshooting.
- Work with AWS Sage Maker, AWS Bedrock, EKS, ECS, and related AWS services.
- Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB.
- Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security.
- Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.
- Implement best practices for security, reliability, availability, and scalability.
CI/CD & Azure-to-AWS Migration
- Build and maintain CI/CD pipelines for ML and AI services.
- Develop and manage Git Hub Actions and Azure Dev Ops pipelines using YAML.
- Migrate repositories and CI/CD workflows from Azure Dev Ops to Git Hub/AWS.
- Automate build, test, containerization, deployment, and release processes.
- Establish deployment strategies across development, staging, and production environments.
ML Service Deployment
- Deploy and manage ML services across AWS EKS/ECS and Sage Maker.
- Build and maintain Docker containers and Kubernetes deployments.
- Manage environment segregation and configuration across Dev, QA, and Production.
- Develop and maintain Kubernetes manifests and Helm charts.
- Troubleshoot ML service deployment, networking, scaling, and runtime issues.
App Runner to EKS Migration
- Lead migration of existing services from AWS App Runner to Amazon EKS.
- Containerize applications and develop Kubernetes manifests/Helm charts.
- Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.
- Ensure minimal service disruption during migration and establish operational best practices on EKS.
Self-Hosted LLM & AI Infrastructure
- Deploy and manage self-hosted Large Language Models and inference services.
- Work with model serving frameworks such as vLLM.
- Design containerized infrastructure for GPU-based model serving and inference.
- Manage model versions, deployments, configurations, and rollback strategies.
- Support migration of ML services from managed APIs/services to self-hosted models.
- Work with engineering teams on API integration and inference infrastructure.
Databricks Administration
- Administer Databricks workspaces, clusters, permissions, and access controls.
- Manage cluster configuration, policies, and resource utilization.
- Support LMI Insights and related ML/AI workloads.
- Troubleshoot Databricks infrastructure and connectivity issues.
- Implement appropriate security and access-control practices.
Elasticsearch Infrastructure
- Design, deploy, and manage Elasticsearch clusters.
- Perform cluster sizing, scaling, configuration, and performance optimization.
- Manage indices, mappings, retention, and data lifecycle requirements.
- Support Kibana configuration, dashboards, and troubleshooting.
- Monitor Elasticsearch health, capacity, and performance.
Monitoring, Reliability & Auto-Scaling
- Implement monitoring and observability for Kubernetes, AWS, and ML services.
- Use Prometheus, Grafana, and AWS Cloud Watch for monitoring and alerting.
- Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.
- Establish proactive alerting for infrastructure and application health.
- Perform capacity planning and resource optimization.
- Identify opportunities for AWS infrastructure and compute cost optimization.
Infrastructure as Code & Automation
- Build and maintain infrastructure using Terraform.
- Develop reusable Terraform modules for AWS and Kubernetes infrastructure.
- Manage Kubernetes deployments using Helm charts.
- Automate infrastructure provisioning, configuration, deployments, and operational tasks.
- Maintain infrastructure documentation and deployment standards.
Cross-Team Collaboration
- Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams.
- Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.
- Coordinate infrastructure requirements for new ML models and services.
- Participate in production incident resolution, root-cause analysis, and continuous improvement.
- Establish engineering standards around deployment, monitoring, security, and operational ownership.
You've got what it takes if you have...
Required Skills & Experience
- 5+ years of experience in Dev Ops, Cloud Infrastructure, SRE, or Platform Engineering.
- Strong hands-on experience with AWS.
- Solid experience administering Amazon EKS and Kubernetes in production.
- Hands-on experience with:AWS EKSAWS Sage MakerAWS BedrockAWS ECSAWS App Runner Kubernetes DockerNGINX / AWS ALB Ingress Cloudflare / Cloudflare Tunnels
- Strong experience with Terraform and Helm.
- Robust experience developing CI/CD pipelines using Git Hub Actions and/or Azure Dev Ops.
- Strong YAML scripting and Git experience.
- Experience migrating CI/CD pipelines and repositories from Azure to AWS/Git Hub.
- Experience deploying and operating ML/AI services.
- Experience with self-hosted LLM/model serving, preferably vLLM.
- Experience with GPU-based workloads is highly desirable.
- Experience with Databricks administration.
- Experience managing Elasticsearch and Kibana.
- Experience with Prometheus, Grafana, and Cloud Watch.
- Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery.
- Strong understanding of cloud networking fundamentals.
- Experience with production troubleshooting, monitoring, capacity planning, and cost optimization.
- Strong understanding of security, IAM, secrets management, and access control.
Preferred / Nice-to-Have Skills
- Experience supporting Generative AI / LLM platforms.
- Experience with GPU infrastructure and NVIDIA/CUDA environments.
- Experience with model lifecycle and model version management.
- Experience migrating workloads between managed cloud services and Kubernetes.
- Experience with AWS networking such as VPC, load balancers, security groups, and Route 53.
- Experience with API gateways and microservice architectures.
- Experience with Python or shell scripting for infrastructure automation.
- Experience working with Data Engineering and ML teams in a production environment.
What You'll Own
- AWS ML/AI infrastructure
- EKS cluster administration and upgrades
- ML service deployment and production operations
- CI/CD automation
- App Runner → EKS migration
- Self-hosted LLM infrastructure and vLLM
- Databricks platform administration
- Elasticsearch infrastructure
- Monitoring and auto-scaling
- Terraform and Helm-based infrastructure automation
- Cloud cost, reliability, and performance optimization
Ideal Candidate
The ideal candidate is a hands-on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD → AWS infrastructure → EKS deployment → monitoring → scaling → production support.
They should be comfortable working across both traditional Dev Ops infrastructure and modern AI/ML infrastructure, and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform.
#LI-Onsite
📌 Senior DevOps Engineer (India)
🏢 Cornerstone OnDemand
📍 India