19 Aug
|
Libra AI
|
Bengaluru
19 Aug
Libra AI
Bengaluru
: DevOps Engineer, AI Infrastructure
Location: Bengaluru, India
Experience: 3+ years
Employment Type: Full-time
Department: Software Engineering
About the Role
We are looking for a DevOps Engineer with strong exposure to AI infrastructure to build, operate, and scale the systems that power our products.
This role goes beyond traditional CI/CD and cloud operations. You will work on Amazon EKS, GPU infrastructure, AI workloads, scalable databases, vector databases, distributed systems, observability, and production reliability . You will partner closely with application, data, and AI engineering teams to ensure that our infrastructure remains secure, performant, cost-efficient, and highly available as usage grows.
Key Responsibilities
- Design, deploy, and manage production infrastructure on AWS.
- Operate and optimize Amazon EKS clusters for microservices, data workloads, and AI/ML applications.
- Manage Kubernetes workloads, including deployments, services, ingress, autoscaling, storage, networking, and cluster upgrades.
- Build infrastructure and deployment automation using Infrastructure as Code.
- Develop and maintain reliable CI/CD pipelines for application and infrastructure releases.
- Manage GPU-enabled infrastructure for AI inference, model serving, training, and batch workloads.
- Optimize GPU utilization, scheduling, provisioning, capacity planning, and cost.
- Configure and maintain GPU-enabled Kubernetes nodes, including drivers, device plugins, node pools, taints, tolerations, and workload isolation.
- Support scaling of relational and non-relational databases for increasing traffic and data volumes.
- Improve database availability, performance, replication, read scaling, backups, failover, and disaster recovery.
- Design and operate scalable vector database infrastructure for embeddings, semantic search, and retrieval-augmented generation workloads.
- Work with distributed vector database architectures, sharding, replication, partitioning, indexing, and horizontal scaling.
- Evaluate and implement solutions for vector databases such as Milvus, Qdrant, Weaviate, OpenSearch, pgvector, Pinecone, or equivalent systems .
- Build observability across infrastructure and applications using metrics, logs, traces, dashboards, and actionable alerts.
- Establish and improve reliability practices, including SLOs, incident response, root-cause analysis,
and operational runbooks.
- Identify infrastructure bottlenecks and improve system performance, availability, and cost efficiency.
- Implement security best practices for IAM, secrets management, network access, container images, and Kubernetes workloads.
- Collaborate with AI and backend engineers to productionize model-serving and data-processing pipelines.
- Automate routine operational tasks and reduce manual intervention across environments.
- Document infrastructure architecture, deployment procedures, troubleshooting guides, and recovery processes.
Required Qualifications
- 3+ years of experience in DevOps, Cloud Infrastructure, Site Reliability Engineering, Platform Engineering, or a related role.
- Solid hands-on experience with AWS services and production cloud environments.
- Practical experience managing Amazon EKS and Kubernetes in production or staging environments.
- Experience with Docker, Kubernetes networking, Helm, autoscaling, persistent volumes, and cluster troubleshooting.
- Experience with CI/CD systems such as GitHub Actions, GitLab CI, Jenkins, Argo CD, or equivalent.
- Proficiency with Infrastructure as Code tools such as Terraform, Pulumi, or CloudFormation.
- Strong Linux administration, networking, security, and troubleshooting skills.
- Experience managing GPU-based workloads or AI infrastructure.
- Understanding of GPU provisioning, utilization monitoring, scheduling, and capacity management.
- Experience scaling databases through replication, partitioning, indexing, caching, connection pooling, or sharding.
- Familiarity with distributed systems concepts such as consistency, replication, partition tolerance, fault tolerance, and horizontal scaling.
- Experience working with at least one vector database or vector search platform.
- Proficiency in scripting or programming using Python, Bash, Go, or a similar language.
- Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, CloudWatch, Datadog, or ELK.
- Strong understanding of production operations, incident management, and root-cause analysis.
Preferred Qualifications
- Experience running AI inference or model-serving workloads on Kubernetes.
- Familiarity with NVIDIA GPUs, CUDA, NVIDIA device plugins, GPU Operator, MIG, or similar technologies.
- Experience with Karpenter, Cluster Autoscaler, or custom Kubernetes autoscaling solutions.
- Experience operating Ray, KServe, Seldon, Triton Inference Server, vLLM, or similar AI-serving systems.
- Experience with distributed vector databases such as Milvus, Qdrant, Weaviate, OpenSearch, or large-scale pgvector deployments.
- Experience with PostgreSQL, Redis, Kafka, Elasticsearch/OpenSearch, or other distributed data systems.
- Knowledge of AWS services such as EC2, EKS, RDS, Aurora, ElastiCache, S3, CloudFront, IAM, VPC, and CloudWatch.
- Experience with GitOps and progressive delivery using Argo CD, Flux, Argo Rollouts, or equivalent.
- Understanding of FinOps and cloud cost optimization, particularly GPU cost management.
- Experience implementing security, compliance, backup, and disaster-recovery controls.
- AWS, Kubernetes, or relevant cloud certifications.
What Success Looks Like In the first six months, you will be expected to:
- Improve the reliability and operational maturity of our EKS environments.
- Establish efficient deployment and rollback workflows.
- Improve GPU utilization, scheduling, monitoring, and cost visibility.
- Support scalable database and vector database architectures.
- Strengthen observability, alerting, incident response, and production documentation.
- Help AI and engineering teams deploy workloads reliably and efficiently.
- Reduce infrastructure bottlenecks and manual operational work.
Ideal Candidate You are a hands-on infrastructure engineer who enjoys solving complex production problems. You understand that AI infrastructure requires more than deploying containers: it requires careful management of GPU capacity, workload scheduling, distributed data systems, database performance, reliability, and cost .
You are comfortable debugging issues across the application, Kubernetes, cloud, networking, and data layers. You take ownership, automate wherever possible, and can balance speed of execution with operational excellence.
📌 DevOps Engineer (Bengaluru)
🏢 Libra AI
📍 Bengaluru