02 Aug
|
KnowledgeWorks Global
|
India
02 Aug
KnowledgeWorks Global
India
Sr. DevOps Engineer
Observability Strategy (Web &
- AI Applications) :
- Design and implement an end-to-end observability & alert stack covering both traditional web applications and AI/ML services.
- Develop and maintain tooling such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalent.
- Build AI-specific observability : model latency/throughput tracking, GPU utilization, token usage, drift detection, and inference quality signals.
Infrastructure Scaling &
- Reliability :
- Design and manage infrastructure capable of scaling across on-premise data centers and public cloud.
- Own capacity planning, load testing, auto-scaling, and cost-optimization initiatives across compute, storage, and networking.
- Implement Infrastructure as Code (Terraform, Ansible, or equivalent) to ensure environments are reproducible, version-controlled, and auditable.
- Lead disaster recovery, backup, and high-availability strategy for critical systems.
DevOps &
- CI/CD Delivery :
- Partner closely with Solution Architects to translate project and system designs into concrete DevOps execution plans.
- Design, build, and maintain CI/CD pipelines (Jenkins, GitHub Actions or ArgoCD/Flux for GitOps) across multiple projects and teams.
- Containerize and orchestrate applications using Docker and Kubernetes, including Helm chart and manifest management.
- Embed security and compliance checks (SAST/DAST, secrets scanning, image scanning) directly into the delivery pipeline (DevSecOps).
GPU Infrastructure &
- Model Deployment :
- Provision, configure, and manage GPU infrastructure (on-prem clusters and cloud GPU instances) for model training and inference.
- Deploy, scale, and monitor ML/LLM models in production using tools such as Triton Inference Server, vLLM, or similar.
- Optimize GPU utilization, cost,
and throughput across multi-tenant workloads; manage CUDA/driver/toolkit versions.
- Collaborate with data science/ML engineering teams on MLOps pipelines model versioning, experiment tracking, and model registry.
Requirements :
- 5-10 years of hands-on DevOps/SRE/Infrastructure engineering experience, including at least 23 years in a senior or lead capacity.
- Deep expertise in Git-based workflows and repository management at scale (GitHub/GitLab/Bitbucket).
- Proven experience designing observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, New Relic, OpenTelemetry).
- Solid background in cloud platforms (AWS, Azure, and/or GCP) and on-premise/hybrid infrastructure.
- Expert-level skills with Infrastructure as Code (Terraform, Ansible, CloudFormation, or Pulumi).
- Strong Kubernetes and Docker experience, including multi-cluster and multi-environment management.
- Hands-on experience building and maintaining CI/CD pipelines end to end.
- Working knowledge of GPU infrastructure (NVIDIA CUDA, drivers, NCCL) and experience deploying ML/AI models to production.
- Proficiency in scripting/automation languages : Python, Bash, and/or Go.
- Solid understanding of networking, load balancing, DNS, and security fundamentals in distributed systems.
- Experience partnering with architects and engineering leads to translate designs into infrastructure and delivery plans.
- Hands-on with SAST, DAST, and SCA tooling (e.g., SonarQube, Snyk, Checkmarx, OWASP DependencyCheck) integrated directly into CI/CD pipelines.
- Container and image security : vulnerability scanning (Trivy, Grype, Clair), minimal/hardened base images, and signed/verified image provenance (Cosign/Sigstore).
- Familiarity with Secrets management and credential hygiene using tools such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault.
- Excellent communication skills and comfort operating cross-functionally with development, data science, and product teams.
📌 KnowledgeWorks Global - Senior DevOps Engineer - Docker/Kubernetes (India)
🏢 KnowledgeWorks Global
📍 India