16 Sep
|
Cloudstrats
|
Delhi
Company Description Cloudstrats is an Indian exponential technology product company that builds Artificial Intelligence platforms powered by analytics and automation. Its solutions use computer vision, speech, and text technologies to address complex real-world challenges for government and enterprise customers. The company’s platforms apply Deep Learning, NLP, and ML to support sustainable development goals across sectors such as citizen services, public safety, education, health, and transport.
Cloudstrats delivers marquee projects including command control centers, e-governance solutions, grievance redressal platforms, data hubs, fraud detection, and health systems. The organization is driven by the vision of Atmanirbhar Bharat and large-scale digital transformation in India and beyond.
Immediate Joiners preferred.
About the Role.
We're looking for a hands-on private-cloud engineer who can take an OpenStack-based platform from bare metal to a working, accepted, and supportable production environment — and then own it through its operational lifecycle.
This is a greenfield deployment covering compute, software-defined storage, GPU enablement, Kubernetes, and an AI/ML platform. You'll lead or contribute to technical delivery, automation, capacity growth, operational improvement, resilience testing, documentation, and training of the internal cloud team.
This is not a role for someone who has only consumed OpenStack through a dashboard. We need someone who has deployed it, broken it, recovered it, automated it, and documented it — using Kolla Ansible as the deployment tooling.
The role combines architecture input with hands-on engineering. You won't be expected to single-handedly deliver every specialist subsystem without support, but you must understand the full platform and be able to troubleshoot across its major layers.
What You will do.
Deployment & Architecture
- Design and deploy a multi-node OpenStack private cloud using Kolla Ansible
- Configure a highly available control plane and define service placement, failure domains, and network segmentation (management, API, provider, tenant, external)
- Deploy and configure core services: Keystone, Nova, Placement, Glance, Neutron, Cinder, Horizon
- Maintain version-controlled inventory, configuration, and secrets-management procedures
- Produce architecture documents, capacity plans, and implementation runbooks
Storage
- Design and deploy a distributed software-defined storage platform (Ceph preferred) as the shared substrate
- Integrate block/image storage with Glance, Cinder, and Nova; stand up an S3-compatible object endpoint
- Define replication/erasure-coding policies, failure domains, and capacity thresholds
- Validate failure, degraded-mode, and recovery behaviour
Networking & Security
- Design and configure management,
API, provider, tenant, storage, and tunnel networks
- Implement VLANs, overlays, routing, firewalling, and security groups per the organisation's addressing standards
- Apply security-hardening baselines and maintain compliance evidence
Compute & Workloads
- Build and maintain golden images and flavour catalogues (CPU, memory, NUMA, huge pages, GPU profiles)
- Configure host aggregates, availability zones, and placement/scheduling policies
- Restore and re-platform existing Windows and Linux workloads into the current environment
GPU Enablement
- Enable GPU access end-to-end: firmware, BIOS, IOMMU, kernel parameters, drivers, device binding
- Configure PCI passthrough or vGPU as required; configure Nova PCI aliases, traits, and scheduler filters
- Expose GPU resources to Kubernetes worker nodes via appropriate device plugins
Kubernetes & AI/ML Platform
- Deploy and operate a highly available Kubernetes cluster on the private cloud
- Configure CNI, ingress, RBAC, CSI-backed persistent storage, and GPU-schedulable nodes
- Deploy an approved AI/ML workspace platform (notebooks, training jobs, model serving) with per-project isolation and quotas
Identity & Enterprise Integration
- Integrate with the organisation's enterprise directory; configure SSO/federation across OpenStack, Kubernetes, and the AI workspace
- Implement role mapping, least-privilege access, and credential-rotation procedures
Monitoring, Logging & Audit
- Implement monitoring/alerting across hosts, OpenStack services, storage, VMs, Kubernetes, and GPU utilisation
- Configure centralised logging, audit collection, and alert escalation
Resilience, Backup & Recovery
- Define and execute resilience tests (compute/controller/storage/network/Kubernetes/GPU/identity failure scenarios)
- Validate backup/restore against agreed RTO/RPO targets
- Produce test plans, execution evidence, and resilience reports
Restricted-Network Delivery
- Build and maintain internal package repositories and private container registries
- Prepare an offline bill of materials (packages, images, charts, drivers, firmware, licenses)
- Establish controlled import, scanning, signing, and promotion procedures
- Ensure deployment, patching, and recovery can operate without dependency on public internet access
Lifecycle Ownership (post go-live)
- Own technical evolution: capacity expansion, tuning, upgrades, automation
- Act as escalation point for platform issues; lead root-cause analysis
- Maintain architecture documents, SOPs, and runbooks
- Train the internal cloud engineering team to independent competency
Experience we Need.
Core
- 6+ years infrastructure/platform-engineering experience; 3+ years hands-on OpenStack
- Demonstrable delivery of at least one production OpenStack environment from install through acceptance
- Evidence of personal technical contribution, not just project oversight
OpenStack & Kolla Ansible
- Strong working knowledge of Keystone, Nova, Placement, Glance, Neutron, Cinder, Horizon
- Hands-on Kolla Ansible: inventory, globals config, secrets handling, bootstrap, prechecks, deployment, validation, upgrades
- Ability to troubleshoot containerised OpenStack services and deployment failures
Storage
- Hands-on experience with a distributed SDS platform (Ceph preferred)
- Experience integrating storage with Glance/Cinder/Nova and Kubernetes CSI
Linux, Compute & Networking
- Deep Linux systems engineering (kernel, drivers, systemd, networking, storage stacks)
- Strong KVM/QEMU understanding
- VLANs, overlays, routing, firewalling, DNS, NTP, load balancing
Kubernetes & AI Platforms
- Production Kubernetes deployment and operation experience
- GPU scheduling via vendor device plugins
- AI/ML platform deployment (notebooks, training, model serving) strongly preferred
GPU Virtualisation
- PCI/IOMMU passthrough, device binding, driver installation and validation in VMs/containers
Identity, Observability & Automation
- Enterprise directory integration, SSO/RBAC, certificate and secrets management
- Strong Python/shell scripting and Ansible or equivalent IaC experience
Restricted-Network Environments
- Demonstrable experience delivering in air-gapped or restricted-network environments
- Comfortable with formal change control, security review, and audit evidence
Desirable.
- Ceph RBD, CephFS, RADOS Gateway
- Octavia, Barbican, Ironic, Heat, Swift, Manila
- Kubeflow, JupyterHub, MLflow, Ray, KServe, vLLM, or comparable AI/ML platforms
- NVIDIA GPU operators or equivalent vendor tooling
- Multi-site disaster recovery experience
- CKA / CKS / Linux / OpenStack certifications (supplementary to delivery evidence, not a substitute
You will own.
High-level and low-level architecture · Capacity-sizing workbook · Kolla Ansible configuration · IP/VLAN/DNS/firewall matrix · Storage architecture and recovery procedures · GPU enablement report · Kubernetes and AI-platform architecture · Restricted-network BOM · Security baseline · Monitoring/logging catalogue · Backup/RTO/RPO design · Migration and cutover plan · Resilience test plan and evidence · As-built documentation · Runbooks and SOPs · Training materials
📌 Platform Engineer - OpenStack (Delhi)
🏢 Cloudstrats
📍 Delhi