Hello All,
Greetings from ZettaMine!!!
Job Title: Senior Cloud Infrastructure / Kubernetes Lead AI Datacenter
Location: Chennai
Experience: 7+ Years
Notice Period: Immediate Joiners
Role Objective
Act as a senior technical lead responsible for managing, operating and scaling AI Datacenter and Cloud infrastructure, ensuring high availability, operational stability, security and compliance while supporting the growth of cloud and AI platforms.
The role requires solid hands-on expertise in Linux administration, Kubernetes, container platforms, cloud operations, infrastructure support, monitoring, troubleshooting and automation, along with experience providing L3 production support.
Key Responsibilities
- Manage and support Linux servers and Kubernetes environments across enterprise cloud and AI Datacenter infrastructure.
- Perform infrastructure deployments, patching, upgrades, monitoring, performance tuning and troubleshooting.
- Provide L3 technical support within the System Frontline Team Cloud Ops Support (L2/L3) function.
- Build, operate, maintain and scale sovereign cloud and AI infrastructure for Datacenter initiatives.
- Ensure high availability, reliability, performance and operational stability of infrastructure platforms.
- Monitor system and application health and proactively identify and resolve infrastructure issues.
- Perform root cause analysis for recurring incidents and implement permanent corrective actions.
- Follow and manage Incident, Problem and Change Management processes in accordance with IT service management practices.
- Ensure SLA compliance and provide timely resolution of critical infrastructure incidents.
- Develop and maintain technical documentation, operational procedures, runbooks and knowledge articles.
- Conduct knowledge transfer sessions and mentor L2/L3 support engineers.
- Drive infrastructure automation, continuous improvement and operational efficiency.
- Identify opportunities to automate repetitive operational and support activities using scripting and DevOps practices.
- Collaborate with infrastructure, cloud, security, application, delivery and platform engineering teams.
- Work closely with business stakeholders and support GTM capability enhancement initiatives.
- Support infrastructure capacity planning, scalability and platform expansion for AI and cloud workloads.
- Ensure infrastructure operations comply with organizational security, governance and regulatory requirements.
- Support regulatory and compliance readiness related to DPDP, RBI and CERT-In requirements.
- Participate in infrastructure audits,
security reviews and compliance assessments.
- Maintain operational dashboards, SLA/KPI reports and infrastructure health metrics.
- Support disaster recovery, business continuity and high-availability requirements for critical infrastructure.
Required Technical Skills
Linux Administration
- Strong hands-on experience in Linux Administration.
- Experience with Linux server installation, configuration, patching, upgrades and troubleshooting.
- Strong understanding of system performance, CPU, memory, storage, networking and process management.
- Experience with shell scripting and Linux automation.
Kubernetes & Container Platforms
- Strong hands-on experience with Kubernetes administration and operations.
- Experience managing Kubernetes clusters, nodes, workloads, pods, services and deployments.
- Knowledge of container technologies such as Docker / containerd.
- Experience with Kubernetes troubleshooting, upgrades, scaling and performance optimization.
- Understanding of container networking, storage and security.
Cloud Operations
- Strong experience in Cloud Operations and Infrastructure Support.
- Experience operating and supporting cloud infrastructure and enterprise platforms.
- Experience with infrastructure availability, capacity management and performance monitoring.
- Exposure to sovereign cloud or private cloud environments is highly desirable.
- Experience supporting large-scale Datacenter infrastructure is preferred.
Monitoring & Troubleshooting
- Strong experience in infrastructure monitoring and observability.
- Ability to troubleshoot complex Linux, Kubernetes, networking and infrastructure issues.
- Experience with monitoring, logging and alerting platforms.
- Strong root cause analysis and problem management skills.
Automation & DevOps
- Experience with infrastructure and operational automation.
- Strong scripting knowledge using Shell/Bash, Python or similar scripting languages.
- Understanding of DevOps and CI/CD practices.
- Experience automating deployments, patching, monitoring and operational tasks.
- Infrastructure-as-Code experience such as Terraform or Ansible is preferred.
IT Operations & Support
- Experience providing L2/L3 production support.
- Strong understanding of Incident, Problem and Change Management.
- Experience working against defined SLAs and operational KPIs.
- Experience with production troubleshooting and critical incident management.
Compliance & Governance
- Understanding of infrastructure security, governance and compliance requirements.
- Experience supporting regulatory and compliance requirements such as:
- DPDP
- RBI
- CERT-In
- Knowledge of security hardening, access management, audit requirements and operational controls is preferred.
Documentation & Knowledge Management
- Strong documentation and knowledge management skills.
- Ability to create and maintain:
- SOPs
- Runbooks
- Troubleshooting guides
- Architecture/operational documentation
- Knowledge articles
- Experience conducting knowledge transfer and mentoring support teams.
Soft Skills
- Strong technical leadership and problem-solving skills.
- Excellent communication and stakeholder management skills.
- Ability to work effectively with cross-functional technical and business teams.
- Strong ownership and ability to manage critical production issues.
- Ability to work in a fast-paced Datacenter and Cloud Operations environment.
- Strong focus on operational excellence, reliability and continuous improvement.
Preferred Skills
- Experience with AI Datacenter / AI Infrastructure.
- Experience with GPU infrastructure and AI/ML platforms.
- Experience with Sovereign Cloud / Private Cloud environments.
- Kubernetes certifications such as CKA / CKAD / CKS.
- Linux certifications such as RHCE / RHCSA.
- Experience with OpenShift or other enterprise Kubernetes platforms.
- Experience with Terraform, Ansible, Helm and GitOps.
- Experience with monitoring tools such as Prometheus, Grafana, ELK/EFK or equivalent observability platforms.
- Experience with cloud platforms such as AWS, Azure or GCP.
- Knowledge of infrastructure security and compliance frameworks.
- Experience with high-availability, disaster recovery and business continuity solutions.
Experience
- Senior-level experience in Linux, Kubernetes, Cloud Infrastructure or Datacenter Operations.
- Strong experience handling enterprise production environments and L2/L3 infrastructure support.
- Proven experience managing critical infrastructure with high availability and SLA requirements.
Interested candidates kindly share your updated CV to
[email protected] or WhatsApp to (phone hidden). Thanks & Regards,
Gangadhar.B
📌 Senior Cloud Infrastructure Kubernetes Lead AI Datacenter (Bengaluru)
🏢 Zettamine Labs
📍 Bengaluru