10 Sep
|
Tata Communications
|
Chennai
10 Sep
Tata Communications
Chennai
About The Company
Tata Communications Redefines Connectivity with Innovation and IntelligenceDriving the next level of intelligence powered by Cloud, Mobility, Internet of Things, Collaboration, Security, Media services and Network services, we at Tata Communications are envisaging a New World of Communications
Broad outline of the Role
- Responsible for managing Linux and Kubernetes infrastructure by handling incidents, service requests, changes, and operational activities while ensuring SLA compliance.
- Diagnose and resolve complex technical issues, maintain platform availability, perform routine maintenance, and support production environments.
- Collaborate with Platform, Network, Storage, Security, and Application teams to ensure timely issue resolution and seamless service delivery.
- Drive operational excellence through proactive monitoring, automation, root cause analysis, documentation, and continuous service improvement.
Minimum Qualifications & Experience
- Graduate with 6-13 years of experience
Other Knowledge & Skills
- Robust expertise in Linux administration, Kubernetes cluster management, and container technologies (Docker/containerd).
- Proficient in troubleshooting Linux, Kubernetes, networking, storage, and application issues, with strong RCA skills.
- Experience with monitoring tools (Prometheus, Grafana, ELK),
shell scripting, and ITSM/ticketing tools ServiceNow.
- Good understanding of security best practices, patch management, Kubernetes upgrades, and cross-functional collaboration.
Key Responsibilities
- Manage and resolve Linux and Kubernetes incidents, service requests, and change requests within SLA.
- Troubleshoot Linux OS issues, including CPU, memory, disk, networking, file systems, and service failures.
- Monitor infrastructure and cluster health using monitoring and logging tools, responding promptly to alerts.
- Perform Linux patching, Kubernetes upgrades, node rollouts, and routine maintenance activities with minimal downtime.
- Conduct root cause analysis (RCA) for recurring issues and implement preventive measures.
- Collaborate with Platform, Network, Storage, and Application teams to resolve complex technical issues.
- Document troubleshooting steps, maintain operational runbooks, and provide regular status updates through the ticketing system.
- Automate routine Linux and Kubernetes operational tasks using shell scripting and standard administration tools.
- Support security and compliance activities, including user access management, certificate renewal, vulnerability remediation, and configuration hardening.
📌 Assistant Manager - Cloud Operations (Chennai)
🏢 Tata Communications
📍 Chennai