16 Sep
|
Tata Communications
|
Tamil Nadu
16 Sep
Tata Communications
Tamil Nadu
Job Description
1 Year Contract role
n
Role: Kubernetes Manager
n
Level: L3 / Senior Kubernetes Administration & Operations
n
Primary Focus: Linux Administration, Kubernetes Administration, Production Operations, SRE, Monitoring, Security, Upgrades and Customer Escalation
n
Mandatory Skills
n
n
- CKA certification (Preferred)
n
- Robust/deep knowledge of Linux administration and troubleshooting
n
- Strong hands-on experience in Kubernetes administration and L3 troubleshooting
n
- Knowledge on virtualization like vSphere, OpenStack, Ovirt , KVM
n
- Production experience in Kubernetes cluster operations, incident management and troubleshooting
n
- Strong knowledge of Kubernetes architecture, control plane, worker nodes, networking, storage and workloads
n
- Ability to handle P1/P2/P3 incidents and customer escalations within SLA
n
- Hands-on experience with Kubernetes upgrades, patching and maintenance
n
- Strong understanding of container runtime technologies:
n
- containerd
n
- CRI-O
n
- Docker
n
- CRI
n
- Experience with container registries, preferably Harbor
n
n
Monitoring & Observability
n
n
- Configure and administer Prometheus
n
- Configure and manage Alert manager
n
- Grafana dashboard creation and monitoring
n
- Kubernetes cluster health monitoring
n
- Alert investigation and resolution within SLA
n
- Alert tuning and noise reduction
n
- Identify duplicate/repeated alerts and optimize alert rules
n
- Experience with ELK/EFK and Fluentd
n
- Log collection, analysis and troubleshooting
n
n
CI/CD & DevOps
n
Hands-on experience with:
n
n
- Jenkins
n
- GitLab CI/CD
n
- Git
n
- Container image build and deployment
n
- Kubernetes deployment automation
n
- CI/CD troubleshooting
n
n
Kubernetes Operations
n
Responsible for:
n
n
- Day-to-day Kubernetes cluster administration
n
- L3 troubleshooting of Kubernetes issues
n
- Kubernetes upgrade and patching activities
n
- Cluster health checks
n
- Node and pod troubleshooting
n
- Control-plane and worker-node troubleshooting
n
- Kubernetes networking troubleshooting
n
- Kubernetes storage troubleshooting
n
- Resource and performance analysis
n
- Incident/problem management
n
- Root Cause Analysis (RCA)
n
- Vulnerability remediation
n
- Production change implementation
n
- Customer-specific Kubernetes requirements
n
n
Security & Vulnerability Management
n
n
- Kubernetes security best practices
n
- Vulnerability identification and remediation
n
- Container image vulnerability management
n
- Registry security
n
- Kubernetes RBAC
n
- Network/security policy concepts
n
- SSL/TLS certificate management
n
- Experience with security tools such as Gatekeeper
n
- Ability to coordinate vulnerability fixes across OS, Kubernetes and container layers
n
n
Networking
n
Strong understanding of:
n
n
- DNS
n
- TCP/IP and L3 networking
n
- Load Balancers
n
- SSL/TLS termination
n
- Ingress
n
- Kubernetes Services
n
- Network troubleshooting
n
- Ingress Controllers
n
- Istio Gateway
n
- Service Mesh concepts
n
- NGINX
n
- Kong API Gateway
n
n
Storage
n
Good understanding of Kubernetes storage and underlying infrastructure:
n
n
- NFS / File storage
n
- Block storage
n
- Object storage
n
- PV/PVC
n
- Storage Class
n
- CSI
n
- Storage troubleshooting
n
- Mount and I/O issues
n
- Storage performance concepts
n
n
Platform Components
n
Working knowledge of:
n
Component Required Knowledge
n
Istio
n
Service Mesh, Gateway, traffic management
n
NGINX
n
Ingress / reverse proxy
n
Kong
n
API Gateway
n
Redis
n
Cache / datastore
n
PostgreSQL
n
Database fundamentals & troubleshooting
n
Kafka
n
Messaging / event streaming
n
RabbitMQ
n
Message broker
n
Keycloak
n
IAM / authentication
n
Gatekeeper
n
Kubernetes policy enforcement
n
Harbor
n
Container registry
n
Prometheus
n
Monitoring
n
Grafana
n
Visualization
n
Alertmanager
n
Alert management
n
ELK
n
Centralized logging
n
Fluentd
n
Log collection
n
Customer & Operations Responsibilities
n
n
- Provide 24×7/on-call customer support as required
n
- Handle customer escalations at OS and Kubernetes levels
n
- Understand and resolve customer production issues
n
- Participate in incident bridges and technical discussions
n
- Perform RCA for critical incidents
n
- Track and resolve tickets within defined SLA
n
- Handle day-to-day customer requirements
n
- Coordinate with development, compute, network, storage, security and application teams
n
- Ensure production changes follow the approved change-management process
n
- Prepare technical documentation and operational procedures
n
📌 Manager (Tamil Nadu)
🏢 Tata Communications
📍 Tamil Nadu