OpenShift & GitLab (Mumbai)

OpenShift & GitLab (Mumbai)

29 Aug
|
3i Infotech
|
Mumbai

29 Aug

3i Infotech

Mumbai

- Monitor OpenShift cluster health, including control plane components, API responsiveness, kubelet, node status, cluster operators, etcd health, CPU, memory, disk, network, pod restarts, and overall platform availability using enterprise monitoring tools.

- Perform daily/shift health checks for nodes, pods, cluster operators, disk pressure, certificate validity, scheduled jobs, and overall platform status.

- Monitor platform alerts, acknowledge incidents within defined SLAs, execute approved L1 runbooks, and coordinate timely incident response.

- Troubleshoot and resolve control plane issues including API instability, etcd quorum/latency, node NotReady conditions, operator failures, CNI/OVN-Kubernetes networking issues, Ingress/Route failures, admission webhook problems, and image registry outages.

- Collect, analyze, and provide diagnostics using oc/kubectl, pod descriptions, logs, events, manifests, must-gather, oc adm inspect, audit logs, API server metrics, etcd metrics, node logs, and journalctl.

- Perform first-line remediation by restarting pods/services, deleting failed pods, scaling workloads, clearing PVC locks, rotating authorized tokens, and executing documented recovery procedures.

- Escalate unresolved or complex issues to L2/L3 with complete diagnostics, perform Root Cause Analysis (RCA), workload impact assessment, and recommend corrective and preventive actions.

- Monitor and administer Machine API, Autoscaler, node lifecycle operations including drain, cordon, uncordon, reimage, node replacement, and cluster resource optimization.

- Plan, coordinate, and execute OpenShift cluster upgrades, z-stream patching, Operator (OLM) upgrades, cluster add-on maintenance, compatibility validation, canary deployments, rollback planning, and post-upgrade verification.

- Support scheduled maintenance by performing pre/post patch validation, node drain/uncordon, controlled reboots, cluster health validation, and maintenance documentation.

- Administer and troubleshoot OpenShift networking including OVN-Kubernetes/CNI, NetworkPolicies, Egress/IPs, ExternalIPs, Routes, Ingress Controllers, HAProxy/load balancing, DNS, Services, EndpointSlices, and network connectivity.





- Administer storage services including CSI drivers, StorageClasses, PVC/PV lifecycle, reclaim policies, storage performance, registry storage, ODF/ODF-LVM/Ceph platforms, and persistent storage troubleshooting.

- Monitor and manage the internal image registry, registry storage, image pull policies, registry trust, image cleanup, garbage collection, artifact repositories, replication, retention, and geo-distribution.

- Validate cluster backups, etcd snapshots, backup reports, restore procedures, Velero integration, disaster recovery readiness, DR runbooks, and recovery testing.

- Execute approved routine platform administration tasks including ConfigMap and Secret updates, Route/Ingress management, namespace/project creation, service account support, and simple configuration changes.

- Administer RBAC, enforce least-privilege access, manage quotas, LimitRanges, Pod Disruption Budgets (PDBs), priority classes, eviction policies, and multi-tenancy best practices.

- Manage TLS certificates including API, Ingress, internal trust chains, certificate rotation, renewal, expiry validation, and certificate lifecycle management.

- Monitor, administer, and optimize CI/CD platforms including Jenkins, GitLab CI, GitLab Runners, Tekton, Argo Workflows, controllers, agents, runners, job queues, artifact retention, caching, and execution capacity.

- Troubleshoot CI/CD pipeline failures involving credentials, registry access, network connectivity, runner capacity, image pull failures, flaky tests, deployment failures, pipeline parameters, and build logs.

- Develop, maintain, and optimize reusable CI/CD pipeline templates, shared libraries, GitOps repository structures, Infrastructure-as-Code automation using Terraform/Ansible, approval workflows, and policy validation.

- Support developers by troubleshooting application deployments, image pull issues,



liveness/readiness probe failures, resource limits, deployment errors, namespaces, service account tokens, and cluster access issues.

- Operate and optimize the monitoring and logging platforms including Prometheus, Alertmanager, Grafana, Loki/EFK, log forwarding, indexing, retention, scrape targets, dashboards, recording rules, SLO/SLA reporting, and alert optimization.

- Enforce platform security by administering SCC/Pod Security Admission (PSA), RBAC reviews, image security, registry trust, service mesh security, Vault/KMS integration, secrets management, agile credentials, key rotation, authentication monitoring, and compliance controls.

- Implement release strategies including blue-green deployments, canary deployments, feature flags, rollout policies, automated rollback, and deployment validation.

- Optimize platform performance through resource requests/limits tuning, HPA/VPA optimization, node right-sizing, workload bin-packing, capacity forecasting, and compute, storage, and network cost optimization.

- Maintain operational documentation, update patch, backup, maintenance, and incident records, ensure compliance with organizational standards, and support change management processes.

- Coordinate with L2/L3 teams during incidents, maintenance windows, platform upgrades, disaster recovery exercises, and production support while ensuring adherence to operational procedures, SLAs, and security policies.

- GitLab Administration & Monitoring: Administer, configure, monitor, secure, troubleshoot, upgrade, patch, and maintain the complete GitLab platform, including GitLab Server, GitLab CI/CD, GitLab Runners, Gitaly, PostgreSQL, Redis, Praefect (where applicable), Container Registry, Package Registry, LDAP/Active Directory integration, SSL/TLS certificates, backups and disaster recovery, repository administration, user and group management, RBAC, authentication, performance tuning, high availability, monitoring, capacity planning, log analysis, security hardening, integrations with external systems, and overall platform health, availability, and lifecycle management.

Contact: [email protected]

📌 OpenShift & GitLab (Mumbai)
🏢 3i Infotech
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: openshift & gitlab (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: openshift & gitlab (mumbai) / mumbai