27 Sep
|
Cognisive Solutions
|
India
27 Sep
Cognisive Solutions
India
Job Summary
This is a hands-on execution role on a live network. Youll be running deployments, upgrades, and changes across pre-production and production Kubernetes clusters - following documented procedures precisely, verifying results, and escalating quickly when something doesnt match expected outcomes. Our production clusters carry live traffic for operator partners at a 99.999% availability bar.
Discipline matters more here than invention: doing the documented thing correctly, at 2am, on the fifth cluster in a sequence, with the same care as the first.
Key Skills
- Kubernetes
- AWS
- GCP
- Terraform
- Ansible
- Python
- Bash
- CI/CD
- Prometheus
Responsibilities
- Execute deployments, rollouts, and rollbacks across pre-production and production Kubernetes clusters, following documented runbooks and change-management procedures
- Run cluster upgrades and node lifecycle operations in a defined sequence, validating health at each gate before proceeding
- Promote changes through environments in order - dev to pre-prod to production - and verify each stage
- Work within approved change windows and maintenance schedules, with escalation paths and rollback plans prepared in advance
- Keep pre-prod a faithful mirror of production, and flag drift between them as a defect
- Participate in a follow-the-sun on-call rotation covering IST hours alongside global teams
- Respond to alerts using established runbooks: triage, mitigate, escalate to the on-call engineer or SME when the runbook runs out
- Communicate clearly during incidents - status, impact, actions taken, and what is needed
- Automate repetitive parts of the operational workload - scripts, pipeline steps, Ansible/Terraform changes - within established team patterns
- Improve runbooks as you execute them; fix unclear or incorrect procedures while they are fresh
- Log every production change accurately with what was done, when, by whom, and verification results
Requirements
- 5+ years in DevOps, SRE, infrastructure, or production operations roles, with substantial time operating production systems
- Solid hands-on Kubernetes operations: deployments, rollouts, troubleshooting pods and nodes, reading logs and events to isolate faults under pressure
- Working experience with AWS and GCP in a production context
- Comfortable with infrastructure as code and configuration management (Terraform, Ansible, or equivalent)
- Scripting proficiency in Python or Bash sufficient to automate operational tasks
- Demonstrated production on-call experience in a change-controlled workplace with formal change processes
- Linux systems fundamentals: networking, storage, processes, systemd, performance basics
- Precision and follow-through: read procedures fully before starting, verify rather than assume, and escalate immediately when something looks wrong
Nice to Have
- Telecom domain knowledge - 5G core, cloud-native network functions, IMS, or telecom operations
- Experience managing databases and applications with 99.999% high availability
- Observability tooling: Prometheus, Grafana, ELK - building dashboards and tuning alerts operationally
- GitOps workflows (ArgoCD, Flux) and progressive delivery
- Edge or on-premises infrastructure alongside public cloud, particularly sites with constrained connectivity
- Formal change management or ITIL-style process experience
- FinOps and cloud cost awareness in day-to-day operational decisions
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Sr DevOps Engineer (India)
🏢 Cognisive Solutions
📍 India