07 Aug
|
Allegis Group
|
Hyderabad
07 Aug
Allegis Group
Hyderabad
Role & responsibilities
- Ensure service readiness, maintain runbooks, SOPs, and operational processes.
- Manage monitoring, alerting, incident response, RCA, and problem management.
- Handle backup, recovery, disaster recovery (DR), and validate RTO/RPO targets.
- Manage security operations, vulnerability remediation, and OS patching (Windows/Linux).
- Administer and support GCP and GKE/Kubernetes environments including upgrades and troubleshooting.
- Automate operational tasks using PowerShell, Bash, Ansible, and Python.
- Design and validate resilience and failover testing for multi-region environments.
Preferred candidate profile
- 12+ years in Platform Engineering, Cloud Operations, SRE, or Production Support.
- Strong hands-on experience with GCP, IAM, networking, and observability.
- Strong Kubernetes/GKE administration and troubleshooting experience.
- Robust Windows & Linux administration skills.
- Scripting/automation experience with PowerShell, Bash, Ansible, Python, or Java.
- Knowledge of high-availability and active-active architectures.
- Understanding of resilience patterns such as retry, timeout, circuit breaker, and failover.
- Experience with monitoring tools like Prometheus, Grafana, and OpenTelemetry.
📌 Platform Engineer (Hyderabad)
🏢 Allegis Group
📍 Hyderabad