29 Aug
|
Nisum
|
Hyderabad
Senior SRE Kubernetes / AWS
Experience: 6+ Years
Location: Hyderabad
Employment Type: Contract-to-Hire (CTH) 1 Year
Job Overview
We are looking for an experienced Senior Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, AWS, Docker, CI/CD, monitoring, automation, and production support.
The ideal candidate should have solid experience managing Kubernetes-based production environments, troubleshooting critical incidents, improving system reliability, and automating operational processes. Exposure to AI-powered SRE tools and Agentic AI for incident monitoring and remediation is required.
Key Responsibilities
- Design, deploy, manage, and troubleshoot Kubernetes-based applications and infrastructure in production.
- Manage cloud infrastructure and services on AWS.
- Work with Docker and Helm for containerization and Kubernetes deployments.
- Develop, maintain, and troubleshoot CI/CD pipelines using Jenkins and Git-based tools.
- Perform production troubleshooting across Linux systems, Kubernetes, applications, databases, and networking.
- Automate operational tasks using Bash/Shell scripting and Python.
- Monitor applications and infrastructure using tools such as Splunk, Dynatrace, Grafana, Prometheus, ELK, and CloudWatch.
- Handle Incident, Problem, Change, Release, Hot Fix, and ECRQ management following ITIL/ITSM processes.
- Work with ServiceNow, Remedy, Jira, and/or Rally for service and incident management.
- Manage certificate renewals using Venafi / CertiS.
- Support database-related troubleshooting involving MySQL and SQL.
- Work with Akamai for traffic routing and related production issues.
- Support messaging/event-driven systems using Apache Kafka / Axon API.
- Participate in production deployments, release activities, on-call support, and root-cause analysis.
- Identify opportunities for automation and continuously improve system reliability, availability, and performance.
AI / SRE Automation Experience – Required
Candidates should have exposure to AI-assisted SRE and automation, particularly:
- GitHub Copilot for CI/CD pipeline scripting, Kubernetes YAML generation, cloud configuration, documentation, and automation.
- Agentic AI for monitoring logs, identifying incidents, performing automated remediation, and escalating issues.
- Exposure to AI-powered fraud/risk platforms such as Visa Advanced Authorization, Mastercard Decision Intelligence, or Stripe Radar is preferred.
Mandatory Technical Skills
- Kubernetes – Strong hands-on production experience
- AWS
- Docker & Helm
- Jenkins / CI-CD
- Linux – AWS Linux / Ubuntu
- Bash / Shell & Python
- Monitoring – Splunk, Dynatrace, Prometheus, Grafana, ELK, CloudWatch
- ITIL / ITSM – Incident, Problem, Change & Release Management
- Git – Bitbucket / GitHub / GitLab
- Apache Kafka
- Akamai
- Venafi / CertiS
- MySQL / SQL
- AI-assisted SRE / Copilot / Agentic AI exposure
📌 Senior SRE (Hyderabad)
🏢 Nisum
📍 Hyderabad