We are looking for an experienced Site Reliability Engineer (SRE) with solid DevOps and production support experience.
Key Responsibilities
Monitor and maintain highly available, scalable applications and infrastructure.
Troubleshoot production issues across applications, databases, and infrastructure.
Manage and monitor Redis and Kafka environments.
Implement monitoring and alerting using Prometheus and Grafana.
Perform root cause analysis (RCA) and resolve performance/reliability issues.
Automate operational tasks and improve system reliability.
Support deployments, incident management, and production releases.
Mandatory Skills
SRE / DevOps
Redis & Kafka
Prometheus & Grafana
Robust production troubleshooting
Monitoring, alerting, RCA, and incident management
Linux/Unix and scripting knowledge