08 Aug
|
Pythian
|
Hyderabad
Role & responsibilities
Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
Automate workflows using Go, Python, and Shell scripting.
Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
Troubleshoot complex networking, storage, and system performance issues.
Participate in on-call rotations and postmortem reviews to improve system resilience.
Partner with AI/ML teams to ensure infrastructure readiness for model training and data pipelines.
Preferred candidate profile
Experience with Google Cloud, plus IaC tools (Terraform).
Solid knowledge of microservices, containers (Kubernetes, Docker), and networking.
SRE mindset with a focus on automation, scalability, and reliability.
Hands-on experience with PKI, service mesh, and Linux systems administration.
📌 Senior Site Reliability Engineer Hyderabad
🏢 Pythian
📍 Hyderabad