- Operate and optimize Kubernetes clusters, Istio service mesh, and Linux-based systems.
- Automate workflows using Go, Python, and Shell scripting.
- Build monitoring and observability solutions with Prometheus, Grafana, and Loki.
- Troubleshoot complex networking, storage, and system performance issues.
- Participate in on-call rotations and postmortem reviews to improve system resilience.
- Partner with AI/ML teams to ensure infrastructure readiness for model training and data pipelines.
What we need from you:
- Experience with Google Cloud, plus IaC tools (Terraform).
- Robust knowledge of microservices, containers (Kubernetes, Docker), and networking.
- SRE mindset with a focus on automation, scalability, and reliability.
- Hands-on experience with PKI, service mesh, and Linux systems administration.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer (Hyderabad)
🏢 ManageServe Technologies
📍 Hyderabad
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.