25 Sep
|
CodeRound AI
|
Bengaluru
25 Sep
CodeRound AI
Bengaluru
Join a fast-growing GenAI startup as an SRE and take ownership of the reliability, scalability, and operational excellence of a platform powering the end-to-end ML lifecycle. You’ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build strong SRE and incident-management practices.
What You'll Do
- Own platform uptime, reliability, scalability, and performance
- Manage Kubernetes clusters, cloud infrastructure, and production environments
- Establish and improve incident response, on-call, RCA, and postmortem processes
- Drive deployment, rollback, and change-management practices
- Handle capacity planning and disaster recovery
- Build and enhance monitoring, alerting, and operational dashboards
- Automate deployments, scaling, and repetitive operational workflows
- Troubleshoot complex production infrastructure and application issues
- Support GPU workloads and model-serving infrastructure
Must have Experience in SRE, DevOps, or Platform Engineering Strong hands-on knowledge of Linux, networking, and Kubernetes Experience with AWS, GCP, or Azure Hands-on experience with Terraform, Helm, and CI/CD Solid troubleshooting, incident management, and problem-solving skills Scripting/programming experience in Python, Bash, or Go
Good to have
Experience with MLOps / AI infrastructure Exposure to GPU clusters and model serving Knowledge of release engineering Understanding of reliability and operational best practices
📌 Site Reliability Engineer (Bengaluru)
🏢 CodeRound AI
📍 Bengaluru