26 Aug
|
CodeRound AI
|
Bengaluru
26 Aug
CodeRound AI
Bengaluru
Job Summary
Provider of platform for machine learning model training and deployment. It offers tools for fine-tuning, deploying, and observing machine learning models. Own platform uptime, reliability, scalability, and performance.
Join a fast-growing GenAI startup as an SRE and take ownership of the reliability, scalability, and operational excellence of a platform powering the end-to-end ML lifecycle. You ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build solid SRE and incident-management practices.
Responsibilities
- Manage Kubernetes clusters, cloud infrastructure, and production environments
- Establish and improve incident response, on-call, RCA, and postmortem processes
- Drive deployment, rollback,
and change-management practices
- Handle capacity planning and disaster recovery
- Build and enhance monitoring, alerting, and operational dashboards
- Automate deployments, scaling, and repetitive operational workflows
- Troubleshoot complex production infrastructure and application issues
- Support GPU workloads and model-serving infrastructure
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer (Bengaluru)
🏢 CodeRound AI
📍 Bengaluru