18 Sep
|
CodeRound AI
|
India
18 Sep
CodeRound AI
India
Job Summary
Provider of platform for machine learning model training and deployment. It offers tools for fine-tuning, deploying, and observing machine learning models. Own platform uptime, reliability, scalability, and performance.
Join a quick-growing GenAI startup as an SRE and take ownership of the reliability, scalability, and operational excellence of a platform powering the end-to-end ML lifecycle. You ll work across Kubernetes, cloud infrastructure, production systems, observability, automation, and model-serving workloads, while helping build robust SRE and incident-management practices.
Responsibilities
Manage Kubernetes clusters, cloud infrastructure, and production environments
Establish and improve incident response, on-call, RCA, and postmortem processes
Drive deployment, rollback,
and change-management practices
Handle capacity planning and disaster recovery
Build and enhance monitoring, alerting, and operational dashboards
Automate deployments, scaling, and repetitive operational workflows
Troubleshoot complex production infrastructure and application issues
Support GPU workloads and model-serving infrastructure
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Engineer Bengaluru (India)
🏢 CodeRound AI
📍 India