22 Sep
|
Mirantis
|
Hyderabad
22 Sep
Mirantis
Hyderabad
Job Description
We are looking for a senior Kubernetes-focused DevOps/SRE engineer to own both the developer platform and a customer-facing production region of a multi-tenant control plane for enterprise GPU infrastructure. You will build and run the environments, pipelines, and infrastructure tooling that our engineering teams across the US, Europe, and APAC depend on to ship daily.
This role spans both sides of the line. You will make our development, test, and pre-production clusters fast and reproducible, harden the Helm and CI/CD path from commit to release, and carry operational ownership — including on-call — for one of our smaller customer-facing production regions, under real availability commitments. That production experience makes you the internal expert on how k0rdent AI is deployed and operated — the person other teams consult, including the teams running our larger regions. Working within an agile framework, you will directly shape how quickly and safely changes reach production, and be accountable for how they behave once there.
Main Responsibilities:
- Own the Kubernetes footprint across development, CI, pre-production, and one customer-facing production region — local kind clusters, shared dev and QA environments, and multi-cluster/multi-region topologies.
- Operate your production region against defined SLOs: capacity and upgrade planning, patching, backup and restore, disaster recovery drills, and participation in an on-call rotation.
- Lead incident response for your region — detection, mitigation, customer-impact assessment, root-cause analysis,
and blameless postmortems that feed fixes back into the platform.
- Build and maintain Helm charts and umbrella releases for the platform's services and dependencies, including versioning, values hygiene, and upgrade paths.
- Own the CI/CD pipelines end to end — build, test, image publishing, chart packaging, release cutting, and hotfix/backport flows.
- Automate workplace bootstrap and seeding so any engineer can bring up a full stack — control plane, identity, gateway, database, workflow engine — with one command.
- Operate and troubleshoot the supporting stack across test and production: PostgreSQL, Temporal, Keycloak, API gateway, message broker, and observability components.
- Build observability and diagnostics — metrics, dashboards, alerting, log and audit access — that serve both engineering environments and production operations.
- Consult with product teams and with the teams operating our larger regions on deployment topology, GPU and resource scheduling, RBAC, networking, and failure modes; validate upgrade and migration procedures and hand over runbooks.
- Enforce security and tenant isolation in production: least-privilege access, secret handling, certificate and TLS lifecycle, image and dependency scanning, and audit evidence for compliance reviews.
- Drive infrastructure as code and repeatability — no snowflake environments, no undocumented manual steps.
- Mentor engineers on Kubernetes and operational practice, and raise the team's bar through review and documentation.
📌 Senior Site Reliability Engineer (Hyderabad)
🏢 Mirantis
📍 Hyderabad