22 Aug
|
Insight Global
|
Hyderabad
22 Aug
Insight Global
Hyderabad
Required Skills & Experience
- 5+ years of experience in DevOps, MLOps, Site Reliability Engineering (SRE), Platform Engineering, or Cloud Engineering, with robust hands-on production Kubernetes experience.
- 2-4 years of experience building, deploying, and supporting containerized applications using Docker, Helm, ArgoCD, CI/CD pipelines, and GitFlow methodologies.
- Experience with Kubeflow or similar machine learning workflow orchestration platforms.
- Experience supporting applications and services within Azure environments, including familiarity with Azure OpenAI, Azure AI Search, RAG architectures, and LLM-powered services.
- Experience operating production systems with monitoring, alerting, incident response, operational support, observability tooling (logs, metrics, dashboards, alerts), and runbook creation.
- Experience with OpenCity or similar internal developer platforms and cloud enablement tools.
- Strong troubleshooting skills across Kubernetes, networking, infrastructure, secrets management, cloud dependencies, and application services.
- Understanding of reliability engineering principles including scaling, health probes, retries, timeouts, failover, recovery, rollback strategies, and graceful degradation.
Key Responsibilities:
- Maintain and support Kubernetes-based application deployments across development and production environments.
- Manage and troubleshoot GitFlow-based CI/CD pipelines used by multiple engineering teams.
- Support containerized workloads using Docker, Helm Charts, ArgoCD, and Kubernetes.
- Perform platform maintenance activities including security upgrades, patching, dependency updates, and infrastructure improvements.
- Partner with engineering teams across North America and Eastern Europe to resolve production issues and operational challenges.
- Troubleshoot application, infrastructure, networking, and deployment issues across Kubernetes clusters.
- Support AI services, model evaluation workflows, RAG pipelines, automation services, and related cloud-native applications.
- Improve reliability through monitoring, alerting, health checks, readiness probes, retries, scaling strategies, and rollback procedures.
- Build and maintain operational dashboards, logging, metrics, alerts, and runbooks.
- Investigate incidents, identify recurring failure patterns, and implement operational improvements to increase platform resiliency.
- Provide day-to-day support for engineering teams during active troubleshooting scenarios, serving as a technical resource when production issues arise.
- Collaborate with platform, AI, and operations teams to establish operational best practices for running production services.
📌 DevOps/MLOps/Platform Engineer (Hyderabad)
🏢 Insight Global
📍 Hyderabad