Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure (Chennai)

Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure (Chennai)

24 Sep
|
HCLTech
|
Chennai

24 Sep

HCLTech

Chennai

Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure

Chennai, Tamil Nadu

Job Summary

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and robust production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Key Responsibilities

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems)



SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Skill Requirements

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement,



toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Other Requirements

Microsoft Azure, Red Hat OpenShift, Red Hat Enterprise Linux, Ansible, Red Hat Cluster, Kubernetes Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

📌 Administrator - RedHat Cluster, Ansible, Kubernetes, Microsoft Azure (Chennai)
🏢 HCLTech
📍 Chennai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: administrator - redhat cluster, ansible, kubernetes, microsoft azure (chennai) / chennai

Subscribe to this job alert:

Get the latest job offers by email for: administrator - redhat cluster, ansible, kubernetes, microsoft azure (chennai) / chennai