13 Sep
|
Trigent Software
|
Noida
13 Sep
Trigent Software
Noida
JD:
Hands-on experience in an SRE, DevOps, Platform Engineering, or Production Support role.
Solid working knowledge of Kubernetes: workloads, networking (services, ingress, CNI), RBAC, Helm, troubleshooting pod/node-level issues, and kubectl proficiency.
Solid experience with AWS services: EC2, EKS, IAM, VPC, S3, CloudWatch, Route53, ELB/ALB, and Auto Scaling.
Deep understanding of SRE principles: SLIs/SLOs/SLAs, error budgets, toil reduction, blameless postmortems, and reliability engineering practices.
Proficiency in observability tooling and concepts: metrics, logs, distributed tracing, alerting, and dashboarding (Prometheus, Grafana, Datadog, Current Relic, ELK, Splunk, or similar).
Strong Linux fundamentals and scripting ability (Bash, Python, or Go).
Familiarity with CI/CD pipelines, Infrastructure as Code (Terraform, CloudFormation), and Git-based workflows.
Excellent debugging mindset and the ability to reason about distributed systems failure modes.
Robust written and verbal communication skills, with the ability to interact with both technical and non-technical users.
📌 Site Reliability Engineering Noida
🏢 Trigent Software
📍 Noida