The Senior Site Reliability Engineer (SRE) will build, automate, and operate Hanwha Vision's global cloud platform. We treat operations as a software engineering problem. Your goal is to eliminate manual tasks and replace them with reliable, automated systems.
Key Responsibilities
Automation: Write clean code (Go or Python) to automate cloud operations and deployment pipelines.
Kubernetes Engineering: Build and manage large-scale Amazon EKS clusters, including networking, security, and scaling (using Karpenter).
Infrastructure as Code: Write and maintain modular Terraform and Helm templates.
Incident Management: Lead troubleshooting for critical system outages. Write clear, blameless post-mortems to prevent issues from happening again.
Observability: Define SLIs/SLOs. Set up monitoring, logging, and on-call alerts using Datadog and PagerDuty.
Security & Compliance: Build secure, self-service tools (like automated JIT AWS access)
to meet SOC2 requirements.
Required Technical Skills
10+ years of professional experience in SRE, DevOps, or Systems/Infrastructure Engineering.
Coding: Solid programming skills in Go or Python.
Containers: Production experience running and scaling Kubernetes (EKS preferred).
IaC: Expert-level knowledge of Terraform.
Monitoring: Experience with Datadog, Prometheus, Grafana, or PagerDuty.
Database/Messaging (Preferred): Basic understanding of DynamoDB, Kakfa/MSK, or caching layers.
Languages: Robust written English proficiency is required to collaborate with global teams.