24 Sep
|
BuildxPartners
|
Sadar Bazaar
24 Sep
BuildxPartners
Sadar Bazaar
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 4–8 years of experience in SRE, Platform Engineering, or DevOps, with a robust senior IC track record of owning production systems.
- Production-grade expertise with Kubernetes, containers, and cloud platforms (AWS preferred) in distributed-systems environments.
- Hands-on experience defining and operating against SLIs, SLOs, and error budgets, along with leading incident response and blameless postmortems.
- Strong experience with observability tools covering metrics, logging, tracing, dashboards, and alerting, such as Datadog, Prometheus/Grafana, or equivalent platforms.
- Proficient in automation and infrastructure-as-code, using technologies such as Python, Go, Shell, and Terraform.
- Comfortable working with GitOps and CI/CD practices for reliable software and infrastructure delivery.
- Strong understanding of Linux/Unix internals, networking, and cloud-native security fundamentals.
- Strong operational rigor, ownership mindset, and ability to communicate clearly in both written and verbal formats, including during high-pressure incidents.
Requirements
Responsibilities
- Own the end-to-end reliability of Syfe’s production platform, ensuring high availability, performance, and operational stability.
- Work with a Kubernetes-native, multi-region platform across Singapore, Hong Kong, and Sydney, supporting a regulated digital wealth-management product.
- Operate as a senior individual contributor, defining measurable reliability standards and building systems, automation, and processes to maintain them.
- Define and drive SLIs, SLOs,
and error budgets across critical services, partnering with product and engineering teams to balance velocity and stability.
- Own the on-call, escalation, and incident-response program, including incident command, blameless postmortems, RCA tracking, and reducing MTTD and MTTR.
- Own the reliability of the AWS EKS-based deployment platform, including GitOps with ArgoCD, Helm-based release configuration, and Infrastructure as Code using Terraform/OpenTofu.
- Ensure deployments are safe, progressive, and reversible, with strong rollout and rollback mechanisms.
- Build and continuously improve the observability stack using tools such as Datadog, Grafana, VictoriaMetrics, and ClickHouse.
- Develop meaningful dashboards, actionable alerts, and monitoring practices while reducing unnecessary alert noise.
- Lead capacity planning, scalability analysis, failure-mode analysis, disaster recovery, and business continuity planning across regions.
- Plan and conduct game days and chaos engineering exercises to validate system resilience and recovery processes.
- Identify operational toil and eliminate it through automation, self-service tooling, and improved engineering practices.
- Strengthen production safety, including deployment guardrails, rollback processes, and secrets management using HashiCorp Vault.
- Partner with engineering teams to improve production readiness and service reliability before systems are deployed to production.
- Drive reliability through hands-on engineering, design reviews, production-readiness reviews, runbooks, and technical documentation.
- Mentor engineers and influence engineering teams to adopt SRE best practices and build a strong culture of production ownership.
📌 Site Reliability Engineer (Sadar Bazaar)
🏢 BuildxPartners
📍 Sadar Bazaar