03 Oct
|
Softobiz
|
Hyderabad
03 Oct
Softobiz
Hyderabad
Site Reliability Engineer (SRE) – Platform Team (AWS + Hybrid)
Purpose of Position
We’re looking for a Site Reliability Engineer to join our Platform team and own the reliability of the systems that keep our company trading — from our AWS-based digital platform to the connectivity and point-of-sale systems in every store. You’ll treat reliability as a product: setting SLOs with engineering and product leaders, managing error budgets, and making the call between shipping faster and holding the line.
This is a hands-on role that spans cloud and ground. One hour you’re tuning a Prometheus federation or an Open Telemetry pipeline; the next you’re troubleshooting a store’s connectivity at peak service. You’ll lead P0/P1 incident response, run blameless post-mortems, and build the observability and on-call practices that let a distributed, multi-site business run calmly through its busiest hours.
We care about reliability and cost — uptime at any price isn’t the goal. We want someone who genuinely understands how a technology failure translates into guest experience, crew workflows,
and revenue, and who builds systems and schedules around how restaurants operate.
Our current stack includes:
Technology we use
Current Stack
- Application Stack: TypeScript, NestJS, TypeORM/Sequelize, React / React Native, MySQL / Aurora.
- Cloud (AWS): ECS/Fargate, Lambda, RDS Aurora, CloudFront, API Gateway, SQS/SNS, IAM, VPC networking, Flow Logs.
- IaC & Automation: Terraform and AWS CDK at scale, CloudFormation, GitHub Actions, Bitrise.
- Observability: Prometheus, Grafana, Open Telemetry / OTel Collector, Datadog, AWS CloudWatch.
- Networking: SD-WAN, MPLS, BGP, VLANs and firewall policy; AWS Direct Connect, Site-to-Site VPN, Transit Gateway.
Reliability & Operations
- SLOs & error budgets: service-level indicators, objectives and agreements shaped with product and engineering leaders.
- Incident response: P0/P1 command, structured blameless post-mortems, and error-budget-driven relea
📌 Site Reliability Engineer (SRE) – Platform Team (AWS + Hybrid)
🏢 Softobiz
📍 Hyderabad