We are looking for an experienced Site Reliability Engineer (SRE) with at least 5 years of hands-on experience in AWS, Kubernetes, Terraform, SQL Server, Datadog, and Java-based service reliability.
You will ensure the stability, performance, and reliability of our cloud-native Hospitality Revenue Management System (RMS) — a multi-tenant SaaS platform serving hotels globally.
This role involves deep collaboration with engineering to fix reliability issues in Java microservices, build resilient cloud infrastructure, and optimize end-to-end performance of forecasting, rate publishing, and PMS/CRS integrations.
What you’ll be doing...
- Reliability & Platform Resilience
- Define and manage SLIs/SLOs for RMS services (latency, availability, job success, data freshness).
- Improve reliability of Java-based microservices by:
- Fixing resource leaks, threading issues, timeouts, connection pool exhaustion
- Improving resilience patterns (retries, backoff, circuit breakers
- Profiling memory/CPU bottlenecks
- Resolve performance issues impacting:
- Forecast generation cycles
- Rate push APIs
- PMS/CRS connector API calls.
- AWS Cloud Engineering
- Operate and optimize AWS services:
- EC2, EKS (Kubernetes), S3, CloudFront
- RDS for SQL Server
- Lambda, SNS/SQS, API Gateway
- VPC, Security Groups, NAT, Transit Gateway
- Implement cost optimization (compute efficiency, right-sizing, egress reduction).
- Infrastructure as Code (Terraform)
- Build, version, and maintain infrastructure using Terraform:
- Reusable modules
- Remote state + state locking
- Drift management
- Multi-workplace automation
- Enforce GitOps principles and automated provisioning.
- Observability & Incident Response (Datadog – Mandatory)
- Build end-to-end observability using Datadog, including:
- Metrics, logs, traces
- JVM profiling (GC, heap, threads)
- Datadog APM for Java services
- Datadog Synthetics for API uptime (PMS, CRS, OTA)
- Datadog RUM (if applicable)
- Configure SLO-based alerting to reduce
📌 Site Reliability Engineer (Pune)
🏢 SAS
📍 Pune