21 Aug
|
Tavant
|
Bengaluru
seeking a Site Reliability Engineer to join a dedicated operational support team for Booking.com's Kafka and infrastructure engineering organization. engagement with potential to evolve into project-based work as the team matures. The role covers an environment of ~600 servers plus Kubernetes pods/services, primarily on-premises bare metal infrastructure with a planned transition toward virtualized environments. The engineer will operate as part of a 4-member India-based team providing first-line operational support, reducing the load on Booking's core Kafka/infra engineering teams, with escalations only routed to Booking when they cannot be resolved internally
Roles and Responsibilities
- Perform vulnerability management and remediation, including patching within defined SLAs (critical vulnerabilities)
- Execute OS and Kafka upgrades, reinstalls, and migrations across the workplace
- Conduct health checks, monitoring, and proactive alert handling across ~600 servers and Kubernetes-based services
- Own incident response and troubleshooting for operational issues, escalating to Booking engineering only when unresolved at the first line
- Handle change requests and day-to-day operational support tasks
- Manage customer support requests and escalations, including on-call rotation coverage
- Coordinate with vendors and internal security teams to assess impact and apply fixes for identified vulnerabilities
- Develop and improve automation and repeatable procedures to reduce manual operational overhead
- Participate in onboarding/knowledge transfer during ramp-up to build operational maturity on Booking's Kafka and infrastructure stack
- Support and administer AWS cloud infrastructure (EC2, VPC, IAM, S3, EKS/Kubernetes, CloudWatch) as the environment transitions from on-premises bare metal toward virtualized/cloud infrastructure
- Design, implement, and maintain infrastructure-as-code (Terraform/CloudFormation) for provisioning and managing AWS resources supporting Kafka and platform services
- Own infrastructure reliability and capacity planning across hybrid (on-prem + AWS) environments, including scaling, failover, and disaster recovery readiness
- Implement and maintain observability tooling (CloudWatch, Prometheus/Grafana, or equivalent) for infrastructure and application-level monitoring
- Ensure AWS environment security posture — IAM least-privilege access, security group/network hardening, and compliance with vulnerability remediation SLAs
- Collaborate with the DevOps/platform team on CI/CD pipeline support for infrastructure changes and deployments
📌 SRE — Kafka/Infrastructure Operational Support (Bengaluru)
🏢 Tavant
📍 Bengaluru