Seeking a Site Reliability Engineer to join a dedicated operational support team for Kafka and infrastructure engineering organization. engagement with potential to evolve into project-based work as the team matures. The role covers an environment of ~600 servers plus Kubernetes pods/services, primarily on-premises bare metal infrastructure with a planned transition toward virtualized environments.
- Perform vulnerability management and remediation, including patching within defined SLAs (critical vulnerabilities)
- Execute OS and Kafka upgrades, reinstalls, and migrations across the environment
- Conduct health checks, monitoring, and proactive alert handling across ~600 servers and Kubernetes-based services
- Handle change requests and day-to-day operational support tasks
- Manage customer support requests and escalations, including on-call rotation coverage
- Coordinate with vendors and internal security teams to assess impact and apply fixes for identified vulnerabilities
- Develop and improve automation and repeatable procedures to reduce manual operational overhead
- Support and administer AWS cloud infrastructure (EC2, VPC, IAM, S3, EKS/Kubernetes, CloudWatch)
as the environment transitions from on-premises bare metal toward virtualized/cloud infrastructure
- Design, implement, and maintain infrastructure-as-code (Terraform/CloudFormation) for provisioning and managing AWS resources supporting Kafka and platform services
- Own infrastructure reliability and capacity planning across hybrid (on-prem + AWS) environments, including scaling, failover, and disaster recovery readiness
- Implement and maintain observability tooling (CloudWatch, Prometheus/Grafana, or equivalent) for infrastructure and application-level monitoring
- Ensure AWS workplace security posture - IAM least-privilege access, security group/network hardening, and compliance with vulnerability remediation SLAs
- Collaborate with the DevOps/platform team on CI/CD pipeline support for infrastructure changes and deployments
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 SRE - Kafka/Infrastructure Operational Support Professional (Bengaluru)
🏢 Tavant
📍 Bengaluru
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.