- Seeking a Site Reliability Engineer to join a dedicated operational support team for Kafka and infrastructure engineering.
- Provide reliable and efficient operational support for production infrastructure and application environments.
- Monitor infrastructure, services and production environments to ensure availability, reliability and stability.
- Troubleshoot infrastructure and application issues and participate in incident resolution and root-cause analysis.
- Support and administer Kafka architecture and Kafka infrastructure.
- Work with Kafka brokers, topics, partitions, replication and configurations.
- Perform Kafka administration, monitoring and troubleshooting in production environments.
- Work with Kubernetes pods and services as part of infrastructure operations.
- Support large-scale server environments and ensure smooth day-to-day infrastructure operations.
- Work with AWS infrastructure and cloud services.
- Collaborate with infrastructure and engineering teams to maintain reliable production systems.
- Contribute to infrastructure modernization initiatives and improvements in system reliability.
Primary Skills & Responsibilities
- Strong experience in Site Reliability Engineering / SRE.
- Strong experience in operational support for production infrastructure.
- Hands-on experience with Kafka architecture and Kafka administration.
- Strong understanding of Kafka components, configuration, monitoring and troubleshooting.
- Strong experience supporting AWS infrastructure and services.
- Hands-on experience with Kubernetes pods and services.
- Experience handling production incidents, troubleshooting and root-cause analysis.
- Experience with Operating Systems and infrastructure environments.
- Experience with monitoring and observability tools.
- Ability to support large-scale infrastructure environments.
- Strong problem-solving and troubleshooting skills.
- Valuable understanding of infrastructure reliability, availability and performance.
- Ability to work effectively with infrastructure, engineering and application teams.