12 Sep
|
Embarkgcc Services
|
Bengaluru
12 Sep
Embarkgcc Services
Bengaluru
About the Role
We are looking for a Site Reliability Engineer / Platform SRE who enjoys working with cloud infrastructure, Kubernetes and production environments.
The ideal candidate should be comfortable troubleshooting real-world production issues, improving platform reliability, monitoring applications and working closely with Platform and DevOps teams.
This is not a coding-heavy role. We are looking for someone who can understand existing code and scripts, troubleshoot issues and make basic changes when required. AI-assisted tools can be used for more complex coding requirements.
Key Responsibilities
- Manage and support Azure-based cloud infrastructure and services.
- Work with Kubernetes for deployments, troubleshooting, scaling and production support.
- Monitor production environments and investigate application/infrastructure issues.
- Analyze logs, metrics and alerts to identify issues and perform root-cause analysis (RCA).
- Work closely with Platform Engineering and DevOps teams to maintain reliable production environments.
- Support and improve CI/CD pipelines and deployment processes.
- Work with Apache Flink and data-processing environments where required.
- Support ETL/ELT pipelines and data-platform workflows.
- Contribute to monitoring and observability using tools such as Datadog, Dynatrace or similar platforms.
- Participate in incident management and help improve system reliability.
- Support vulnerability identification, remediation and other basic security-related activities.
- Assist with API performance/load testing and identifying performance bottlenecks.
- Use automation and AI-assisted tools to improve operational efficiency.
Required Skills :
- Hands-on experience with Microsoft Azure.
- Strong practical experience with Kubernetes.
- Good understanding of SRE / Platform Engineering concepts.
- Robust production troubleshooting and RCA experience.
- Experience with CI/CD and deployment processes.
- Good understanding of monitoring, logging and observability.
- Exposure to Datadog, Dynatrace or similar observability tools.
- Experience with Apache Flink is preferred.
- Understanding of ETL/ELT and data pipelines.
- Experience working with Platform/DevOps team
📌 Senior Site Reliability Engineer (Bengaluru)
🏢 Embarkgcc Services
📍 Bengaluru