13 Sep
|
Embarkgcc Services
|
Bengaluru
13 Sep
Embarkgcc Services
Bengaluru
About the Role
We are looking for a Site Reliability Engineer / Platform SRE who enjoys working with cloud infrastructure, Kubernetes and production environments.
The ideal candidate should be comfortable troubleshooting real-world production issues, improving platform reliability, monitoring applications and working closely with Platform and DevOps teams.
This is not a coding-heavy role. We are looking for someone who can understand existing code and scripts, troubleshoot issues and make basic changes when required. AI-assisted tools can be used for more complex coding requirements.
Key Responsibilities
Manage and support Azure-based cloud infrastructure and services.
Work with Kubernetes for deployments, troubleshooting, scaling and production support.
Monitor production settings and investigate application/infrastructure issues.
Analyze logs, metrics and alerts to identify issues and perform root-cause analysis (RCA).
Work closely with Platform Engineering and DevOps teams to maintain reliable production environments.
Support and improve CI/CD pipelines and deployment processes.
Work with Apache Flink and data-processing environments where required.
Support ETL/ELT pipelines and data-platform workflows.
Contribute to monitoring and observability using tools such as Datadog, Dynatrace or similar platforms.
Participate in incident management and help improve system reliability.
Support vulnerability identification, remediation and other basic security-related activities.
Assist with API performance/load testing and identifying performance bottlenecks.
Use automation and AI-assisted tools to improve operational efficiency.
Required Skills :
Hands-on experience with Microsoft Azure.
Strong practical experience with Kubernetes.
Positive understanding of SRE / Platform Engineering concepts.
Robust production troubleshooting and RCA experience.
Experience with CI/CD and deployment processes.
Good understanding of monitoring, logging and observability.
Exposure to Datadog, Dynatrace or similar observability tools.
Experience with Apache Flink is preferred.
Understanding of ETL/ELT and data pipelines.
Experience working with Platform/DevOps team
📌 Senior Site Reliability Engineer Bengaluru
🏢 Embarkgcc Services
📍 Bengaluru