01 Sep
|
UST
|
Thiruvananthapuram
01 Sep
UST
Thiruvananthapuram
Role Description
Experience: 8-15 Years
Job Location: Chennai, Bangalore, Hyderabad, Kochi, Trivandrum, Noida, Pune
DevOps SRE Engineer / Platform EngineerRole Overview
We are looking for highly hands-on DevOps SRE Engineers with strong Platform Engineering experience supporting application and API hosting platforms in AWS and Kubernetes environments.
The ideal candidate should have strong recent hands-on technical expertise in production environments, with the ability to operate, troubleshoot, optimize, and enhance platforms within a multi-vendor ecosystem involving multiple teams.
This is a hands-on individual contributor role and not a people-management position. Strong communication skills are essential, along with the ability to lead technical discussions with Tier 2 teams, stakeholders, and cross-functional engineering groups.
Key Responsibilities
- Maintain, troubleshoot, and enhance CI/CD pipelines, with GitLab preferred.
- Manage and support AWS infrastructure and cloud-native services.
- Deploy, operate, and troubleshoot Kubernetes workloads and platforms.
- Perform production incident triage, debugging, and root cause analysis.
- Use logs, metrics, monitoring tools, and Splunk for effective production troubleshooting.
- Develop and enhance Terraform-based infrastructure and automation.
- Support repository restructuring, platform modernization, and engineering transformation initiatives.
- Drive improvements in platform stability, reliability, observability, scalability, and performance.
- Support platforms that host and enable application and API services.
- Collaborate closely with the GCC team and UST BFF leads.
- Work effectively across a multi-vendor engineering ecosystem and coordinate technical resolution across teams.
- Drive technical discussions, identify recurring platform issues,
and implement sustainable solutions.
- Contribute to continuous improvement of platform engineering practices, automation, and operational processes.
Must-Have Skills
- Robust hands-on experience in DevOps / Site Reliability Engineering (SRE).
- Strong AWS experience, including production infrastructure management and troubleshooting.
- Strong hands-on Kubernetes experience.
- Strong experience with CI/CD pipelines, preferably GitLab CI/CD.
- Hands-on experience with Terraform and Infrastructure as Code (IaC).
- Strong production troubleshooting and incident management experience.
- Experience with Splunk, logs, metrics, and monitoring/observability tools.
- Strong understanding of Linux/Unix environments and troubleshooting.
- Experience supporting application and API hosting platforms.
- Experience with platform reliability, availability, performance, and observability improvements.
- Strong debugging and root cause analysis (RCA) capabilities.
- Ability to work independently as a hands-on technical contributor.
- Strong communication and stakeholder-management skills.
- Ability to drive technical discussions with Tier 2 support, engineering teams, stakeholders, and cross-functional teams.
- Experience working in a multi-vendor / distributed engineering environment.
Good-to-Have Skills
- Experience with AWS EKS or other managed Kubernetes services.
- Experience with GitLab administration or advanced GitLab CI/CD.
- Advanced Terraform modules, state management, and automation experience.
- Experience with Helm and Kubernetes deployment automation.
- Experience with observability platforms such as Prometheus, Grafana, CloudWatch, or similar tools.
- Experience with Docker/containerization.
- Experience with scripting/automation using Python, Shell, or Bash.
- Knowledge of API platforms, microservices, and cloud-native architectures.
- Experience with platform modernization and repository restructuring.
- Experience implementing SRE practices, SLIs/SLOs, error budgets, and reliability engineering principles.
- Experience with security, IAM, networking, and cost optimization in AWS.
- Experience coordinating technical initiatives across multiple vendors and engineering teams.
Experience Range 6–12 years of overall experience in DevOps, SRE, Cloud Engineering, Platform Engineering, or related roles.
Preferred: 4+ years of strong hands-on experience with AWS and Kubernetes in production environments.
Candidates with extensive recent hands-on expertise and strong production troubleshooting experience will be preferred over candidates with primarily managerial or coordination experience.
Preferred Candidate Profile The ideal candidate is a strong hands-on engineer who can operate production platforms, troubleshoot complex incidents, automate infrastructure, improve reliability, and work across multiple engineering/vendor teams. The role requires someone who can not only execute technically but also confidently drive technical conversations and influence stakeholders toward effective platform solutions.
Skills
Site Reliability Engineering, Splunk, Terraform, Kubernetes
📌 Lead II - DevOps Engineering (Thiruvananthapuram)
🏢 UST
📍 Thiruvananthapuram