- Design, implement, and maintain reliable, scalable, and secure infrastructure solutions.
- Build and enhance automation frameworks using Python, Shell scripting, and YAML.
- Develop and maintain CI/CD pipelines using GitHub Actions, Jenkins, or GitLab.
- Automate deployment, monitoring, and operational processes to improve efficiency and reliability.
- Manage and optimize Kubernetes environments, particularly Amazon EKS.
- Implement Infrastructure as Code (IaC) using Terraform.
- Troubleshoot platform, infrastructure, and application-related issues.
- Collaborate with development and operations teams to integrate automation throughout the software development lifecycle.
- Monitor system performance, availability, and operational health.
- Work with incident management and ticketing tools such as ServiceNow and JIRA.
Required Skills
Must Have
- 4+ years of experience in Site Reliability Engineering, DevOps, or related roles.
- Proficiency in Shell Scripting or Python and YAML.
- Hands-on experience with CI/CD tools:
- GitHub Actions
- Jenkins
- GitLab
- Experience integrating automation into the software development lifecycle.
- Strong SQL database knowledge.
- Experience troubleshooting platform and infrastructure issues.
- Understanding of code quality and development best practices.
- Hands-on experience with:
- ServiceNow
- JIRA
Preferred Skills
- Experience with Cloud Platforms (AWS preferred).
- Knowledge of Kubernetes administration and container technologies.
- Familiarity with observability and monitoring tools.
- Solid problem-solving and analytical skills.
- Excellent communication and stakeholder management capabilities.