18 Sep
|
Encardio Rite
|
India
18 Sep
Encardio Rite
India
About the Role
We are looking for a Senior DevOps Engineer with strong hands-on experience in building, operating, and continuously improving production infrastructure on AWS . The ideal candidate will have genuine ownership of production environments and should be comfortable working across cloud infrastructure, Kubernetes, infrastructure-as-code, CI/CD, observability, security, and automation.
The role will involve designing and maintaining reliable, scalable, and secure infrastructure, while enabling effective and consistent software delivery through Terraform, Kubernetes, Helm, GitOps, and CI/CD automation .
The candidate should be a strong problem solver who can independently troubleshoot production issues, understand infrastructure dependencies, and automate repetitive operational activities using Python or Bash .
Key Responsibilities
- Own and operate production infrastructure on AWS, ensuring reliability, availability, scalability, and performance.
- Design, provision, and manage cloud infrastructure using Terraform, including reusable modules, state management, imports, and infrastructure changes.
- Review Terraform plans carefully and assess the potential impact of infrastructure changes before implementation.
- Manage and operate production Kubernetes environments, including cluster upgrades, ingress, RBAC, resource limits, workload troubleshooting, and operational maintenance.
- Develop, maintain, and troubleshoot Helm charts for application deployment and configuration management.
- Implement and maintain GitOps / declarative delivery practices using tools such as ArgoCD, Flux, or equivalent platforms.
- Design, maintain, and improve CI/CD pipelines, preferably using GitHub Actions, to enable reliable and repeatable software delivery.
- Troubleshoot Linux systems, networking issues, application deployments, containers, and infrastructure-related production incidents.
- Work closely with engineering and application teams to understand infrastructure requirements and ensure smooth deployment and operation of services.
- Develop automation scripts using Python and/or Bash to eliminate repetitive operational tasks and improve engineering efficiency.
- Monitor system health, performance, availability, and reliability and proactively identify potential production issues.
- Contribute to infrastructure security, access management, secrets management, and operational hardening.
- Participate in incident investigation, root-cause analysis, and implementation of preventive measures.
- Continuously evaluate infrastructure cost, capacity, scalability, and operational efficiency.
Key Deliverables
- Stable, scalable, and highly reliable AWS production infrastructure.
- Infrastructure managed through Terraform and declarative configuration practices.
- Well-maintained and production-ready Kubernetes clusters and workloads.
- Standardized and reusable Helm charts for application deployments.
- Reliable GitOps-based deployment workflows with appropriate controls and traceability.
- Efficient and maintainable CI/CD pipelines supporting frequent and reliable releases.
- Effective monitoring, alerting, troubleshooting,
and incident-resolution practices.
- Automation of repetitive operational activities through Python/Bash scripting.
- Improved infrastructure security, access control, and operational resilience.
- Continuous optimization of cloud cost, capacity, performance, and reliability.
Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical discipline.
- 3+ years of hands-on experience managing production infrastructure, preferably on AWS.
Technical Skills: Must-Have Cloud & Infrastructure
- AWS – production infrastructure management
- AWS networking and core infrastructure concepts
- Linux administration
- Networking fundamentals
Infrastructure as Code
- Terraform
- Terraform modules
- Terraform state management, imports, and infrastructure changes
- Ability to interpret and validate Terraform plans before applying changes
- Terragrunt – added advantage
Kubernetes & Containers
- Production Kubernetes operations
- Kubernetes upgrades
- Ingress
- RBAC
- Resource requests and limits
- Workload troubleshooting and debugging
- Containerized application environments
Helm
- Helm chart development
- Helm chart maintenance and customization
- Application configuration and deployment through Helm
GitOps / Declarative Delivery
- ArgoCD / Flux / equivalent GitOps platform
- Declarative deployment and configuration management
CI/CD
- CI/CD pipeline design and implementation
- GitHub Actions preferred
- Automated build, test, deployment, and release workflows
Scripting & Automation
- Python and/or Bash
- Automation of repetitive operational tasks
- Troubleshooting and operational scripting
📌 DevOps Engineer (India)
🏢 Encardio Rite
📍 India