18 Sep
|
Agile Dna
|
Kochi
IN Infrastructure Specialist - AWS DevOps
Location: Kochi
Number of Positions: 1
Total Experience: 7+ Years
Relevant Experience: 6+ Years
Role Overview
We are looking for an experienced AWS Cloud & DevOps Engineer with robust expertise in AWS cloud services, Terraform, infrastructure automation, Python scripting, CI/CD, cloud operations, and production support.
The ideal candidate will be responsible for designing, implementing, automating, and supporting scalable and reliable cloud infrastructure and application environments. The role requires strong hands-on experience in AWS, Infrastructure as Code, DevOps practices, troubleshooting, incident management, monitoring, and SRE principles.
The candidate should have a strong automation mindset and be comfortable working in production environments, resolving complex infrastructure and application issues, performing root cause analysis, and collaborating with development, operations, security, and other technical teams.
Key Responsibilities
AWS Cloud Infrastructure
- Design, provision, configure, and manage AWS cloud infrastructure.
- Work hands-on with EC2, VPC, IAM, S3, RDS, CloudWatch, ECS, and EKS.
- Ensure AWS environments are secure, scalable, highly available, and reliable.
- Implement appropriate IAM roles, policies, permissions, and security controls.
- Support cloud infrastructure upgrades, configuration changes, and optimization activities.
Infrastructure as Code & Automation
- Develop and maintain infrastructure using Terraform and Infrastructure as Code (IaC) practices.
- Build reusable and modular Terraform configurations for multiple environments.
- Manage Terraform state, remote backends, variables, modules, and environment-specific configurations.
- Automate infrastructure provisioning, configuration, deployment, and operational activities.
- Identify opportunities to eliminate manual processes through automation.
Python & Scripting
- Develop Python scripts for infrastructure automation, operational tooling, monitoring, reporting, and cloud management.
- Create Shell/Linux scripts to automate routine operational activities.
- Develop tools and utilities to improve operational efficiency and reduce manual intervention.
CI/CD & DevOps
- Design, implement, maintain, and optimize CI/CD pipelines.
- Work with CI/CD technologies such as Jenkins, GitHub Actions, GitLab CI, and Azure DevOps.
- Automate application build, testing, deployment, and release processes.
- Integrate infrastructure and application deployment processes into CI/CD workflows.
- Implement DevOps best practices across development, test, staging, and production environments.
Containers & Orchestration
- Build, manage, and deploy Dockerized applications.
- Manage containerized workloads using Kubernetes and Amazon EKS/ECS.
- Troubleshoot container, pod, node, cluster, and deployment-related issues.
- Support scalable and highly available container platforms.
Cloud Operations & Production Support
- Provide hands-on support for AWS cloud and production environments.
- Monitor infrastructure and applications to ensure availability, performance, and reliability.
- Perform advanced troubleshooting of cloud, infrastructure, networking, operating system, and application issues.
- Participate in production support, incident response, change management, and service restoration activities.
- Ensure operational activities meet defined SLAs and service expectations.
Incident Management & RCA
- Investigate and resolve production incidents and recurring technical issues.
- Perform detailed Root Cause Analysis (RCA) for major incidents.
- Identify corrective and preventive actions to avoid recurrence.
- Collaborate with application, database, network, security, and infrastructure teams during incident resolution.
- Maintain incident documentation and operational knowledge bases.
Monitoring & Observability
- Implement and maintain monitoring, logging, alerting, and observability solutions.
- Work with tools such as Amazon CloudWatch, Prometheus, Grafana, Datadog, and Splunk.
- Monitor system health, application performance, resource utilization, availability, and reliability.
- Define appropriate alerts and dashboards for production environments.
- Analyze logs and metrics to proactively identify potential issues.
Linux/Unix & Networking
- Perform Linux/Unix administration and troubleshooting.
- Troubleshoot CPU, memory, disk, process, service, and connectivity issues.
- Apply strong networking fundamentals including DNS, TCP/IP, load balancing, subnets, routing, and security groups.
- Troubleshoot network connectivity and communication issues across AWS environments.
SRE & Reliability
- Apply Site Reliability Engineering (SRE) principles to cloud and application environments.
- Support the definition and monitoring of SLAs, SLOs, and reliability objectives.
- Identify and eliminate reliability risks and operational bottlenecks.
- Contribute to capacity planning, availability, scalability, and operational resilience.
- Promote automation and continuous improvement across operational processes.
Security & Best Practices
- Implement AWS cloud security and operational best practices.
- Follow least-privilege access principles and secure infrastructure configurations.
- Support security, compliance, vulnerability management, and operational governance activities.
- Ensure infrastructure and deployment processes follow organizational security standards.
Mandatory Skills
- AWS Cloud Services: EC2, VPC, IAM, S3, RDS, CloudWatch, ECS/EKS
- Terraform & Infrastructure as Code (IaC)
- Python Scripting & Automation
- Cloud Operations, Production Support & Advanced Troubleshooting
- DevOps & CI/CD: Jenkins, GitHub Actions, GitLab CI, or Azure DevOps
- Containerization & Orchestration: Docker, Kubernetes/EKS
- Monitoring & Observability: CloudWatch, Prometheus, Grafana, Datadog, or Splunk
- Linux/Unix Administration
- Networking: DNS, TCP/IP, Load Balancing, Security Groups
- Incident Management & Root Cause Analysis
- SRE principles, SLAs/SLOs, and reliability practices
- Cloud Security and operational best practices
Good to Have
- AWS certifications such as AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps Administrator.
- Experience with advanced Kubernetes/EKS administration.
- Experience with AWS ECS and containerized production workloads.
- Experience implementing automated remediation and self-healing solutions.
- Experience with centralized logging and distributed observability.
- Experience with Datadog, Splunk, Prometheus, Grafana, or similar observability platforms.
- Experience with cloud cost optimization.
- Experience working in Agile/DevOps environments.
📌 IN Infrastructure Specialist - AWS DevOps (Kochi)
🏢 Agile Dna
📍 Kochi