- Design, implement, and maintain scalable cloud infrastructure primarily in AWS
- Build and manage CI/CD pipelines for application deployment automation
- Maintain Kubernetes/EKS clusters and containerized workloads
- Manage Infrastructure as Code (IaC) using Terraform and related tooling
- Improve platform reliability, availability, disaster recovery, and observability
- Automate operational processes and reduce manual intervention
- Implement security best practices across cloud and deployment pipelines
- Collaborate with Security teams on vulnerability remediation and compliance initiatives
- Support SSO, IAM, RBAC, secrets management, and Zero Trust architecture
- Monitor production systems and respond to incidents/outages
- Optimize cloud costs, performance, and resource utilization
- Lead root cause analysis and post-incident remediation efforts
- Support multi-setting deployments (Dev, QA, Staging, Production)
- Maintain logging, monitoring, and alerting systems
- Mentor junior engineers and establish DevOps standards/best practices
Qualifications
- 5+ years of DevOps/SRE/Cloud Engineering experience
Required Skills Required Skills & Experience
- Cloud & Infrastructure
- 5+ years of DevOps/SRE/Cloud Engineering experience
- Strong experience with AWS services:
- EC2
- EKS
- ECS
- VPC
- Route53
- CloudFront
- S3
- IAM
- RDS
- Lambda
- CloudWatch
- WAF
- AWS Organizations
- Experience with hybrid and multi-account AWS environments
- Containers & Orchestration
- Strong Kubernetes administration experience
- Docker containerization expertise
- Helm charts and deployment automation
- Experience troubleshooting distributed systems
- CI/CD & Automation
- GitHub Actions / GitLab CI / Jenkins
- Terraform and Infrastructure as Code
- Ansible or configuration management tooling
- Bash/Python scripting automation
- Monitoring & Security
- Experience with:
- Datadog
- ELK/OpenSearch
- Prometheus/Grafana
- CloudWatch