29 Sep
|
Everforth Apex Systems
|
Bengaluru
29 Sep
Everforth Apex Systems
Bengaluru
Key Responsibilities
- Design, implement, and manage highly available and scalable infrastructure on AWS.
- Build and maintain CI/CD pipelines for application deployment and infrastructure changes.
- Implement Infrastructure as Code using Terraform or CloudFormation.
- Manage containerized workloads using Docker and Kubernetes/EKS.
- Automate infrastructure provisioning, configuration, deployments, and operational tasks.
- Monitor production environments and ensure availability, reliability, and performance.
- Implement observability using tools such as Prometheus, Grafana, CloudWatch, ELK/OpenSearch, or similar platforms.
- Define and monitor SLIs, SLOs, and SLAs.
- Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews.
- Improve system reliability through automation, capacity planning, performance optimization, and proactive monitoring.
- Implement backup, disaster recovery, high availability, and business continuity strategies.
- Manage AWS services including EC2, EKS, ECS, Lambda, S3, RDS, VPC, IAM, CloudWatch, Route 53, ALB/NLB, and related services.
- Implement and maintain security best practices including IAM, secrets management, network security, encryption, and least-privilege access.
- Work closely with development, QA,
security, and infrastructure teams to improve deployment and operational processes.
- Develop scripts and automation using Python, Bash, or similar languages.
- Maintain technical documentation, runbooks, architecture diagrams, and operational procedures.
Required Skills
- Strong hands-on experience with AWS cloud infrastructure.
- Positive understanding of Linux administration and troubleshooting.
- Strong knowledge of CI/CD concepts and tools such as Jenkins, GitHub Actions, GitLab CI, or AWS CodePipeline.
- Experience with Terraform or other Infrastructure-as-Code tools.
- Experience with Docker and Kubernetes, preferably AWS EKS.
- Strong understanding of networking concepts including VPC, subnets, routing, security groups, NACLs, DNS, load balancers, and TLS.
- Experience with Git and branching strategies.
- Experience with monitoring, logging, alerting, and observability.
- Strong scripting/automation skills in Python and/or Bash.
- Experience troubleshooting production issues and participating in on-call/incident management.
- Understanding of high availability, scalability, fault tolerance, and disaster recovery.
📌 SRE (Bengaluru)
🏢 Everforth Apex Systems
📍 Bengaluru