21 Aug
|
Innova ESI
|
Bengaluru
21 Aug
Innova ESI
Bengaluru
Job Title: Senior SRE Engineer (AWS & DevOps)
Experience:
7+ Years – Sal2
Location:
Bangalore / Hyderabad
Job Summary
We are seeking a highly skilled and experienced Site Reliability Engineer (SRE) with robust AWS and DevOps expertise to join our engineering team. The ideal candidate will have extensive experience in building, automating, monitoring, and maintaining highly scalable, reliable, and secure cloud infrastructure on AWS. This role requires a strong background in infrastructure automation, CI/CD, observability, incident management, and cloud-native technologies.
The candidate will work closely with Development, Platform Engineering, Security, and Operations teams to ensure high availability, performance, and reliability of mission-critical applications.
Key Responsibilities
Site Reliability Engineering
- Ensure high availability, reliability, and performance of production systems.
- Define and maintain SLI, SLO, and SLA metrics across applications and infrastructure.
- Lead incident response, root cause analysis (RCA), and postmortem reviews.
- Implement proactive monitoring, alerting, and capacity planning strategies.
- Drive automation to eliminate repetitive operational tasks.
AWS Cloud Management
- Design, deploy, and manage cloud infrastructure on AWS.
- Hands-on experience with:
- EC2
- ECS/EKS
- Lambda
- RDS
- DynamoDB
- S3
- VPC
- Route53
- CloudFront
- IAM
- CloudWatch
- SNS/SQS
- Implement AWS security best practices and cost optimization initiatives.
DevOps & Automation
- Develop and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
- Implement Infrastructure as Code (IaC) using Terraform and CloudFormation.
- Automate provisioning, deployment, and configuration management processes.
- Build deployment strategies including Blue-Green, Canary, and Rolling Deployments.
Containerization & Kubernetes
- Deploy and manage containerized workloads using Docker.
- Administer Kubernetes clusters (EKS preferred).
- Troubleshoot Kubernetes networking, storage, and scaling issues.
- Experience working with Helm Charts and Kubernetes Operators.
Monitoring & Observability
- Implement monitoring solutions using:
- Prometheus
- Grafana
- Datadog
- ELK Stack
- Splunk
- CloudWatch
- Create dashboards, alerts, and performance reports.
- Drive observability initiatives across distributed systems.
Security & Compliance
- Collaborate with security teams to implement DevSecOps practices.
- Manage IAM policies, secrets management, and vulnerability remediation.
- Ensure compliance with industry standards and security frameworks.
Required Skills
Cloud Platforms
- Strong hands-on AWS experience (5+ years)
- Multi-account AWS environment management
- AWS Well-Architected Framework knowledge
DevOps Tools
- Terraform
- CloudFormation
- Jenkins
- GitHub Actions
- GitLab CI/CD
- ArgoCD
Containers & Orchestration
- Docker
- Kubernetes (EKS preferred)
- Helm
Monitoring Tools
- Prometheus
- Grafana
- Datadog
- ELK Stack
- CloudWatch
Scripting Languages
- Python
- Shell/Bash
- Go (Good to Have)
Version Control
- Git
- GitHub
- GitLab
- Bitbucket
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, or equivalent.
- 7+ years of experience in SRE/DevOps/Cloud Engineering.
- Strong experience in AWS cloud infrastructure.
- Experience supporting large-scale production environments.
- Deep understanding of Linux systems administration.
- Expertise in automation and Infrastructure as Code.
Preferred Certifications
- AWS Certified Solutions Architect – Professional
- AWS Certified DevOps Engineer – Professional
- Certified Kubernetes Administrator (CKA)
- HashiCorp Terraform Associate
📌 Senior SRE Engineer (AWS & DevOps) (Bengaluru)
🏢 Innova ESI
📍 Bengaluru