03 Sep
|
Solutions By Text
|
Bengaluru
03 Sep
Solutions By Text
Bengaluru
Position Summary
We are seeking a proactive Site Reliability Engineer (SRE) with 2 to 5 years of hands-on experience to maintain high availability, scalability, and performance across our AWS cloud ecosystem.
In this role, you will bridge software engineering and operations by building automated delivery pipelines, implementing proactive monitoring, managing containerized applications, and driving incident response protocols.
Key Responsibilities
- Reliability & Incident Management: Monitor production systems, lead operational triage, analyze root causes (RCA), and implement preventative measures to ensure system availability and performance.
- Cloud Infrastructure & Containerization: Manage and support AWS core compute and networking services (EC2, ECS Fargate, S3, VPC, Route 53). Maintain containerized workflows using Docker.
- Automation & IaC: Write and maintain Infrastructure as Code using AWS CloudFormation or Terraform to automate infrastructure provisioning and reduce operational toil.
- CI/CD Pipeline Support: Maintain and optimize automated build and deployment pipelines using Jenkins and shell/OS scripting.
- Observability & Alerting: Configure metrics, log aggregation, and alerting frameworks using Datadog or AWS CloudWatch to maintain proactive platform health monitoring.
- Cloud Security & Quality Governance: Maintain cloud perimeter security using AWS WAF and AWS Guard Duty, while assisting with code quality and license compliance tools.
QUALIFICATIONS & REQUIREMENTS
Must-Have Qualifications & Skills: • Experience: 2 to 5 years of hands-on experience in SRE, DevOps,
or Cloud Infrastructure Operations.
- Scripting Knowledge: Strong scripting proficiency (Bash, PowerShell, or Python) for task automation, system support, and reducing operational toil.
- Containerization: Hands-on expertise with Docker for containerizing applications.
- Infrastructure as Code (IaC): Working experience with AWS CloudFormation or Terraform.
- AWS Core Services & Networking: Practical experience with VPC, Route 53, EC2, ECS Fargate, S3, AWS WAF, and AWS GuardDuty.
- CI/CD & Automation: Experience configuring and maintaining deployment pipelines in Jenkins.
- Observability: Practical experience establishing metrics, logs, and alerts in Datadog or AWS CloudWatch.
Good-To-Have Skills: • AI in DevOps / Automation: Familiarity with integrating AI tools, protocols, or agents into DevOps workflows and operational automation. • Identity & Access Management: Experience with Active Directory management. • Software Development
Background: Minimum 1 year of hands-on software development experience in any programming language (C# or Python preferred).
- Code Quality & Security Tools: Familiarity or administration experience with SonarQube and Mend (formerly WhiteSource).
- Modern CI/CD: Exposure to GitHub Actions.
- Enterprise Governance: Basic understanding of AWS Control Tower and AWS Security Hub. CORE COMPETENCIES • Strong analytical and troubleshooting mindset for incident triage.
- Drive to automate repetitive manual tasks and eliminate operational toil.
- Transparent communication skills to collaborate with development teams and operational stakeholders.
📌 Site Reliability Engineer (Bengaluru)
🏢 Solutions By Text
📍 Bengaluru