27 Aug
|
Nam info
|
Bengaluru
27 Aug
Nam info
Bengaluru
Associate Architect – Site Reliability Engineer (SRE)
Experience: 10–14 Years
Location: Bangalore
Required Skills
- 8–14+ years of experience in Site Reliability Engineering (SRE), DevOps, or Cloud Engineering.
- Strong hands-on experience with AWS (EKS, EC2, S3, VPC) and Kubernetes.
- Expertise in Docker, Terraform, Ansible, and Linux administration.
- Experience building and maintaining CI/CD pipelines using Jenkins or GitHub Actions.
- Solid programming skills in Python, Java, or TypeScript.
- Experience with Monitoring & Observability tools like Prometheus, Grafana, Datadog, New Relic, Honeycomb, or ELK.
- Hands-on experience defining and managing SLIs, SLOs, SLAs, proactive monitoring, and incident management.
- Good knowledge of SQL/NoSQL, OpenSearch, Elasticsearch, and AWS S3.
- Experience troubleshooting production issues, RCA, and automation of operational tasks.
- Strong communication skills and ability to work closely with development teams.
Key Responsibilities
- Design, build, and maintain highly available, scalable cloud infrastructure.
- Manage and optimize Kubernetes (EKS) clusters and containerized applications.
- Develop Infrastructure as Code (IaC) using Terraform and Ansible.
- Automate deployments, monitoring,
and operational tasks using Python and CI/CD tools.
- Implement SLIs, SLOs, monitoring dashboards, alerting, and observability solutions.
- Troubleshoot production incidents, perform Root Cause Analysis (RCA), and reduce MTTR.
- Support developers with infrastructure architecture, scalability, and reliability improvements.
- Participate in on-call support and production issue resolution.
AI / GenAI (Preferred)
- Experience implementing AIOps solutions for intelligent monitoring and incident management.
- Hands-on experience using GenAI/LLMs (OpenAI, Azure OpenAI, Amazon Q, Vertex AI, etc.) for:
- Automated log analysis
- Incident troubleshooting
- Root Cause Analysis (RCA)
- Runbook automation
- Predictive monitoring and alerting
- Experience integrating AI solutions into DevOps/SRE workflows is an added advantage.
Good to Have
- Istio / Service Mesh
- GitHub Actions
- Datadog / Honeycomb / New Relic
- Elasticsearch / OpenSearch
- Experience in cloud architecture and platform engineering
- Exposure to AI-based infrastructure automation and observability
Skills:- DevOps, Site reliability, SRE, Docker, Ansible, Linux/Unix, Python, Monitoring, Observability, prometheus and grafana
📌 Associate Architect â Site Reliability Engineer (SRE) (Bengaluru)
🏢 Nam info
📍 Bengaluru