28 Aug
|
Nam Info
|
Bengaluru
28 Aug
Nam Info
Bengaluru
Associate Architect – Site Reliability Engineer (SRE)
Experience: 10–14 Years
Location: Bangalore
Required Skills
8–14+ years of experience in Site Reliability Engineering (SRE), DevOps, or Cloud Engineering.
Strong hands-on experience with AWS (EKS, EC2, S3, VPC) and Kubernetes.
Expertise in Docker, Terraform, Ansible, and Linux administration.
Experience building and maintaining CI/CD pipelines using Jenkins or GitHub Actions.
Solid programming skills in Python, Java, or TypeScript.
Experience with Monitoring & Observability tools like Prometheus, Grafana, Datadog, New Relic, Honeycomb, or ELK.
Hands-on experience defining and managing SLIs, SLOs, SLAs, proactive monitoring, and incident management.
Good knowledge of SQL/NoSQL, OpenSearch, Elasticsearch, and AWS S3.
Experience troubleshooting production issues, RCA, and automation of operational tasks.
Robust communication skills and ability to work closely with development teams.
Key Responsibilities
Design, build, and maintain highly available, scalable cloud infrastructure.
Manage and optimize Kubernetes (EKS) clusters and containerized applications.
Develop Infrastructure as Code (IaC) using Terraform and Ansible.
Automate deployments, monitoring,
and operational tasks using Python and CI/CD tools.
Implement SLIs, SLOs, monitoring dashboards, alerting, and observability solutions.
Troubleshoot production incidents, perform Root Cause Analysis (RCA), and reduce MTTR.
Support developers with infrastructure architecture, scalability, and reliability improvements.
Participate in on-call support and production issue resolution.
AI / GenAI (Preferred)
Experience implementing AIOps solutions for intelligent monitoring and incident management.
Hands-on experience using GenAI/LLMs (OpenAI, Azure OpenAI, Amazon Q, Vertex AI, etc.) for:
Automated log analysis
Incident troubleshooting
Root Cause Analysis (RCA)
Runbook automation
Predictive monitoring and alerting
Experience integrating AI solutions into DevOps/SRE workflows is an added advantage.
Good to Have
Istio / Service Mesh
GitHub Actions
Datadog / Honeycomb / Recent Relic
Elasticsearch / OpenSearch
Experience in cloud architecture and platform engineering
Exposure to AI-based infrastructure automation and observability
Skills:- DevOps, Site reliability, SRE, Docker, Ansible, Linux/Unix, Python, Monitoring, Observability, prometheus and grafana
📌 Associate Architect  Site Reliability Engineer Sre Bengaluru
🏢 Nam Info
📍 Bengaluru