SRE / Cloud Engineer (Bengaluru)

SRE / Cloud Engineer (Bengaluru)

08 Sep
|
MetAntz
|
Bengaluru

08 Sep

MetAntz

Bengaluru

Cloud & Infrastructure Site Reliability Engineer (SRE)

Role Overview

We are looking for a Cloud & Infrastructure Site Reliability Engineer (SRE) to build, operate, automate, and continuously improve highly available cloud-native infrastructure powering SaaS platform. You will work across public cloud, Kubernetes, compute, storage, networking, observability, and infrastructure automation - with a strong focus on reliability, scalability, security, and reducing manual operations. The ideal candidate combines strong infrastructure fundamentals with a software engineering and automation mindset - someone who proactively identifies operational issues and engineers permanent solutions rather than relying on repetitive manual processes.

Key Responsibilities

- Cloud & Infrastructure Operations: Design, deploy, operate, and troubleshoot cloud-native infrastructure across AWS, Azure, and GCP - including compute, storage, networking, and containers. Handle provisioning, upgrades, patching, and capacity management; resolve connectivity, DNS, and load-balancing issues.
- Site Reliability Engineering: Define and improve SLIs, SLOs, monitor for reliability risks, and drive incident response, root-cause analysis, and post-incident reviews. Automate away recurring operational work and improve system resilience through redundancy and automated recovery.
- Kubernetes & Microservices: Deploy, manage, and troubleshoot containerized microservices on Kubernetes (EKS, AKS, GKE), including ingress, storage, secrets, and cluster health.
- Infrastructure as Code & Automation: Automate provisioning and operational workflows with Terraform, Ansible, Python, and Bash; build reusable IaC modules integrated into CI, CD.
- Monitoring & Observability: Implement monitoring, logging, tracing, and alerting (Prometheus, Grafana, OpenTelemetry, Datadog, CloudWatch, etc.) with meaningful dashboards and low-noise alerting.
- Networking & Security: Troubleshoot cloud networking (VPC, VNet,



DNS, VPN, firewalls, load balancers); apply IAM, RBAC, secrets, and least-privilege security practices.
- Incident & Production Support: Participate in on-call rotations, own incidents end-to-end, and maintain runbooks and automated remediation procedures.
- Cost & Performance Optimization: Identify underutilized resources and support FinOps initiatives (rightsizing, scheduling, lifecycle management) while balancing cost with reliability.

Required Skills

- Strong hands-on experience with at least one major public cloud - AWS, Azure, or GCP.
- Solid understanding of Linux administration and infrastructure fundamentals.
- Experience with Kubernetes and containerized microservices.
- Experience with Infrastructure as Code (Terraform and, or Ansible).
- Strong automation, scripting skills in Python and Bash (PowerShell a plus).
- Positive understanding of networking, DNS, load balancing, firewalls, and cloud routing.
- Experience with monitoring, logging, alerting, and observability platforms.
- Understanding of CI, CD, Git, APIs, and modern DevOps practices.
- Experience troubleshooting distributed production environments.
- Understanding of security, IAM, RBAC, secrets, certificates, and infrastructure security.
- Familiarity with database operations a plus.

Preferred Skills Experience & Qualifications

- Experience operating hybrid-cloud or multi-cloud environments.
- Experience with VMware, Nutanix, OpenShift, or similar private-cloud technologies.
- Experience building automated remediation or self-healing workflows.




- Familiarity with AIOps, event correlation, anomaly detection, and AI-assisted operations.
- Understanding of FinOps and cloud cost optimization.
- Experience with ITSM platforms such as ServiceNow.
- Cloud, Kubernetes, Linux, or Terraform certifications (AWS, Azure, GCP) are a solid plus. SRE Mindset
- Automate before repeating, and engineer permanent fixes instead of resolving the same incident twice.
- Treat infrastructure and operational configuration as code.
- Design for failure and recovery, and measure reliability using data.
- Build reusable automation and operational standards.
- Reduce operational toil and unnecessary tickets.
- Continuously improve reliability, performance, security, and cost efficiency.
- Bachelor's or Master's degree in Computer Science or a related field, with minimum of 70%.
- SRE Engineer - 2–6 years of relevant Cloud, Infrastructure, DevOps, SRE experience.
- Cloud certifications (AWS, Azure, or GCP) are a strong plus. Roles and responsibilities
- Design, deploy, operate, and troubleshoot cloud-native infrastructure across AWS, Azure, and GCP. Handle provisioning, upgrades, patching, and capacity management.
- Define and improve SLIs, SLOs, monitor for reliability risks, and drive incident response.
- Automate away recurring operational work.
- Deploy, manage, and troubleshoot containerized microservices on Kubernetes.
- Automate provisioning and operational workflows with Terraform, Ansible, Python, and Bash.
- Implement monitoring, logging, tracing, and alerting.
- Troubleshoot cloud networking and apply security practices.
- Participate in on-call rotations and own incidents end-to-end.
- Identify underutilized resources and support FinOps initiatives. Experience and education
- Bachelor's or Master's degree in Computer Science or a related field, with minimum of 70% Location Remote

📌 SRE / Cloud Engineer (Bengaluru)
🏢 MetAntz
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sre / cloud engineer (bengaluru) / bengaluru