30 Sep
|
HyperVerge
|
Bengaluru
30 Sep
HyperVerge
Bengaluru
Role: Senior DevOps / Infrastructure & Reliability Engineer
Experience Level: 3–5 Years
Location: Bangalore
Employment Type: Full-time
About the Role
We are seeking a proactive and skilled Senior DevOps Engineer with 3 to 5 years of hands-on experience in building, scaling, and maintaining high-availability cloud infrastructure. In this role, you will play a key part in driving our infrastructure automation, improving platform reliability, and enhancing system observability.
You will work closely with software engineering teams to streamline CI/CD pipelines, optimize containerized workloads on Kubernetes, and ensure our AWS environment remains secure, resilient, and cost-effective.
Key Responsibilities
1. Infrastructure as Code (IaC) & Cloud Engineering
- Design, build, and maintain scalable, secure, and multi-region AWS cloud infrastructure using Terraform .
- Enforce IaC best practices, including modular design, state management, dry principles, and automated infrastructure testing.
- Optimize AWS resources (EKS, EC2, S3, RDS, IAM, VPC, CloudFront) for performance, security, and cost efficiency (FinOps).
1. Containerization & Kubernetes Platform Engineering
- Manage and maintain operational excellence of production-grade AWS EKS / Kubernetes (k8s) clusters.
- Manage cluster add-ons, ingress controllers, service meshes (e.g., Istio, Linkerd), and storage integration.
- Optimize Kubernetes workloads for auto-scaling (HPA/Karpenter/Cluster Autoscaler), resource requests/limits, and node provisioning.
1. Continuous Integration & Delivery (CI/CD)
- Design, enhance, and manage robust CI/CD pipelines (GitLab CI, GitHub Actions, ArgoCD, or Jenkins) for continuous deployment.
- Implement GitOps workflows to manage Kubernetes cluster states and application deployments seamlessly.
1. Observability & System Reliability
- Architect, implement, and maintain centralized observability systems using modern telemetry stacks (Prometheus, Grafana, OpenTelemetry, Datadog, ELK/EFK, or AWS CloudWatch).
- Establish SLIs, SLOs, and Error Budgets alongside software engineering teams to quantify and improve platform reliability.
- Configure intelligent alerting and dashboarding to proactive detect, diagnose, and resolve system anomalies before they impact end users.
- Lead root-cause analysis (RCA) and post-mortem reviews for critical incidents to drive long-term reliability improvements.
Key Qualifications
- Experience: 3–5 years of dedicated experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure Engineering.
- AWS Expertise: In-depth knowledge of core AWS services (EKS, VPC, IAM, EC2, RDS, S3, Route53) and cloud security best practices.
- Terraform: Proven experience writing modular, reusable, and maintainable Terraform code for infrastructure provisioning.
- Kubernetes (k8s): Solid hands-on experience deploying, troubleshooting, and managing containerized applications on production-grade Kubernetes.
- Observability: Solid experience setting up and managing monitoring, logging, and tracing tools (e.g., Prometheus, Grafana, Datadog, Jaeger/OpenTelemetry).
- CI/CD & GitOps: Hands-on expertise building CI/CD automation pipelines and utilizing GitOps tools (e.g., ArgoCD, Flux).
- Scripting & Automation: Proficiency in Python , Bash , or Go for operational automation and task scripting.
- Linux Fundamentals: Strong Linux system administration, networking, and security concepts (DNS, TCP/IP, SSL/TLS, SSH).
📌 Senior Devops Engineer (Bengaluru)
🏢 HyperVerge
📍 Bengaluru