18 Aug
|
Smart IMS
|
Bengaluru
18 Aug
Smart IMS
Bengaluru
Roles & Responsibilities
As a Site Reliability Engineer, your day-to-day responsibilities will include:
- Infrastructure Orchestration: Architect, maintain, and scale production-grade Kubernetes (K8s) clusters; continuously monitor and troubleshoot control plane mechanics, cluster networking, and workload scheduling.
- Deployment Automation: Design, optimize, and maintain robust CI/CD pipelines to streamline application delivery.
- Progressive Delivery Management: Safely execute and manage advanced application rollout strategies, including Canary deployments and Blue/Green deployment patterns, minimizing user downtime.
- Cloud Operations: Manage, provision, and secure cloud resources within a major public cloud workplace (AWS, GCP, or Azure), utilizing a software-defined infrastructure approach.
- Eliminate Administrative Toil: Actively identify manual, repetitive operational tasks and eliminate them by writing clean, reusable automation scripts and internal tools using Python, Go, or Bash.
- Network Troubleshooting: Diagnose and resolve complex system interconnectivity, routing, load balancing, and communication bugs across core protocols (TCP/IP, DNS, HTTP).
- Cross-Functional Collaboration: Bridge the gap between software development and infrastructure operations to ensure system reliability, visibility, and optimal performance scaling.
Preferred Candidate Profile
The ideal candidate brings a blend of systems engineering expertise, an automation-first mindset, and production-level troubleshooting skills:
- Experience: 3 to 6 years of proven experience working as an SRE, DevOps Engineer, or Cloud Infrastructure Engineer managing production environments.
- Deep Kubernetes Expertise: Strong conceptual and operational mastery of Kubernetes architecture. You should know K8s troubleshooting paradigms "inside out."
- Pipeline Proficiency: Hands-on experience with modern CI/CD tools (e.g., Jenkins, GitLab CI, GitHub Actions, or ArgoCD) and a proven track record of fixing pipeline friction or lifecycle failures.
- Cloud & Security Competency: Solid working experience with at least one major cloud platform (AWS, GCP, or Azure), including a strong grasp of native service provisioning and identity/access management (IAM).
- Strong Coding/Scripting Skills: Proficiency in Python, Bash, or Go. You write clean, production-grade code to build automation tools rather than just running manual commands.
- Solid Networking Foundation: Deep understanding of core networking principles, routing, load balancers, and protocols (TCP/IP, DNS, HTTP) to quickly pinpoint why services aren't communicating.
- Education: Bachelors degree in Computer Science, Information Technology, Engineering, or a related technical field (or equivalent practical experience).
📌 Site Reliability Engineer (Bengaluru)
🏢 Smart IMS
📍 Bengaluru