SRE Site Reliability Engineer | Cloud | Kubernetes | Observability | Automation
Company: Spark Techwave InnovationZ Pvt Ltd
Job Type: Full-Time
Experience: 27 Years
Location: Chennai / Remote / Hybrid
Department: Cloud Infrastructure / Site Reliability Engineering
Job Summary
Spark Techwave InnovationZ Pvt Ltd is looking for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, availability, and performance of production systems and applications.
The ideal candidate should have hands-on experience with cloud infrastructure, Linux, Kubernetes, CI/CD, monitoring, observability, automation, and incident management. The candidate will work closely with development, DevOps, cloud, and security teams to build highly reliable and scalable systems.
Key Responsibilities
- Design, implement, and maintain highly available and scalable production infrastructure.
- Monitor system availability, performance, capacity, and reliability.
- Define and track SLIs, SLOs, and SLAs for critical services.
- Develop automation to reduce manual operational work and improve system reliability.
- Manage and troubleshoot production environments across cloud and on-premises infrastructure.
- Support containerized workloads using Docker and Kubernetes.
- Build and maintain CI/CD pipelines for reliable application deployments.
- Implement monitoring, logging, alerting, and observability solutions.
- Investigate production incidents and perform root-cause analysis (RCA).
- Participate in incident response, troubleshooting, escalation, and service recovery.
- Develop and maintain runbooks, operational procedures, and incident documentation.
- Identify performance bottlenecks and implement system optimization.
- Perform capacity planning and infrastructure scaling.
- Implement disaster recovery, backup, high-availability, and fault-tolerance strategies.
- Conduct post-incident reviews and implement preventive actions.
- Improve deployment reliability through automation and Infrastructure as Code.
- Collaborate with software developers to improve application reliability and production readiness.
- Support security, compliance, and infrastructure best practices.
Required Skills
- Robust experience with Linux administration and troubleshooting.
- Hands-on experience with AWS, Azure, or GCP.
- Strong knowledge of Docker and Kubernetes.
- Experience with CI/CD pipelines.
- Experience with Infrastructure as Code such as Terraform.
- Strong understanding of monitoring, logging, alerting, and observability.
- Experience with tools such as Prometheus, Grafana, ELK, or equivalent.
- Strong scripting/programming skills using Python, Bash, or similar.
- Understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, and load balancing.
- Experience troubleshooting production systems and distributed applications.
- Knowledge of high availability, scalability, fault tolerance, and disaster recovery.
- Strong incident management and problem-solving skills.
- Experience with Git and version-control systems.
Good to Have
- AWS EKS / Azure AKS / Google GKE.
- Helm.
- Argo CD / Flux.
- Ansible.
- Jenkins / GitHub Actions / GitLab CI.
- OpenTelemetry.
- Datadog / Recent Relic / Dynatrace.
- Grafana Loki / Promtail.
- Kafka.
- Redis.
- Service mesh technologies such as Istio.
- CloudFormation / ARM / Bicep.
- Chaos engineering.
- Experience with microservices architecture.
- Knowledge of SRE practices and Google SRE principles.
- Cloud certifications.
Education
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
Preferred Candidate Profile
- 2–7 years of experience in SRE, DevOps, Cloud Infrastructure, Platform Engineering, or Production Engineering.
- Strong hands-on experience managing production systems.
- Good understanding of reliability and observability principles.
- Strong troubleshooting and incident-management capabilities.
- Ability to automate repetitive operational tasks.
- Comfortable working with development, cloud, infrastructure, and security teams.
- Strong communication, documentation, and problem-solving skills.
Key Skills for Naukri
SRE, Site Reliability Engineer, DevOps, Cloud, AWS, Azure, GCP, Linux, Kubernetes, Docker, Terraform, CI/CD, Prometheus, Grafana, Observability, Monitoring, Logging, Alerting, Incident Management, Production Support, Root Cause Analysis, RCA, SLI, SLO, SLA, Python, Bash, Infrastructure as Code, Microservices, High Availability, Disaster Recovery
Suggested Naukri Job Titles
- Site Reliability Engineer
- SRE Engineer
- Site Reliability Engineer – Cloud
- DevOps/SRE Engineer
- Cloud SRE
- Production Engineer
- Reliability Engineer
- Platform Reliability Engineer
- Infrastructure Reliability Engineer
- DevOps Engineer – SRE
Send your resume at
[email protected]
📌 Site Reliability Engineer (India)
🏢 Spark Tech Wave Innovation
📍 India