Site Reliability Engineer (India)

Site Reliability Engineer (India)

09 Oct
|
Spark Tech Wave Innovation
|
India

09 Oct

Spark Tech Wave Innovation

India

SRE Site Reliability Engineer | Cloud | Kubernetes | Observability | Automation

Company: Spark Techwave InnovationZ Pvt Ltd
Job Type: Full-Time
Experience: 27 Years
Location: Chennai / Remote / Hybrid
Department: Cloud Infrastructure / Site Reliability Engineering

Job Summary

Spark Techwave InnovationZ Pvt Ltd is looking for a Site Reliability Engineer (SRE) to ensure the reliability, scalability, availability, and performance of production systems and applications.

The ideal candidate should have hands-on experience with cloud infrastructure, Linux, Kubernetes, CI/CD, monitoring, observability, automation, and incident management. The candidate will work closely with development, DevOps, cloud, and security teams to build highly reliable and scalable systems.

Key Responsibilities

- Design, implement, and maintain highly available and scalable production infrastructure.
- Monitor system availability, performance, capacity, and reliability.
- Define and track SLIs, SLOs, and SLAs for critical services.
- Develop automation to reduce manual operational work and improve system reliability.
- Manage and troubleshoot production environments across cloud and on-premises infrastructure.
- Support containerized workloads using Docker and Kubernetes.
- Build and maintain CI/CD pipelines for reliable application deployments.
- Implement monitoring, logging, alerting, and observability solutions.
- Investigate production incidents and perform root-cause analysis (RCA).
- Participate in incident response, troubleshooting, escalation, and service recovery.
- Develop and maintain runbooks, operational procedures, and incident documentation.
- Identify performance bottlenecks and implement system optimization.
- Perform capacity planning and infrastructure scaling.




- Implement disaster recovery, backup, high-availability, and fault-tolerance strategies.
- Conduct post-incident reviews and implement preventive actions.
- Improve deployment reliability through automation and Infrastructure as Code.
- Collaborate with software developers to improve application reliability and production readiness.
- Support security, compliance, and infrastructure best practices.

Required Skills

- Robust experience with Linux administration and troubleshooting.
- Hands-on experience with AWS, Azure, or GCP.
- Strong knowledge of Docker and Kubernetes.
- Experience with CI/CD pipelines.
- Experience with Infrastructure as Code such as Terraform.
- Strong understanding of monitoring, logging, alerting, and observability.
- Experience with tools such as Prometheus, Grafana, ELK, or equivalent.
- Strong scripting/programming skills using Python, Bash, or similar.
- Understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, and load balancing.
- Experience troubleshooting production systems and distributed applications.
- Knowledge of high availability, scalability, fault tolerance, and disaster recovery.
- Strong incident management and problem-solving skills.
- Experience with Git and version-control systems.

Good to Have

- AWS EKS / Azure AKS / Google GKE.
- Helm.
- Argo CD / Flux.
- Ansible.
- Jenkins / GitHub Actions / GitLab CI.




- OpenTelemetry.
- Datadog / Recent Relic / Dynatrace.
- Grafana Loki / Promtail.
- Kafka.
- Redis.
- Service mesh technologies such as Istio.
- CloudFormation / ARM / Bicep.
- Chaos engineering.
- Experience with microservices architecture.
- Knowledge of SRE practices and Google SRE principles.
- Cloud certifications.

Education

- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.

Preferred Candidate Profile

- 2–7 years of experience in SRE, DevOps, Cloud Infrastructure, Platform Engineering, or Production Engineering.
- Strong hands-on experience managing production systems.
- Good understanding of reliability and observability principles.
- Strong troubleshooting and incident-management capabilities.
- Ability to automate repetitive operational tasks.
- Comfortable working with development, cloud, infrastructure, and security teams.
- Strong communication, documentation, and problem-solving skills.

Key Skills for Naukri

SRE, Site Reliability Engineer, DevOps, Cloud, AWS, Azure, GCP, Linux, Kubernetes, Docker, Terraform, CI/CD, Prometheus, Grafana, Observability, Monitoring, Logging, Alerting, Incident Management, Production Support, Root Cause Analysis, RCA, SLI, SLO, SLA, Python, Bash, Infrastructure as Code, Microservices, High Availability, Disaster Recovery

Suggested Naukri Job Titles

- Site Reliability Engineer
- SRE Engineer
- Site Reliability Engineer – Cloud
- DevOps/SRE Engineer
- Cloud SRE
- Production Engineer
- Reliability Engineer
- Platform Reliability Engineer
- Infrastructure Reliability Engineer
- DevOps Engineer – SRE

Send your resume at [email protected]

📌 Site Reliability Engineer (India)
🏢 Spark Tech Wave Innovation
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (india) / india