10 Sep
|
GSPANN
|
Hyderabad
Job Title: Senior Site Reliability Engineer
Experience Required: 5+ Years
Job Type: Full-Time
Location: Hyderabad/Pune/Gurgaon
Domain: eCommerce / Site Reliability Engineering (SRE) / Production Support
About the Role
We are seeking a highly motivated System Engineer eCommerce SRE to join our Enterprise Support & Reliability Engineering team. This role is focused on ensuring the reliability, availability, performance, and operational stability of large-scale, highly transactional eCommerce platforms.
As part of the ESRE team, you will be responsible for supporting business-critical production systems, executing deployments, troubleshooting incidents, automating operational processes, and driving continuous reliability improvements. You will work closely with Engineering, DevOps, Product, Infrastructure, and Business teams to ensure seamless customer experiences while minimizing production issues and downtime.
This is an excellent prospect for professionals with strong experience in Unix/Linux Administration, Production Support, Shell Scripting, CI/CD, Rundeck, Kubernetes, Monitoring, Automation, and SRE practices.
Key Responsibilities
Production Support & Reliability
- Provide hands-on support for enterprise eCommerce platforms, microservices, Kubernetes environments, and integrated business applications.
- Monitor production applications, infrastructure, interfaces, jobs, and services to proactively identify and resolve issues.
- Respond to critical production incidents, ensuring minimal customer and business impact.
- Participate in incident management, problem management, and root cause analysis (RCA) activities.
- Analyze application logs, infrastructure metrics, and operational data to troubleshoot production issues.
- Support high-availability systems and contribute to reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
Deployments & Release Management
- Execute production deployments using Rundeck and/or CI/CD tools.
- Support release activities including deployment validation, rollback planning, post-deployment verification, and production readiness reviews.
- Troubleshoot deployment issues and ensure smooth release execution.
- Maintain compliance with change management and release governance processes.
Automation & Scripting
- Develop and enhance shell scripts to automate operational processes and repetitive tasks.
- Automate deployment workflows, health checks, monitoring activities, data processing, and operational procedures.
- Drive continuous process improvements to reduce manual intervention and improve operational efficiency.
Monitoring & Observability
- Monitor application and infrastructure health using enterprise monitoring and logging platforms.
- Improve monitoring coverage, alerting mechanisms, system visibility, and operational dashboards.
- Analyze alerts, logs, metrics, and system behavior to proactively prevent service disruptions.
- Support observability initiatives and reliability engineering best practices.
Collaboration & Continuous Improvement
- Collaborate with Engineering, DevOps, Infrastructure, Network, and Database teams to resolve complex issues.
- Create and maintain runbooks, knowledge articles, troubleshooting guides, and operational documentation.
- Support disaster recovery, business continuity, and application resiliency initiatives.
- Contribute to operational excellence through automation, standardization, and continuous improvement efforts.
Required Technical Skills
Production Support & Reliability Engineering
- Strong experience supporting business-critical production applications.
- Hands-on experience in Incident Management, Problem Management, RCA, and Service Restoration.
- Exposure to SRE, DevOps, and Reliability Engineering practices.
- Understanding of availability, scalability, resiliency, observability, and operational excellence.
Unix/Linux Administration
- Strong hands-on Unix/Linux administration experience.
- Expertise in Linux processes, filesystems, permissions, networking, services, jobs, and resource management.
- Experience troubleshooting application and system-level issues in production environments.
- Proficiency with Linux diagnostic and troubleshooting tools.
Shell Scripting & Automation
- Strong scripting skills using Bash/Shell scripting.
- Experience automating operational workflows, deployments, monitoring activities, and health checks.
- Ability to identify manual processes and implement automation solutions.
Deployment & CI/CD
- Hands-on experience with Rundeck and/or enterprise CI/CD tools.
- Experience performing production deployments, application restarts, workflow execution, and release validation.
- Knowledge of deployment controls, rollback strategies, and change management processes.
Monitoring & Troubleshooting
- Robust troubleshooting skills across application, infrastructure, integration, and platform layers.
- Experience working with monitoring, logging, and alerting tools.
- Ability to respond effectively to production alerts and business-critical incidents.
Kubernetes & Platform Support
- Experience supporting Kubernetes-based environments and containerized applications.
- Understanding of microservices architecture and platform operations.
Preferred Skills
- Experience supporting large-scale eCommerce platforms.
- Knowledge of Docker and Kubernetes administration.
- Experience with application monitoring, observability, and logging tools.
- Exposure to cloud platforms such as AWS, Azure, or GCP.
- Experience with traffic management platforms such as Akamai and Cloudflare.
- Familiarity with Infrastructure as Code (IaC) and automation frameworks.
- Understanding of SLIs, SLOs, Error Budgets, and Service Health Monitoring.
- Experience with batch processing, scheduling systems, and job monitoring.
Education & Experience
- Bachelor's Degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Minimum 5+ years of experience in System Engineering, Production Support, Site Reliability Engineering (SRE), DevOps, or related technical roles.
- Proven experience supporting mission-critical enterprise systems in high-availability environments.
Key Competencies
- Strong ownership mindset and accountability for production services.
- Excellent analytical and problem-solving abilities.
- Ability to troubleshoot complex issues independently.
- Strong communication and stakeholder management skills.
- Customer-focused approach with a commitment to reliability and uptime.
- Ability to collaborate effectively across Engineering, DevOps, Product, Infrastructure, and Business teams.
📌 Senior Site Reliability Engineer (Hyderabad)
🏢 GSPANN
📍 Hyderabad