21 Aug
|
Rhysley
|
New Delhi
Role Overview
We are looking for a highly experienced and hands-on Senior Server Administrator who will take end-to-end ownership of our production server infrastructure.
This is not a routine support or maintenance role. The role is focused on ensuring server availability, stability,
security, performance, scalability, and reliability. The ideal candidate should be capable of independently managing production infrastructure, handling critical incidents, implementing high-availability solutions, and proactively preventing infrastructure failures.
The candidate should have strong hands-on server administration experience along with an SRE-oriented approach to reliability, monitoring, automation, and proactive problem-solving.
Key Responsibilities
1. Production Server Ownership
• Own and manage production servers end-to-end.
- Ensure high availability, stability, performance, and minimum downtime.
- Take complete responsibility for server health and reliability.
- Monitor server capacity, performance, and resource utilization.
- Manage server configuration, upgrades, patching, and lifecycle activities.
1. High Availability & Failover
• Design and maintain highly available multi-server environments.
- Configure and manage load balancing using NGINX / HAProxy.
- Implement automatic failover and redundancy for critical services.
- Ensure infrastructure remains available during server or service failures.
- Regularly test failover and recovery procedures.
1. Monitoring & Incident Management
• Implement and manage server monitoring using Prometheus, Grafana, ELK, or equivalent tools.
- Configure real-time alerts for server, application, network, and service issues.
- Respond quickly to production incidents and restore services with minimum impact.
- Perform Root Cause Analysis (RCA) for major incidents.
- Identify recurring issues and implement permanent solutions.
1. Server Security & Hardening
• Perform server security hardening and secure configuration.
- Manage SSH, firewalls, Fail2ban, access permissions, and related security controls.
- Perform regular security patching and system updates.
- Implement secure access mechanisms, including MFA/RBAC where required.
- Identify and address server-level security vulnerabilities.
1. Backup & Disaster Recovery
• Manage and monitor server backup systems.
- Ensure backup integrity and availability.
- Maintain disaster recovery procedures for critical systems.
- Conduct regular backup restoration and recovery testing.
- Ensure critical production services can be restored effectively after failures.
1. Deployment & Reliability
• Manage production deployments with minimal or zero downtime.
- Implement deployment validation and rollback strategies.
- Coordinate with development teams for reliable production releases.
- Identify performance and reliability issues before they affect users.
- Continuously improve infrastructure stability and availability.
1. Performance & Capacity Management
• Monitor CPU, memory, disk, network, and other server resources.
- Identify and resolve performance bottlenecks.
- Plan server capacity based on business and application requirements.
- Recommend infrastructure scaling and optimization initiatives.
1. Automation
• Automate repetitive server administration and operational activities.
- Use Bash, Python, or equivalent scripting for automation.
- Reduce manual intervention and operational dependency.
- Improve efficiency, reliability, and consistency through automation.
1. Documentation & Operational Excellence
• Maintain server architecture documentation, SOPs, recovery procedures, and configuration records.
- Document major incidents, RCA findings, and preventive actions.
- Maintain clear operational processes for production infrastructure.
- Ensure knowledge is properly documented for business continuity.
Required Technical Skills
- 8+ years of hands-on experience in Senior Server Administration / System Administration /
Infrastructure Administration.
- Strong hands-on experience managing production servers.
- Strong experience with Linux server administration.
- Experience with High Availability and failover architecture.
- Hands-on experience with NGINX / HAProxy or equivalent load-balancing technologies.
- Experience with Prometheus, Grafana, ELK, or equivalent monitoring tools.
- Strong troubleshooting skills across server, system, and networking issues.
- Valuable understanding of TCP/IP, DNS, HTTP/HTTPS, SSL/TLS, firewalls, and networking fundamentals.
- Experience with server security hardening, SSH, firewall, Fail2ban, patching, and access management.
- Strong understanding of backup and disaster recovery.
- Experience with production deployments, rollback, and minimal-downtime strategies.
- Good scripting skills in Bash and/or Python.
- Understanding of server performance, capacity planning, and infrastructure scalability.
Mandatory Experience
Candidates must have demonstrated hands-on experience in:
- Production server ownership and administration
- Linux server administration
- High Availability and automatic failover
- NGINX / HAProxy or equivalent load balancing
- Infrastructure/server monitoring and alerting
- Production incident management and RCA
- Server security hardening
- Backup and Disaster Recovery
- Server automation using Bash/Python
- Performance optimization and capacity management
- Independent production troubleshooting
📌 Senior Server Administrator (New Delhi)
🏢 Rhysley
📍 New Delhi