- Work to understand any arising issues and overall application performance by enacting monitoring solutions.
- Conduct consistent and thorough analysis of current systems and work to reduce the quantity of existing problems, suggesting recent solutions to help upgrade & refine such systems.
- Provide support across a broad range of areas including monitoring, processes & tools, architecture, and Root Cause Analysis.
- Develop and maintain monitoring and alerting systems to proactively detect and resolve issues.
- Automate routine tasks to improve system efficiency and reduce downtime.
- Troubleshoot and resolve incidents and outages.
- Develop and implement automation scripts and tools to improve the efficiency and effectiveness of system management tasks.
- Identifying areas for improvement, and designing solutions that are scalable, reliable, and easy to maintain.
- Monitoring & acting on Alerts to avoid production outages, Incidents.
- Upkeeping of Run books for the Alerts
Skills
- 6-9 years of sysadmin experience in handling large-scale distributed system software deployments in cloud or in an on-premises environment.
- Strong cloud management foundation.
- Unix shells, Python & Go programming proficiency.
- Experience in MySQL or PostgresQL in database.
- Outstanding teammate who can collaborate and influence in a multifaceted environment.
- Excellent interpersonal, and written communication skills.
- Excellent debugging and troubleshooting skills.
- Ability to define standard operating procedures for supported platform features.
- Experience working with observability tools and practices(Prometheus, Grafana).
- Experience in troubleshooting and resolving incidents.
- Cloud experience in AWS (preferred) including hands-on experience with AWS-CLI.
- Hands-on experience in the Orchestration and Containerisation like Kubernetes, Containers.
- Experience with CI /CD (i. e. Jenkins, ArgoCD).
- Solid Understanding of Networking (firewall, connectivity, routing, iptables, subnet config, etc.).
- Experience with Linux OS and Shell/Python Scripting.
- Experience in programming with Python, Go
- Experience with API Gateway like Kong, Nginx based systems.
- Experience with security best practices and technologies.
- BS degree in Computer Science or a related technical field involving coding, or equivalent practical experience
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Site Reliability Engineer (Karnataka)
🏢 Zeta
📍 Karnataka
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.