Role & responsibilities
Monitor and maintain the health and performance of the servers & applications
Respond promptly and effectively to technical issues reported by end users, identifying root causes and implementing solutions.
Collaborate with the development team to prioritize, diagnose, and resolve incidents, aiming for minimal disruption to operations.
Participate in an on-call rotation to provide support for critical issues, ensuring timely incident resolution.
Monitor system alerts and logs to proactively identify potential problems, taking preventive actions when necessary.
Perform regular system checks, updates, and patches to ensure the application's security and stability.
Coordinate with internal or external vendors or partners if necessary to address issues related to data feeds we depend on.
Document incident details, troubleshooting steps, and resolution outcomes for future reference
Communicate effectively with agents and stakeholders, providing updates on ongoing incidents and issue resolutions
Contribute to the continuous improvement of the support process by identifying recurring issues and suggesting process enhancements
Preferred candidate profile
Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent experience)
More than 4 years of experience with AWS Cloud (S3, EC2, Transfer family, Secret Manager,) & on-Premise infrastructure setting (Windows, Linux Redhat,).
Proven experience in technical support or production support roles in various settings.
Familiarity with IoT concepts and data streams, with the ability to troubleshoot data Network connectivity issues (TCP DUMP, flow analysis)