- Contribute to incident response and on-call rotation - own and resolve escalated incidents, identifying and mitigating root causes associated with distributed computing architecture (client server, intranet/internet), h/w platforms and resources: CPU, memory, virtualization, clustering and Cloud computing.
- Support the migration of markets applications infrastructure to Google Cloud Platform (GCP).
- Collaborate with systems engineers, SRE''s and Product engineering teams to monitor, maintain and troubleshoot our Markets systems.
- Collaborate with stakeholders to improve observability and alerting of our infrastructure to enable data-driven business decisions, faster issue detection and incident resolution.
- Take accountability for delivery of moderately-complex incidents, problems and changes.
- Lead technical discussions for your own work and present solutions and approaches.
- Assist in automating tasks to enhance system scalability and reliability.
- Collaborate with cross-functional teams to improve system performance and efficiency
- Act as a mentor to L2 and L1 colleagues.
Skills:
- Minimum of 5 years experience with Linux Windows-based Operating Systems support and administration.
- Minimum of 3 years supporting Cloud-based platform(s).
- Strong problem-solving and analytical abilities.
- Excellent communication and teamwork skills.
- Experience and knowledge of working with distributed systems.
- Exposure to working with metrics monitoring: Splunk, Grafana etc.