Why Mizuho
At Mizuho, we provide the stability of an international industry leader with the career trajectory of a growing business. Our steady, strategic growth gives our people at all levels rewarding degrees of responsibility and richer work experience than a boutique firm or an established giant could offer alone.
It's the local expertise of our employees that makes our global network so powerful. By collaborating with colleagues and clients who have the same ambition and drive, you can amplify your sphere of influence and base of knowledge as part of one of the largest—and growing—banks in the world.
Role Objective:
- Senior reliability engineering for Infrastructure Operations. This position will be leveraging observability data, automation, and modern SRE practices. Working knowledge of Ops Ramp or equivalent, Info Sight, VMware Aria Operations (v Sphere/Horizon), Workspace ONE, Net App, and Service Now to effectively interpret signals, performance, and resiliency data. This role includes a strong AI understanding with automation and toil reduction mindset. MELT knowledge is critical.
Key Responsibilities:
- Apply SRE principles (SLIs/SLOs, alert quality, incident reduction) across Infrastructure Operations.
- Use working knowledge of VDI and VSI environments and infrastructure platforms to interpret health and performance signals from VMware Aria Operations.
- Diagnose cross-platform incidents, including:
- VMware v Sphere / ESXi / v Center
- Horizon components (brokers, agents)
- Log Analytics
- Windows & Linux workloads
- Appliance based workloads
- Translate infrastructure signals into:
- Actionable metrics and alerts
- Observability pipelines (APIs, exporters, telemetry flows) into Data Mart
- Platform-level insights for SRE consumption
- Development of AI Based resolution with solid automated based approaches.
- Partner with SRE / Operations teams to:
- Eliminate false positives and alert noise
- Improve dashboard accuracy and usability
- Drive proacti