We are seeking an L2 Support Engineer to join our Cloud & Data Operations team. In this role, you will be responsible for maintaining, troubleshooting, and supporting production Big Data pipelines and AI/ML workflows hosted on Google Cloud Platform (GCP). You will act as the second line of technical escalation, resolving incident tickets, analyzing logs, debugging PySpark jobs, and leveraging contemporary AI productivity tools to accelerate root-cause analysis and operational maintenance.
Key Responsibilities
- Incident & Ticket Management: Monitor operational queues via ticketing systems (e.g., ServiceNow, Jira); prioritize, investigate, and resolve escalated L2 incidents within defined SLA boundaries.
- PySpark & Pipeline Support: Troubleshoot and fix failing PySpark scripts, Dataproc/Dataflow jobs, batch ETL failures, and data flow anomalies across distributed Big Data environments.
- Log Analysis & Monitoring: Perform deep-dive log analysis using GCP Cloud Logging (Log Explorer), Datadog, or Splunk to identify memory leaks, execution bottlenecks, or infrastructure failures.