27 Sep
|
Marsh Mclennan
|
Delhi
27 Sep
Marsh Mclennan
Delhi
What can you expect:
The Databricks Platform Support Engineer (L2) supports the reliability, performance, and security of Marshs Databricks analytics platform across AWS and Azure, including Unity Catalog governance. This role provides day-to-day operational support for Databricks workspaces and compute, resolves incidents and service requests, assists with standard changes and upgrades, and partners with Cloud and Data Engineering teams to deliver a stable, compliant, and cost-effective platform.
We will count on you to:
Daily operations responsibilities include following items/activities:
- Provide Level 2 support for Databricks workspaces, clusters, jobs/workflows, pools, libraries, runtime versions, and configurations across dev/test/prod.
- Monitor platform health and job execution; respond to alerts and production issues within defined SLAs.
- Troubleshoot common operational issues (cluster start failures, job failures, library conflicts, connectivity errors, quota/limit constraints, permission issues).
- Escalate complex issues to L3/engineering/vendor with complete diagnostics, logs, and reproduction steps.
- Support Unity Catalog operations
- Handle incidents, service requests, and problems through ITSM processes; maintain accurate ticket notes and customer communications.
- Contribute to Root Cause Analysis (RCA), corrective/preventive actions, and continuous improvement.
- Perform standard changes (runtime upgrades, library updates, cluster policy changes, UC permission updates) using change control and validation procedures.
- Maintain and update runbooks/SOPs, onboarding guides,
and troubleshooting knowledge articles.
- Familiar with features like Open Share
- Participate in on-call rotation and/or shift coverage for production workloads.
What you need to have:
- Bachelors degree in Computer Science/Engineering or equivalent practical experience.
- 3–6 years of experience in platform operations, cloud operations, or data platform support, including exposure to Databricks and/or Apache Spark ecosystems.
- Hands-on troubleshooting experience with:
- Job/workflow failures, cluster start/termination issues, library and dependency management
- Spark configuration basics and performance triage (at an L2 level)
- Working knowledge of both AWS and Azure fundamentals (IAM/identity, networking concepts, storage services, encryption/key management basics).
- Familiarity with security practices: access controls, secrets handling, audit logging concepts.
- Scripting/automation skills (Python and/or Bash/PowerShell) for operational tasks.
- Familiarity with governance concepts (e.g., Unity Catalog basics) and data access controls.
- Experience following ITIL-aligned practices (incident, problem, and change management) and maintaining operational documentation.
What makes you stand out:
- Strong customer support mindset; ability to communicate clearly with technical and non-technical stakeholders.
- Structured troubleshooting and incident handling; good ticket hygiene and documentation habits.
- Ability to prioritize multiple issues in a production environment and meet SLAs.
- Cooperative approach across Cloud, Security, and Data Engineering teams.
📌 Databricks Engineer (Delhi)
🏢 Marsh Mclennan
📍 Delhi