04 Aug
|
Apidel Technologies
|
Pune
04 Aug
Apidel Technologies
Pune
Role - Data Platform Observability & Monitoring Engineer
Role Purpose -
The Data Platform Observability & Monitoring Engineer will design, implement and operate the monitoring and observability capabilities for enterprise data and AI services. The role requires a strong data-engineering foundation and practical experience with Azure and enterprise observability platforms. The engineer will establish actionable telemetry across Azure Databricks, Azure Data Factory, Azure AI services, data pipelines and supporting infrastructure, enabling proactive detection, faster incident resolution, service-level reporting and cost-aware platform operations.
Key Responsibilities -
- Design and implement an enterprise observability architecture for data pipelines, Databricks workloads, ADF orchestration, AI services and supporting Azure infrastructure.
- Onboard and standardize logs, metrics, traces and operational events from Azure services into approved monitoring and analytics platforms.
- Configure and manage Event Hubs, diagnostic settings, Log Analytics workspaces, Application Insights, Azure Monitor and enterprise log-ingestion routes.
- Develop and maintain dashboards, alerts, service-health views and management reporting for availability, failures, latency, throughput, data freshness and resource consumption.
- Administer or support enterprise observability platforms such as Splunk, including indexes, source types, data inputs, retention, role access, dashboards and alerts.
- Define telemetry standards for data pipelines, including correlation IDs, run IDs, project metadata, error classification, business context and data-quality outcomes.
- Work with data engineers to embed structured logging, metrics and traceability into Azure Databricks, Azure Data Factory and streaming/batch solutions.
- Create monitoring patterns for end-to-end lineage and operational correlation across source systems, orchestration, processing, storage and downstream consumption.
- Investigate incidents using logs and telemetry,
perform root-cause analysis and recommend preventive engineering improvements.
- Define and report service-level indicators and objectives for critical data and AI services.
- Implement monitoring-as-code and automated deployment of alerts, dashboards, diagnostic settings and observability configuration.
- Monitor ingestion volume, retention and platform usage to maintain observability coverage while controlling cost.
- Support security, audit and compliance requirements by ensuring appropriate log coverage, retention, access controls and evidence.
- Produce runbooks, support documentation and training for engineering and operational teams.
Required Experience and Skills -
- 5+ years of experience in data engineering, cloud engineering, platform engineering or a closely related technical role.
- Robust data-engineering knowledge, including data pipelines, orchestration, batch/stream processing, data quality and production support concepts.
- Hands-on experience with at least one enterprise observability platform such as Splunk, Azure Monitor, Datadog, Dynatrace, Grafana/Prometheus or Elastic.
- Practical Azure experience, including diagnostic settings, Log Analytics, Application Insights, Event Hubs, storage and identity/access concepts.
- Experience designing actionable dashboards and alerts rather than collecting logs without defined operational use cases.
- Strong SQL skills and working scripting/development skills in Python, PowerShell or a comparable language.
- Experience investigating distributed-system incidents and performing structured root-cause analysis.
- Understanding of logs, metrics and traces, including telemetry quality,
cardinality, retention, sampling and correlation.
- Knowledge of secure access, data sensitivity, auditability and least-privilege controls for monitoring platforms.
- Strong written and verbal communication skills and the ability to work with engineering, operations, security and management stakeholders.
Preferred / Good to Have -
- Hands-on experience monitoring Azure Databricks, Unity Catalog, Spark workloads and Azure Data Factory pipelines.
- Splunk administration experience, including index design, source types, role-based access, dashboards, alerts and ingestion troubleshooting.
- Experience with OpenTelemetry and distributed tracing across applications, agents, APIs and data pipelines.
- Experience with Azure AI Foundry, machine-learning platforms, GenAI applications or AI-agent observability.
- Experience with infrastructure-as-code and CI/CD for monitoring configuration, such as Terraform, Azure DevOps or Git-based deployment.
- Knowledge of FinOps monitoring, consumption analytics and cost optimisation for Azure and Databricks.
- Experience defining SLI/SLO models, operational scorecards and incident-management processes.
- Relevant Azure, Databricks, Splunk or observability certification.
Professional Competencies
- Analytical and evidence-driven, with the ability to distinguish symptoms from root causes.
- Strong engineering mindset; able to improve instrumentation and code, not only create dashboards.
- Focused on actionable monitoring that supports operations, governance and management decisions.
- Able to work across multiple platforms and teams while maintaining common telemetry standards.
- Proactive in identifying monitoring gaps, noisy alerts and unnecessary ingestion cost.
- Clear communicator who can explain platform health and operational risk to technical and non-technical audiences.
Note No amount of fees or money will be asked during the time of interview or after selection
📌 Data Platform Observability & Monitoring Engineer (Pune)
🏢 Apidel Technologies
📍 Pune