08 Sep
|
Palo Alto Labs
|
Pune
08 Sep
Palo Alto Labs
Pune
Role: Data Engineer
Location: Remote
Employment Type: Full-Time
Experience: 4-7 years
KEY RESPONSIBILITIES
· Design and implement CI/CD pipelines for data pipelines (ingestion jobs) and transformation projects (e.g., dbt on Databricks, SQL, notebooks)
· Orchestrate Databricks jobs and workflows end-to-end (ingestion, transformation, quality checks)
· Integrate automated testing into CI/CD, including schema and contract checks for data models and tables
· Implement FinOps best practices to support cost monitoring and allocation across the EDP
· Automate platform operations for Databricks and related services, such as workspace and cluster provisioning, library and runtime mgmt. and job deployment/config
· Implement and maintain identity and access management for Databricks and supporting cloud resources, including workspace- and cluster-level permissions, table- and view-level access controls (e.g., Unity Catalog or equivalent), service principals, groups, and roles for automated workloads, and RBAC & TBAC models in collaboration with EDP Architect
· Provide patterns, templates, and reusable modules (Terraform modules, Airflow DAG patterns, Databricks job templates) to accelerate onboarding of new projects
· Continuously evaluate and improve tooling, pipelines, and platform architecture to increase reliability, security, and developer productivity on Databricks
· Define the code promotion process to minimize impacts across domains as code is promoted to production
· Manage end-to-end orchestration using managed Airflow.
Contribute to defining and tracking SLA/SLO/SLIs for the platform and participate in incident response (triage, root cause analysis)
· Practical knowledge of IT Infrastructure technologies, cloud computing Azure), cybersecurity, and disaster recovery.
· Working knowledge of Azure ecosystem, including hands-on experience designing, building, and optimizing scalable data pipelines within cloud-native environments.
QUALIFICATIONS
· Hands-on experience with Databricks in production environment, including workspace and cluster management, jobs/workflows and integrations with orchestration tools
· Solid experience with CI/CD pipelines (e.g., GitHub Actions, GitLab CI, Azure DevOps, or similar) and Git-based workflows.
· Strong experience with Infrastructure as Code (IaC) and orchestration tools for provisioning and managing Databricks and cloud infrastructure.
· Experience implementing automated tests and quality gates in CI/CD pipelines.
· Ability to partner effectively with data engineers, analytics engineers, architects, and security teams.
MINIMUM EXPERIENCE & EDUCATION
· Bachelor’s degree in Computer Science, Information Technology, or related field preferred, or equivalent work experience.
· 4–7+ years in DevOps, Cloud Engineering, Site Reliability Engineering, or Platform Engineering, with at least 2+ years supporting data/analytics platforms.
· Experience operating production data workloads, including monitoring, logging, performance tuning, and incident response.
· Scripting skills (e.g., Python, Bash, PowerShell) for automation and integration.
· Experience in CPG, retail, manufacturing, or distribution environments preferred.
📌 Data Engineer- DataBricks (Pune)
🏢 Palo Alto Labs
📍 Pune