25 Aug
|
Nameless
|
Chennai
Data Engineer (Spark/Scala)
About the Role
We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on-premises and cloud environments. This role requires deep hands-on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high-performance data pipelines. You will work extensively on complex on-prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring efficient, reliable data movement across heterogeneous environments.
Key Responsibilities
-
Design, develop, and maintain large-scale data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark
-
Build and support complex on-premises data workflows, including migration/hybrid on-prem-to-cloud integration patterns
-
Integrate data across diverse file systems (on-prem file shares, NAS, HDFS, S3) and formats JSON, Parquet, Fixed-Length, CSV, Excel, Avro
-
Write efficient, optimized SQL for data extraction, transformation, and loading across relational databases
-
Connect to and extract data efficiently from various source databases, tuning queries and pipelines for performance at scale
-
Develop and maintain workflow orchestration using Airflow (or similar schedulers) for reliable, monitored pipeline execution
-
Write clean, production-grade Python code for data processing, automation, and tooling
-
Build and maintain unit/integration tests for data pipelines to ensure data quality and reliability
-
Create and maintain clear technical documentation for pipelines, data flows, and system architecture
-
Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues in complex, multi-system workflows
-
Collaborate with cross-functional teams (data science, analytics, application engineering) to support downstream data consumption
-
Support cloud integration efforts, particularly with Azure, as workloads evolve from on-prem to hybrid/cloud architectures
Required Qualifications
Primary Skills:
-
Strong hands-on experience with Apache Spark and Databricks for large-scale data processing
-
Proficiency with Amazon S3 for data storage and pipeline integration
-
Strong SQL skills - query optimization, complex joins, performance tuning
-
Proven experience integrating data across various file systems and formats: JSON, Parquet, Fixed-Length, CSV, Excel, Avro, etc.
-
Strong knowledge of Scala Spark and PySpark for distributed data processing
-
Strong Python programming skills for scripting, automation, and data engineering tasks
-
Strong experience connecting to and efficiently extracting data from databases (relational/other), including performance-conscious extraction strategies
-
Demonstrated experience working on complex on-prem data workflows (multi-system integration, legacy system data extraction, hybrid on-prem/cloud pipelines)
-
Experience leveraging coding assistant tools and implementing AI agents to enhance development productivity and task execution.
Secondary Skills:
-
Experience with Azure cloud services (storage, compute, data services)
-
Experience with Apache Airflow for workflow orchestration and scheduling
-
Experience writing automated tests for data pipelines (unit, integration, data quality checks)
-
Strong documentation skills able to clearly document pipelines, data lineage, and technical designs
Good to Have:
-
Working knowledge of Java
-
Familiarity with React for building internal tooling/dashboards
-
Experience with Prefect for workflow orchestration
-
PBM (Pharmacy Benefit Management) / Healthcare domain knowledge
Data Engineer (Spark/Scala)
About the Role
We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on-premises and cloud environments. This role requires deep hands-on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high-performance data pipelines. You will work extensively on complex on-prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring efficient, reliable data movement across heterogeneous environments.
Key Responsibilities
-
Design, develop, and maintain large-scale data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark
-
Build and support complex on-premises data workflows, including migration/hybrid on-prem-to-cloud integration patterns
-
Integrate data across diverse file systems (on-prem file shares, NAS, HDFS, S3) and formats JSON, Parquet, Fixed-Length, CSV, Excel, Avro
-
Write efficient, optimized SQL for data extraction, transformation, and loading across relational databases
-
Connect to and extract data efficiently from various source databases, tuning queries and pipelines for performance at scale
-
Develop and maintain workflow orchestration using Airflow (or similar schedulers) for reliable, monitored pipeline execution
-
Write clean,
production-grade Python code for data processing, automation, and tooling
-
Build and maintain unit/integration tests for data pipelines to ensure data quality and reliability
-
Create and maintain transparent technical documentation for pipelines, data flows, and system architecture
-
Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues in complex, multi-system workflows
-
Collaborate with cross-functional teams (data science, analytics, application engineering) to support downstream data consumption
-
Support cloud integration efforts, particularly with Azure, as workloads evolve from on-prem to hybrid/cloud architectures
Required Qualifications
Primary Skills:
-
Strong hands-on experience with Apache Spark and Databricks for large-scale data processing
-
Proficiency with Amazon S3 for data storage and pipeline integration
-
Strong SQL skills - query optimization, complex joins, performance tuning
-
Proven experience integrating data across various file systems and formats: JSON, Parquet, Fixed-Length, CSV, Excel, Avro, etc.
-
Strong knowledge of Scala Spark and PySpark for distributed data processing
-
Strong Python programming skills for scripting, automation, and data engineering tasks
-
Strong experience connecting to and efficiently extracting data from databases (relational/other), including performance-conscious extraction strategies
-
Demonstrated experience working on complex on-prem data workflows (multi-system integration, legacy system data extraction, hybrid on-prem/cloud pipelines)
-
Experience leveraging coding assistant tools and implementing AI agents to enhance development productivity and task execution.
Secondary Skills:
-
Experience with Azure cloud services (storage, compute, data services)
-
Experience with Apache Airflow for workflow orchestration and scheduling
-
Experience writing automated tests for data pipelines (unit, integration, data quality checks)
-
Strong documentation skills able to clearly document pipelines, data lineage, and technical designs
Good to Have:
-
Working knowledge of Java
-
Familiarity with React for building internal tooling/dashboards
-
Experience with Prefect for workflow orchestration
-
PBM (Pharmacy Benefit Management) / Healthcare domain knowledge.
Data Engineer (Spark/Scala)
About the Role
We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on-premises and cloud environments. This role requires deep hands-on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high-performance data pipelines. You will work extensively on complex on-prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring efficient, reliable data movement across heterogeneous environments.
Key Responsibilities
-
Design, develop, and maintain large-scale data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark
-
Build and support complex on-premises data workflows, including migration/hybrid on-prem-to-cloud integration patterns
-
Integrate data across diverse file systems (on-prem file shares, NAS, HDFS, S3) and formats JSON, Parquet, Fixed-Length, CSV, Excel, Avro
-
Write efficient, optimized SQL for data extraction, transformation, and loading across relational databases
-
Connect to and extract data efficiently from various source databases, tuning queries and pipelines for performance at scale
-
Develop and maintain workflow orchestration using Airflow (or similar schedulers) for reliable, monitored pipeline execution
-
Write clean, production-grade Python code for data processing, automation, and tooling
-
Build and maintain unit/integration tests for data pipelines to ensure data quality and reliability
-
Create and maintain clear technical documentation for pipelines, data flows, and system architecture
-
Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues in complex, multi-system workflows
-
Collaborate with cross-functional teams (data science, analytics, application engineering) to support downstream data consumption
-
Support cloud integration efforts, particularly with Azure, as workloads evolve from on-prem to hybrid/cloud architectures
Required Qualifications
Primary Skills:
-
Strong hands-on experience with Apache Spark and Databricks for large-scale data processing
-
Proficiency with Amazon S3 for data storage and pipeline integration
-
Strong SQL skills - query optimization, complex joins, performance tuning
-
Proven experience integrating data across various file systems and formats: JSON, Parquet, Fixed-Length, CSV, Excel, Avro, etc.
-
Strong knowledge of Scala Spark and PySpark for distributed data processing
-
Strong Python programming skills for scripting, automation, and data engineering tasks
-
Strong experience connecting to and efficiently extracting data from databases (relational/other), including performance-conscious extraction strategies
-
Demonstrated experience working on complex on-prem data workflows (multi-system integration, legacy system data extraction, hybrid on-prem/cloud pipelines)
-
Experience leveraging coding assistant tools and implementing AI agents to enhance development productivity and task execution.
Secondary Skills:
-
Experience with Azure cloud services (storage, compute, data services)
-
Experience with Apache Airflow for workflow orchestration and scheduling
-
Experience writing automated tests for data pipelines (unit, integration, data quality checks)
-
Strong documentation skills able to clearly document pipelines, data lineage, and technical designs
Good to Have:
-
Working knowledge of Java
-
Familiarity with React for building internal tooling/dashboards
-
Experience with Prefect for workflow orchestration
-
PBM (Pharmacy Benefit Management) / Healthcare domain knowledge
Data Engineer (Spark/Scala)
About the Role
We are seeking an experienced Data Engineer to design, build, and optimize complex data workflows across on-premises and cloud environments. This role requires deep hands-on expertise in Apache Spark, Databricks, and Scala/PySpark, along with strong SQL and Python skills, to build robust, high-performance data pipelines. You will work extensively on complex on-prem workflows, integrating data across multiple file systems and formats, migrating and modernizing legacy processes, and ensuring efficient, reliable data movement across heterogeneous environments.
Key Responsibilities
-
Design, develop, and maintain large-scale data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark
-
Build and support complex on-premises data workflows, including migration/hybrid on-prem-to-cloud integration patterns
-
Integrate data across diverse file systems (on-prem file shares, NAS, HDFS, S3) and formats JSON, Parquet, Fixed-Length, CSV, Excel, Avro
-
Write efficient, optimized SQL for data extraction, transformation, and loading across relational databases
-
Connect to and extract data efficiently from various source databases, tuning queries and pipelines for performance at scale
-
Develop and maintain workflow orchestration using Airflow (or similar schedulers) for reliable, monitored pipeline execution
-
Write clean, production-grade Python code for data processing, automation, and tooling
-
Build and maintain unit/integration tests for data pipelines to ensure data quality and reliability
-
Create and maintain clear technical documentation for pipelines, data flows, and system architecture
-
Troubleshoot and resolve data pipeline failures, performance bottlenecks, and data quality issues in complex, multi-system workflows
-
Collaborate with cross-functional teams (data science, analytics, application engineering) to support downstream data consumption
-
Support cloud integration efforts, particularly with Azure, as workloads evolve from on-prem to hybrid/cloud architectures
Required Qualifications
Primary Skills:
-
Strong hands-on experience with Apache Spark and Databricks for large-scale data processing
-
Proficiency with Amazon S3 for data storage and pipeline integration
-
Strong SQL skills - query optimization, complex joins, performance tuning
-
Proven experience integrating data across various file systems and formats: JSON, Parquet, Fixed-Length, CSV, Excel, Avro, etc.
-
Strong knowledge of Scala Spark and PySpark for distributed data processing
-
Strong Python programming skills for scripting, automation, and data engineering tasks
-
Strong experience connecting to and efficiently extracting data from databases (relational/other), including performance-conscious extraction strategies
-
Demonstrated experience working on complex on-prem data workflows (multi-system integration, legacy system data extraction, hybrid on-prem/cloud pipelines)
-
Experience leveraging coding assistant tools and implementing AI agents to enhance development productivity and task execution.
Secondary Skills:
-
Experience with Azure cloud services (storage, compute, data services)
-
Experience with Apache Airflow for workflow orchestration and scheduling
-
Experience writing automated tests for data pipelines (unit, integration, data quality checks)
-
Strong documentation skills able to clearly document pipelines, data lineage, and technical designs
Good to Have:
-
Working knowledge of Java
-
Familiarity with React for building internal tooling/dashboards
-
Experience with Prefect for workflow orchestration
-
PBM (Pharmacy Benefit Management) / Healthcare domain knowledge
📌 Data Engineer-Spark,Scala (Chennai)
🏢 Nameless
📍 Chennai