08 Aug
|
Accenture
|
Bengaluru
08 Aug
Accenture
Bengaluru
Project Role : Custom Software Engineer
Project Role Description : Lead the effort to design, build and configure applications, acting as the primary point of contact.
Must have skills : PySpark
Positive to have skills : Python (Programming Language)
Minimum 5 Year(s) Of Experience Is Required
Educational Qualification : 15 years full time education
Summary:
We are seeking an experienced Senior Data Engineer to design, develop, and optimize enterprise-scale data products capable of processing high-volume, complex datasets. The ideal candidate has deep expertise in PySpark, distributed data processing, and performance optimization, with a strong understanding of modern Lakehouse architectures and cloud-native data platforms.
Roles & Responsibilities:
Design, develop, and maintain scalable data products using PySpark and Spark SQL.
Build reusable, modular, and production-ready data pipelines supporting enterprise analytics and AI use cases.
Develop data models for structured, semi-structured, and streaming data.
Performance Engineering
Optimize complex PySpark transformations and Spark SQL queries processing billions of records.
Improve application performance through:
Efficient partitioning strategies
Data skew mitigation
Broadcast joins
Bucketing and sorting
Caching and persistence
Predicate pushdown
Adaptive Query Execution (AQE)
File compaction and optimization
Analyze Spark execution plans and identify performance bottlenecks.
Optimize memory utilization, shuffle operations, executor configuration, and cluster resource consumption.
Large-Scale Data Processing
Build pipelines capable of handling TB to PB-scale datasets.
Process batch and near real-time data efficiently while maintaining SLA commitments.
Ensure scalability, resiliency, and fault tolerance of distributed workloads.
Data Quality & Reliability
Implement automated data validation and reconciliation checks.
Develop monitoring, alerting, and logging for production pipelines.
Perform root cause analysis for production failures and implement preventive improvements.
Cloud & Lakehouse Engineering
Develop solutions on cloud-based data platforms.
Work with Delta Lake, Iceberg, or Parquet-based architectures.
Optimize storage layout, partitioning, and file management for improved query performance.
Collaboration
Partner with architects, product owners, analysts, and data scientists to translate business requirements into scalable data products.
Participate in design reviews, code reviews, and architecture discussions.
Mentor junior engineers on PySpark best practices and performance tuning.
Professional & Technical Skills:
Python
PySpark
Spark SQL
SQL (Advanced)
Additional Information:
Big Data
Apache Spark
Distributed data processing
Data partitioning
Spark optimization
Shuffle optimization
Memory tuning
📌 Custom Software Engineer (Bengaluru)
🏢 Accenture
📍 Bengaluru