Data Engineer - Big Data & PySpark (Bengaluru)

Data Engineer - Big Data & PySpark (Bengaluru)

19 Sep
|
Apexon
|
Bengaluru

19 Sep

Apexon

Bengaluru

Job Summary

We are looking for an experienced Software / Data Engineer to join our technology team. In this role, you will be responsible for designing, building, and maintaining high-throughput data processing pipelines that handle billions of records. You will evaluate incoming data, perform specialized hygiene and standardization, and build robust software integrations supporting client data onboarding and identity matching services.

Key Responsibilities

- Data Pipeline &
- Architecture:

Design, implement, and maintain scalable Big Data ETL pipelines using Apache Spark, Hadoop, and cloud infrastructure (AWS).

- Software Integration: Develop software components that integrate with keying processes, client datasets, and identity graph systems.

- Database &

- Query Optimization:

Write complex SQL queries, perform query tuning, optimize data schemas, and ensure high performance in a high-volume processing setting.

- Data Quality &

- Hygiene:

Implement fuzzy logic matching, perform data research, and conduct trend analysis to ensure data integrity, standardization, and quality.





- System Scalability &

- Capacity:

Participate in capacity monitoring, performance tuning, and cross-functional design to ensure systems scale effectively.

- Compliance &

- Security:

Adhere to security standards and regulatory requirements (including ePHI/data privacy guidelines) and contribute to technical documentation.

Technical Competencies:

- Big Data Technologies: 4+ years of hands-on experience with Apache Spark, Hadoop, and distributed computing environments.

- Programming Proficiency: Strong coding skills in Python, Scala, Java, or Shell scripting.

- Database &
- SQL:

Advanced proficiency in SQL, data modeling, database design, and query optimization.

- Cloud Infrastructure: Solid hands-on experience with AWS services and cloud-based data processing architectures.

- Data Matching / Logic: Familiarity or prior experience with fuzzy logic algorithms, data deduplication, or identity resolution tools is a strong plus.

📌 Data Engineer - Big Data & PySpark (Bengaluru)
🏢 Apexon
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: data engineer - big data & pyspark (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: data engineer - big data & pyspark (bengaluru) / bengaluru