Big Data Engineer (Bengaluru)

Big Data Engineer (Bengaluru)

06 Oct
|
EPAM Systems
|
Bengaluru

06 Oct

EPAM Systems

Bengaluru

Roles and Responsibilities

- Design robust batch and real-time data processing pipelines using Apache Spark/PySpark and Databricks.
- Design and implement scalable ETL/ELT pipelines for structured, semi-structured, and unstructured data.
- Define data ingestion, transformation, orchestration, storage, and serving patterns.
- Establish data architecture standards covering scalability, reliability, security, performance, and maintainability.
- Design data models, data processing frameworks, and reusable data engineering components.
- Optimize large-scale data processing workloads for performance and cost.

Databricks & Lakehouse

- Architect and implement Databricks Lakehouse solutions.
- Develop production-grade Spark/PySpark workloads using Databricks.
- Design data pipelines supporting analytics, ML, and AI workloads.
- Work with Delta Lake and up-to-date Lakehouse architecture patterns.
- Implement data quality, validation, lineage, monitoring, and governance mechanisms.
- Optimize Spark jobs, cluster configurations, partitioning, caching, and data storage strategies.
- Integrate Databricks with AWS services and enterprise data platforms.

AWS Cloud

- Design cloud-native data solutions using AWS.
- Work extensively with services such as:




- Amazon S3
- AWS Lambda
- Amazon EKS
- DynamoDB
- API Gateway
- IAM
- Event-driven AWS services
- Design secure and highly available data architectures.
- Implement cloud-native patterns for scalability, fault tolerance, and disaster recovery.
- Drive AWS cost optimization / FinOps for data workloads.

Data Pipelines & Integration

- Build reliable and reusable data ingestion and transformation frameworks.
- Integrate data from APIs, databases, files, event streams, and enterprise applications.
- Implement incremental processing, CDC, schema evolution, error handling, retries, and reconciliation.
- Establish pipeline monitoring and operational processes.
- Build data pipelines capable of supporting both analytical and ML workloads.

Candidates with experience in the following areas will be preferred:

- MLflow and model lifecycle management.
- Data pipelines supporting model training and inference.
- Retrieval-Augmented Generation (RAG) data pipelines.
- Vector databases such as pgVector, Pinecone, or Weaviate.
- Semantic search and embedding pipelines.
- GenAI data preparation

📌 Big Data Engineer (Bengaluru)
🏢 EPAM Systems
📍 Bengaluru

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: big data engineer (bengaluru) / bengaluru

Subscribe to this job alert:

Get the latest job offers by email for: big data engineer (bengaluru) / bengaluru