24 Sep
|
Pinnacle Services And Solutions
|
Pune
24 Sep
Pinnacle Services And Solutions
Pune
Key Responsibilities
- Design and build scalable ingestion pipelines for structured, semi-structured, unstructured, and vector data from diverse sources, including Oracle Database, SQL Server, PostgreSQL, Pinecone, MongoDB, REST APIs, JSON, XML, CSV, Excel, logs, documents, SaaS applications, ERP systems, SharePoint, Jira, NetSuite, and Microsoft Dynamics 365.
- Implement appropriate ingestion methods, including full loads, incremental loads, Change Data Capture (CDC), batch, micro-batch, and streaming, with watermarking, checkpointing, upserts, and deletion handling.
- Design and develop reliable ETL/ELT pipelines using SQL, Python, PySpark, Databricks, Spark SQL, Delta Lake, and medallion data layers.
- Transform raw data into clean, standardized, enriched, and analytics-ready datasets while managing validation, deduplication, schema evolution, reconciliation, retries, error handling, recovery, and pipeline monitoring.
- Develop and optimize complex SQL queries, stored procedures, views, indexes, partitions, and other database objects.
- Administer databases, including user access, security, backup and recovery, monitoring, patching, migration, high availability, and performance tuning.
- Prepare data for analytics, machine learning, and Generative AI applications, including embeddings, vectorization, semantic search, vector databases, and Retrieval-Augmented Generation (RAG).
- Implement data-quality checks, metadata management,
security controls, and data-governance standards.
- Troubleshoot production database and data-pipeline issues and perform root-cause analysis.
- Collaborate with data architects, application teams, analysts, data scientists, and AI engineers.
- Mentor junior team members and contribute to data-engineering and database best practices.
Required Skills
- Strong hands-on experience with Oracle Database, SQL, and PL/SQL.
- Experience with relational database design, query optimization, indexing, partitioning, and performance tuning.
- Working knowledge of MongoDB, including document modelling, aggregation, indexing, and administration.
- Experience with Databricks, Apache Spark, Spark SQL, PySpark, and Delta Lake.
- Experience building batch, micro-batch, streaming, or near-real-time data pipelines.
- Knowledge of database administration, including backup and recovery, security, monitoring, migration, and high availability.
- Proficiency in PySpark, Python, or another data-engineering scripting language.
- Knowledge of vector databases, embeddings, semantic search, and RAG-based data preparation.
- Solid analytical, troubleshooting, communication, and documentation skills.
This is an immediate hiring, initial part-time. Depending on performance it will be a full-time.
Pay: Up to ₹900,000.00 per year
Work Location: In person
📌 Data Engineer (Pune)
🏢 Pinnacle Services And Solutions
📍 Pune