07 Aug
|
Hudson Data
|
Gurugram
07 Aug
Hudson Data
Gurugram
Role & responsibilities
About the Role
Hudson Data is seeking a Data Engineer with strong expertise in Python, SQL, GCP, and BigQuery to build scalable data platforms that power analytics, machine learning, and AI-driven products.
You will design reliable ETL/ELT pipelines, develop cloud-based data models, and work closely with data scientists, analysts, and product teams to make high-quality data available for advanced analytics and AI use cases.
Key Responsibilities
- Design and maintain scalable ETL/ELT pipelines using Python, SQL, and GCP services.
- Build optimized data models and analytics-ready datasets in BigQuery.
- Integrate data from APIs, databases, files, and third-party platforms.
- Support AI and machine-learning workflows through reliable feature and training datasets.
- Implement data-quality checks, reconciliation, monitoring, logging, and error handling.
- Optimize SQL queries, pipeline performance, and BigQuery cost efficiency.
- Orchestrate batch and near-real-time workflows using Airflow or Cloud Composer.
- Maintain data governance, security, lineage, and access controls.
- Collaborate with data scientists, analysts, engineers, and business stakeholders.
Required Skills
- Strong hands-on experience with Python for data processing and automation.
- Advanced SQL, including complex joins, window functions, CTEs, query optimization, and performance tuning.
- Strong experience with Google Cloud Platform.
- Hands-on expertise in BigQuery, including data modeling, partitioning, clustering, and cost optimization.
- Experience building production-grade ETL/ELT pipelines.
- Familiarity with Airflow or Cloud Composer.
- Understanding of data warehousing, star schema, and dimensional modeling.
- Experience handling large datasets and implementing data-quality controls.
- Working knowledge of Linux/Unix and Git.
AI-Oriented Experience
- Experience preparing datasets for machine learning and generative AI applications.
- Familiarity with feature engineering, model pipelines, and MLOps concepts.
- Exposure to Vertex AI, embeddings, vector databases, or LLM-powered applications is preferred.
- Understanding of model monitoring, retraining workflows, and responsible AI data practices is a plus.
Preferred Qualifications
- Bachelors or Masters degree in Computer Science, Data Engineering, Data Science, Mathematics, or a related field.
- Google Cloud certification, especially Skilled Data Engineer, is preferred.
- Experience with dbt, Dataflow, Pub/Sub, Cloud Storage, or Cloud Functions is an advantage.
Preferred candidate profile
At Hudson Data, you will work at the intersection of data engineering, cloud analytics, AI, and machine learning. You will contribute to consulting engagements and proprietary AI products while solving real-world business problems for global clients.
📌 Data Engineer AI & CLOUD Analytics (Gurugram)
🏢 Hudson Data
📍 Gurugram