06 Aug
|
Iss Corporate Solutions
|
Gurugram
06 Aug
Iss Corporate Solutions
Gurugram
Let s be #BrilliantTogether
Overview:
We are looking for a Senior Data Engineer to join our Index Engineering team in Gurgaon. You will be part of a team that builds and evolves critical data platforms on a modern cloud-native stack using dbt, BigQuery, and Apache Airflow on Google Cloud Platform. You will work with large-scale financial datasets, including securities master data, corporate actions, and market data from external vendors designing robust ELT pipelines, improving data models, and driving platform enhancements using modern engineering practices. This is a hands-on role for someone with strong data engineering fundamentals someone who understands how to design scalable pipelines, model data effectively, and build reliable systems that serve business-critical workloads.
Responsibilities:
Data Modelling Transformation
- Design and build data models that accurately represent business domains applying dimensional modelling, slowly changing dimensions, and normalisation/denormalisation trade-offs appropriate to the use case.
- Develop and maintain data transformation logic using dbt on BigQuery leveraging models, macros, incremental strategies, and tests to keep transformations modular, version-controlled, and well-documented.
- Define and enforce naming conventions, modelling standards, and layering practices (staging, intermediate, marts) across the data warehouse.
Data Pipeline Engineering
- Design, build, and maintain ELT pipelines that ingest, transform, validate, and serve financial data at scale with a focus on reliability, idempotency, and observability.
- Build ingestion frameworks for external data vendor feeds handling diverse file formats, schema variations, validation rules, reconciliation, and error recovery.
- Orchestrate pipeline workflows using Apache Airflow (Cloud Composer), managing dependencies, retries, SLAs, and alerting.
- Design and implement full- and incremental-load strategies, backfill mechanisms, and pipeline-recovery patterns.
Data Quality Reconciliation
- Implement data reconciliation processes to verify accuracy across upstream sources and internal datasets building automated checks for row counts, value matches, and business rule compliance.
- Define and enforce data quality standards through automated testing, validation layers,
and monitoring treating data quality as a first-class engineering concern.
- Set up monitoring and alerting for pipeline health and data freshness using tools such as Datadog.
Platform Performance
- Design partitioning, clustering, materialisation, and caching strategies to optimise query performance and manage storage costs in BigQuery.
- Build and support RESTful APIs (FastAPI) for internal and external data consumption.
- Support and improve CI/CD pipelines for data platform components.
- Participate in disaster recovery planning, testing, and documentation for data infrastructure.
Collaboration Continuous Improvement
- Collaborate with operations, product, and other engineering teams to translate business requirements into well-designed technical solutions.
- Establish and maintain data lineage, documentation, and cataloguing practices so that pipelines and models are understandable and auditable.
- Explore and apply Generative AI capabilities (e.g., LLM-based tooling, RAG patterns) to improve engineering workflows, documentation, and developer productivity.
- Troubleshoot production data issues, perform root-cause analysis, and implement fixes with a sense of urgency.
Qualifications:
- Financial services or fintech domain experience particularly in securities master data, corporate actions, index calculations, or market data vendor feeds.
- Experience with data platform modernisation rebuilding legacy pipelines using modern ELT approaches.
- Understanding of exchange calendars, business day logic, and how they affect data processing schedules.
- 9+ years of experience in data engineering, database development, or a related role.
- Bachelor s or Master s degree in Computer Science, Information Technology, or a related field.
Technical Skills:
Data Modelling SQL
- Strong data modelling skills dimensional modelling, star/snowflake schemas, slowly changing dimensions,
and the ability to design models that balance analytical performance with maintainability.
- Deep SQL expertise complex queries, window functions, CTEs, recursive queries, query plan analysis, and performance tuning on large datasets.
- Good understanding of data formats (Parquet, Avro, JSON, CSV), serialisation trade-offs, and working with structured and semi-structured data.
Modern Data Stack
- Hands-on experience with dbt modelling, transformations, tests, documentation, macros, and incremental models.
- Experience with a cloud data warehouse BigQuery preferred, or Snowflake/Redshift with willingness to work on BigQuery.
- Experience building data pipelines using Python and a workflow orchestration tool such as Apache Airflow or Cloud Composer.
- Solid understanding of ELT/ETL design patterns full vs. incremental loads, idempotent pipelines, backfill strategies, and dependency management.
Data Engineering Fundamentals
- Positive understanding of data lake and data warehouse architectures, and lakehouse concepts.
- Experience with data reconciliation building validation frameworks that compare data across sources and flag discrepancies.
- Understanding of data governance principles lineage, cataloguing, access control, and data quality management.
- Familiarity with version control (Git) and CI/CD practices.
Cloud Infrastructure (Nice to Have)
- Experience with Google Cloud Platform services beyond BigQuery Cloud Run, Cloud Composer, Cloud SQL, Cloud Storage.
- Experience building or working with REST APIs (FastAPI, Flask, or similar).
- Familiarity with API gateway platforms such as Apigee.
- Experience with monitoring and observability tools such as Datadog.
- Knowledge of relational databases such as SQL Server or PostgreSQL including stored procedures, indexing, and query execution plans.
- Familiarity with legacy ETL tools (SSIS, Informatica, or similar).
- Awareness of Generative AI concepts large language models, retrieval-augmented generation (RAG), agentic AI patterns and interest in applying them to data engineering and automation use cases.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Senior Data Engineer Python & SQL (Gurugram)
🏢 Iss Corporate Solutions
📍 Gurugram