21 Aug
|
Skysoft
|
Gurugram
THE ROLE
We are building an enterprise data lake and analytics platform on Google Cloud for a large asset management firm. The platform does not exist yet — you would be one of the first engineers on it, setting up ingestion, modeling and orchestration from scratch and defining the standards the rest of the team builds to. The stack is BigQuery, Python, dbt and Cloud Composer (Airflow), with Terraform underneath. This is a hands-on engineering role with real ownership of architecture decisions.
REQUIRED SKILLS
- Python — production-grade code, not scripts. Modules and packaging, testing (pytest), error handling and retries, API and file ingestion. pandas or polars for local processing.
- SQL — advanced. Window functions, CTEs, complex joins, and the ability to reason about query cost and execution.
- BigQuery — partitioning and clustering, incremental load patterns, materialized and authorized views, slots versus on-demand, and cost control. Equivalent depth in Snowflake, Redshift or Databricks SQL is acceptable.
- dbt — layered project structure, generic and singular tests, incremental models and strategies, snapshots for SCD Type 2, macros, and documentation.
- Airflow / Cloud Composer — DAG design, dependencies, sensors, retries, idempotency, backfills and SLAs.
- Data modeling — dimensional modeling, fact and dimension design, choosing and defending a grain, slowly changing dimensions, and handling late-arriving and restated data.
- Google Cloud fundamentals — Cloud Storage and data lake layout, Parquet and Avro, IAM, service accounts, Secret Manager.
- Engineering practice — Git, code review, CI/CD, and infrastructure as code with Terraform.
- Data quality — designing checks, reconciliation against source systems, monitoring and alerting.
- Communication — able to work directly with business users to define requirements and translate them into models.
WHAT YOU WILL DO
1. Build batch and CDC ingestion from portfolio accounting, order management, custodian, administrator and market data sources into BigQuery.
2. Model the core data — securities, positions, transactions, prices, benchmarks, client and account hierarchies — in dbt, with proper grain and history.
3. Orchestrate everything in Cloud Composer with pipelines that are idempotent, re-runnable and safe to backfill.
4. Build data quality and reconciliation checks, with alerting that catches breaks before the business does.
5. Deliver curated marts for performance, risk, client and regulatory reporting.
6. Own the platform layer: Terraform, CI/CD, environments, access control and BigQuery cost management.
POSITIVE TO HAVE
Technical
- Datastream or other CDC from Oracle / SQL Server
- Dataflow or Apache Beam; Pub/Sub and streaming
- Dataplex, Data Catalog, BigLake
- Looker and LookML
- Great Expectations, Soda, or data contracts
- Iceberg or Delta alongside BigQuery
- VPC Service Controls, CMEK, regulated-data access design
Domain ( Not Mandatory )
- Asset management, investment management or capital markets
- Platforms such as SimCorp, Aladdin, Charles River, Eagle or Geneva
- Market and reference data: Bloomberg, LSEG, FactSet, MSCI
- Security master and identifiers (ISIN, CUSIP, SEDOL, FIGI)
- Positions, transactions, corporate actions, NAV, benchmarks
- Performance and attribution, risk, or regulatory reporting
📌 Senior Data Engineer (Gurugram)
🏢 Skysoft
📍 Gurugram