Job Description
About Fam (previously FamPay)
n
Fam is India’s first payments app for everyone above 11. FamApp helps make online and offline payments through UPI and FamCard. We are on a mission to raise a new, financially aware generation, and drive 250 million+ young users in India to kickstart their financial journey super early in their life.
n
n
We’re reimagining how the next generation experiences fintech - going beyond payments to build a lifestyle brand that blends money, identity, and everyday experiences into one seamless, intuitive journey.
n
n
Founded in 2019 by IIT Roorkee alumni, Fam is backed by some of the most respected investors around the world like Elevation Capital, Y-Combinator, Peak XV (Sequoia Capital) India, Venture Highway, Global Founder’s Capital and the likes of Kunal Shah, Amrish Rao as angel investors.
n
n
About the role:-
n
n
We are looking for a Data Engineer II (SDE-2) to join our data team. The ideal candidate will be a play a key role to develop of high performant and scalable Data Lake-house, moving us toward a world of sub-minute data latency and unified batch/streaming compute. This is an engineering-heavy role where you will manage complex CDC flows, optimize distributed query engines and leverage AI to accelerate our development lifecycle.
n
n
Must Have Requirements-
n
n
n
- Experience: 3–5 years in Data Engineering, specifically with distributed systems and cloud-native architectures.
n
- Coding: Expert-level Python/PySpark and SQL. Familiarity with Go/Java/Scala is a plus
n
- Infrastructure: Hands-on experience with AWS (S3, EKS, MSK) and Infrastructure-as-Code.
n
- Orchestration: Experience with Airflow or Temporal for complex workflow management.
n
- AI-Native: Proficiency in using AI tools (Claude, Codex, Copilot) to write, test, and document code efficiently.
n
- Systems Thinking: Ability to explain the trade-offs between different storage formats and processing frameworks.
n
- Tech Execution :
Drive key tech initiatives by preparing TRD and actively involve in design reviews.
n
- Domain Modelling - Should be hands on in designing Domain models for OLAP like Fact, Dimension, Cumulative , types of SCD’s and OBT pattern tables.
n
- Self Starter - Lead the team technically and bring in new ideas to contribute to the growth of the charter.
n
- Stakeholder Interaction - Interact with the Product & Key Stakeholders & help them by adding value to the business workflow with data & analytics.
n
n
n
Valuable to have -
n
n
n
- Real-time CDC: Ownership of high-throughput ingestion from RDBMS to Lakehouse using Debezium, PeerDB.
n
- Lakehouse Architecture: Designing and optimizing table formats (Iceberg, Delta, Hudi) for both performance and storage efficiency.
n
- Unified Compute: Developing robust ETL/ELT frameworks in PySpark and Flink (handling both batch and streaming workloads).
n
- Infrastructure & Ops: Managing data workloads on AWS (EMR, EKS, MSK, S3) and automating everything via Gitlab/Github Actions.
n
- Query & BI: Tuning Trino or Clickhouse to power real-time dashboards in Metabase, Superset, and PowerBI.
n
n
n
Our Tech Stack-
n
n
n
- Ingestion & CDC: OLake and PeerDB for near real-time sync from production systems into Alchemy; Debezium/Kafka for CDC-heavy use cases, with support for Kafka and S3-based sources.
n
- Lakehouse / Storage: Alchemy on Apache Iceberg with S3 as the data lake storage layer and AWS Glue Catalog for metadata; exposure to Delta/Hudi is a plus.
n
- Processing & Compute:
PySpark on EMR/EKS for batch and streaming workloads; Flink and Spark Structured Streaming fundamentals for low-latency pipelines.
n
- Streaming Platform: MSK / Kafka for event-driven ingestion, CDC propagation, replay, backfills, and operational monitoring through Kafka UI.
n
- Query & Serving Layer: Trino over Alchemy/Iceberg for lakehouse analytics, ClickHouse for high-throughput operational and real-time dashboards, and BigQuery exposure where applicable.
n
- Workflow Orchestration: Airflow for scheduled data pipelines, DQ/reconciliation DAGs, backfills, and SLA-driven jobs; Temporal for durable workflow execution in ingestion services.
n
- Data Quality & Governance: DQ checks, freshness/SLA monitoring, source-to-lake reconciliation, deduplication, schema evolution handling, and cataloging/lineage through OpenMetadata.
n
- Infrastructure & DevOps: AWS (S3, EKS, EMR, MSK), Kubernetes, Terraform/IaC, GitLab/GitHub Actions, observability via Grafana/CloudWatch, and production runbook discipline.
n
- BI & Analytics: Metabase, Superset, Tableau, and PowerBI for business dashboards; strong ability to model curated marts, fact/dimension tables, SCDs, and OBT patterns for stakeholder-facing analytics.
n
n
n
Perks That Go Beyond the Paycheck
n
n
n
- Relocation assistance to make your move seamless.
n
- Free office meals (lunch & dinner).
n
- Generous leave policy, including birthday leave, period leave, paternity and maternity support, and more.
n
- Salary advance and loan policies for any financial help.
n
- Quarterly rewards and recognition programs, and a referral program with great incentives.
n
- Access the latest gadgets and tools.
n
- Comprehensive health insurance for you and your family, mental health support.
n
- Tax benefits with options like food coupons, phone allowances, car/device leasing.
n
- Retirement perks like PF contribution, leave encashment and gratuity.
n
n
📌 Data Engineer (Bengaluru)
🏢 Fam
📍 Bengaluru