Location: Coimbatore | Work Mode: Full time, In-Office
Millions of businesses and products are only useful when the data can be trusted.
About Pepagora
Pepagora is trust infrastructure for global B2B growth. We operate a verification-first ecosystem that reduces trust friction, improves deal velocity, and enables SMEs to expand across borders with confidence. Our core product system - TruBadge, ESR Score, Janu AI, and OneInbox - powers signal-driven decision-making for businesses worldwide. Incorporated in Singapore, Pepagora is at a Pre-Series A stage and scaling rapidly across international markets.
About the Role
Build the data foundation behind ingestion, enrichment, entity resolution, analytics, search and AI. Make data observable, governed and usable by product systems - not trapped in one-off scripts.
What You'll Be Doing
Design ingestion and transformation pipelines
Build reliable canonical datasets and data-quality checks
Support entity resolution, enrichment and feature generation
Create datasets/interfaces for search, analytics and AI
Implement lineage, monitoring, backfills and failure recovery
Partner with backend and AI engineers on data contracts
Who We're Looking For:
Typically 4 - 7 years in data engineering or data-platform roles
Strong SQL and Python
Experience building production ETL/ELT pipelines
Solid data modelling and data-quality fundamentals
Cloud data/storage experience and operational ownership
Positive to Have:
Airflow/Dagster/dbt or similar
Spark/large-scale processing
Kafka/event streaming
Data lake/warehouse architectures
Search/vector data pipelines