Location: Coimbatore | Work Mode: Full-time, In-Office
Millions of businesses and products are only useful when the data can be trusted.
About Pepagora
Pepagora is trust infrastructure for global B2B growth. We operate a verification-first ecosystem that reduces trust friction, improves deal velocity, and enables SMEs to expand across borders with confidence. Our core product system - TruBadge, ESR Score, Janu AI, and OneInbox - powers signal-driven decision-making for businesses worldwide. Incorporated in Singapore, Pepagora is at a Pre-Series A stage and scaling rapidly across international markets.
About the Role
Build the data foundation behind ingestion, enrichment, entity resolution, analytics, search and AI. Make data observable, governed and usable by product systems - not trapped in one-off scripts.
What You'll Be Doing
- Design ingestion and transformation pipelines
- Build reliable canonical datasets and data-quality checks
- Support entity resolution, enrichment and feature generation
- Create datasets/interfaces for search, analytics and AI
- Implement lineage, monitoring, backfills and failure recovery
- Partner with backend and AI engineers on data contracts
Who We're Looking For:
- Typically 4 - 7 years in data engineering or data-platform roles
- Strong SQL and Python
- Experience building production ETL/ELT pipelines
- Solid data modelling and data-quality fundamentals
- Cloud data/storage experience and operational ownership
Positive to Have:
- Airflow/Dagster/dbt or similar
- Spark/large-scale processing
- Kafka/event streaming
- Data lake/warehouse architectures
- Search/vector data pipelines