The Role
You own the data substrate and the batch runtime of a product we are building from scratch. This is a data-specialist seat, not a general backend seat. Routine application development is guided in weekly reviews; what cannot be substituted is your depth in data modeling, set-based computation, and pipeline correctness. You will work self-directed, writing your own specs and decisions, with your tests and documentation carrying the quality bar between reviews.
Stack: Python, PostgreSQL, Django, server-rendered frontend (htmx).
What You Will Own
Core Responsibilities / Engineering Focus
- Deep third-party data ingestion
- Build reliable ingestion pipelines for external data sources
- Account for silently dropped webhooks through reconciliation with source-of-truth totals
- Large-scale batch computation
- Use set-based SQL for processing large datasets
- Ensure jobs complete within fixed nightly processing windows
- Multi-tenant reliability
- Maintain strict tenant isolation
- Ensure one tenant’s bad data does not impact another tenant’s processing
- Backfill & replay capabilities
- Design systems to support backfills and replays as standard capabilities
- Avoid relying on one-off emergency scripts
- Append-only & auditable data
- Maintain immutable, auditable records
- Ensure every automated decision can be reconstructed and traced
- Resilient vendor API integrations
- Build for failures and unreliable external systems
- Implement retries, circuit breakers and reconciliation mechanisms
- Operational reliability
- Build strong alerting and monitoring
- Handle database partitioning effectively
- Design and maintain idempotent jobs
Required
What We Are Looking For
Ideal Candidate Profile
- 5–8 years of experience building and operating production systems — years are a proxy; the skills below are the actual bar.
- Strong PostgreSQL expertise
- Schema design
- Query planning and optimization
- Batch performance at real scale
- Production data pipeline ownership
- Ingestion
- Transformation
- Reconciliation
- Data serving
- Experience should go beyond application CRUD/ORM work.
- Strong SQL & data thinking
- Thinks in sets, not loops
- Multi-step SQL transformations
- Window functions
- Statistical aggregates, percentiles & distributions
- Incremental computation
- Experience building cohort, retention or funnel metrics from raw event/order data
- Production Python backend experience
- Experience with any Python backend framework is fine
- Able to keep business logic cleanly separated from framework-specific code
- Strong testing mindset
- Tests are part of the development process, not an afterthought
- Especially experienced with testing batch jobs and reconciliation systems where data errors can remain unnoticed
- Production problem-solving / “war stories”
- Has dealt with real incidents such as:
- Queues backing up
- Webhooks silently dropping
- Data loss or inconsistencies
- Reconciliation catching issues that went unnoticed
- Strong written communication
- Comfortable with documentation-driven development
- Can write explicit specs, decisions and technical documentation
- High ownership
- Comfortable building systems from scratch
- Can independently own outcomes with minimal supervision
Good to have
- Production Django experience
- E-commerce / D2C platform experience
- Streaming & event-driven architecture
- Kafka or similar systems
- Familiarity with probabilistic data structures
- HyperLogLog
- Bloom Filters
- t-digest
- Comfort with server-rendered applications/screens when required
Skills: postgresql,django,python
📌 Data Engineer (India)
🏢 Skillyt
📍 India