19 Sep
|
Gradientflo Labs
|
Hyderabad
19 Sep
Gradientflo Labs
Hyderabad
Role Overview
You’ll be the custodian of truth at Vibecoderz. Every decision — from TutorAgent personalization to growth funnels — depends on the data pipelines, analytics frameworks, and event taxonomies you design.
This is not about reporting dashboards alone. This is about building a data foundation that fuels personalization, insights, and growth at scale. You’ll own the Developer Graph, Firestore events, Neo4j pipelines, and Amplitude/Mixpanel dashboards, and ensure everyone — from engineers to CMO — can trust the numbers they see.
Your mission is to make data a first-class citizen across Vibecoderz: accessible, accurate, and actionable.
Key Responsibilities
1. Data Architecture & Pipelines
- Design and own Vibecoderz’ data stack: Firestore (events), Neo4j (Developer Graph), Redis (cache), BigQuery (analytics warehouse).
- Build real-time pipelines from TutorAgent sessions to Developer Graph updates.
- Ensure data quality, lineage, and governance across all systems.
- Analytics & Event Taxonomy
- Define Vibecoderz’ event schema (sign-ups, activation, Byte Course completions, artifact usage, retention markers).
- Partner with PM/Head of Growth to implement cohort analysis, funnels, and churn prediction.
- Enforce a consistent taxonomy across product, marketing, and community data.
- Personalization & Developer Graph
- Architect and optimize the Developer Graph in Neo4j.
- Design pipelines for real-time knowledge updates and personalized course recommendations.
- Partner with AI/Prompt Engineers to use graph data for context injection and personalization.
- Growth & Retention Analytics
- Provide the CMO/Head of Growth with conversion, retention, and LTV dashboards.
- Run A/B test analysis with statistical rigor.
- Measure engagement and learning outcomes for every course flow.
- Team Leadership & Data Culture
- Build and mentor a data team (data engineers, analysts, ML engineers).
- Establish data-driven culture across Vibecoderz: every decision tied to metrics.
- Champion “single source of truth” dashboards for leadership and investors.
- Compliance & Privacy
- Ensure compliance with GDPR, SOC 2, and data privacy regulations.
- Partner with Security & Compliance head on data encryption and anonymization policies.
Success Metrics
90 Days (Probation):
- Define Vibecoderz’ event taxonomy and implement tracking for 3 core flows (Text → Course, Voice Tutor, Mini-App).
- Deliver first growth dashboard (activation, retention, conversion) in Amplitude/Mixpanel.
- Deploy Developer Graph MVP in Neo4j with sample learner profiles.
12 Months:
- Developer Graph fully powering hyper-personalization in TutorAgent.
- Real-time pipelines handle 100K+ DAUs with <1s latency for graph updates.
- Growth metrics tracked daily: LTV, churn, cohort retention, funnel drop-offs.
- Vibecoderz achieves data-driven OKR execution with >80% adoption of analytics dashboards by leadership.
Must-Haves
- 10+ years in data engineering/analytics leadership at product-first companies.
- Deep expertise in event-driven architecture, pipelines, and real-time analytics.
- Proven ability to scale data systems to 100K+ DAUs.
- Hands-on with Neo4j (Graph DB),
Firestore, Redis, BigQuery.
- Robust background in cohort analysis, churn modeling, LTV/CAC analysis.
Nice-to-Haves
- Experience in EdTech or developer platforms.
- Contributions to open-source data tools or analytics frameworks.
- Prior founding data leader or startup scaling experience.
- Familiarity with AI-driven personalization and recommendation systems.
Tech Stack Visibility
- Data Sources: Firestore (events), GitHub, TutorAgent logs
- Graph DB: Neo4j Aura
- Pipelines: Pub/Sub, Cloud Tasks, Cloud Workflows
- Storage/Warehouse: BigQuery, Firestore, Redis
- Analytics: Amplitude, Mixpanel, GA4, custom dashboards
- Infra: GitHub Actions, Terraform, GCP
- Compliance: GDPR, SOC 2, ISO 27001
Assessment
Objective: Validate ability to design data pipelines + analytics framework for Vibecoderz.
Assessment (3-Part):
1. Data Architecture Design (Written)
- Draft a data architecture doc for Vibecoderz covering:
- Event taxonomy for TutorAgent (session start, artifact generated, quiz taken).
- Pipeline flow from Firestore → Pub/Sub → Neo4j → BigQuery.
- Data governance policies (lineage, retention, anonymization).
- Analytics Dashboard (Practical)
- Build a sample dashboard in Amplitude/Mixpanel/Looker showing:
- Funnel: Sign-up → Course Start → Artifact Created → Certification.
- Retention: 1-day, 7-day, 30-day.
- Conversion: Free → Pro → Teams.
- Coding Assessment
- Implement a Python ETL pipeline that:
- Listens to Pub/Sub events.
- Writes processed data into Neo4j + BigQuery.
- Handles error retries + logging.
Deliverables:
- Architecture doc (3–4 pages).
- Dashboard screenshots or link.
- GitHub repo with ETL code + instructions.
📌 Head of Data Analytics (Hyderabad)
🏢 Gradientflo Labs
📍 Hyderabad