Role Summary
We are hiring a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the Enterprise AI platform:
What You'll Own
- Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON).
- AWS-based data architecture: S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed.
- Data validation frameworks: GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting.
- Agent run logging & observability: designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback).
- AI Factory monitoring dashboards: operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent.
- ML data pipeline support: dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues.
- APIs: designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools.
- Data quality & testing discipline: idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production.
Key Skills — Non-Negotiable (Must-Have, Strong Level)
- Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting.
- SQL — strong hands-on ability,
including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries).
- AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale.
- Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions."
- ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built.
- Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV).
Key Skills — Good to Have
- Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable.
- FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs.
- ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself.
- Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats).
- Large-scale metadata querying — experience making file discovery quick across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue.
Skills:- Generative AI, LangGraph, ETL, databricks, Retrieval Augmented Generation (RAG) and FastAPI
📌 Data Science Engineer (Bengaluru)
🏢 J&F
📍 Bengaluru