17 Sep
|
Space Inventive
|
Hyderabad
17 Sep
Space Inventive
Hyderabad
About the Role
We are seeking a Senior Data Engineer to join the RBQM Production Pod. In this role, you will design, build, and maintain data pipelines that power AI/GenAI-driven applications for Risk-Based Quality Management (RBQM) in clinical trials. You will be responsible for RAG document ingestion, vector database indexing, and developing scalable data APIs to support AI applications.
Robust expertise in up-to-date AI data architectures is essential, as traditional ETL and data warehouse experience alone will not meet the requirements of this role.
Key Responsibilities
Design and build RAG document ingestion pipelines (chunking, embedding, vector indexing) for clinical trial quality data
Build and manage vector databases (AWS OpenSearch) for RAG-powered AI workflows
Develop batch and streaming ETL/ELT pipelines from scratch for unstructured clinical data (PDF, DOCX, clinical reports)
Build and expose data APIs for AI application consumption
Optimize chunking strategies, embedding generation,
and retrieval performance for RAG architectures
Manage data quality, lineage, and governance for AI/ML data pipelines
Deploy and maintain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB)
Collaborate with Data Scientists and Backend Developers in an integrated pod team
Requirements
5–7+ years hands-on data engineering at scale
RAG document ingestion pipelines (chunking, embedding, vector indexing) — traditional ETL/DWH experience alone is not sufficient
AWS OpenSearch (vector database for RAG workflows)
Python (advanced) including SQL/Spark SQL
Unstructured data transformation (PDF, DOCX) for RAG/LLM applications
AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB
Containerization: Docker
Must demonstrate building custom pipelines from scratch — not just configuring out-of-the-box services
📌 Data Engineer Rag Hyderabad
🏢 Space Inventive
📍 Hyderabad