19 Sep
|
Sonata Software
|
Bengaluru
19 Sep
Sonata Software
Bengaluru
Job Summary:
Key Responsibilities
- Build and extend PDF text & layout extraction using PyMuPDF for text-based PDFs, and Claude Vision for scanned, multi-column, or complex pages, producing bounding-box-anchored regions.
- Generate multilingual/cross-lingual embeddings (via Bedrock) to shortlist candidate PDF regions matching each English instrument item.
- Integrate a vision-capable Claude model on Amazon Bedrock as a verifier to confirm the correct source region for each item, with confidence scoring and a small-model fallback chain; ensure unmatched or low-confidence items are flagged rather than guessed.
- Build deterministic placement logic (AWS Lambda) that clones the English skeleton, injects matched text, and validates structural integrity — string IDs, ordering, and formatting tags.
- Design and build the agent orchestration layer using AWS Agent Core / MCP Server, connecting extraction, retrieval, verification, and placement as distinct tools with transparent interfaces.
- Implement the extract → retrieve → verify → place → validate workflow and guardrails using AWS Strands (Agentic SDK), ensuring structure-locked output, zero invented text, and full source traceability.
- Support the linguist review & approval tool, showing each output string alongside its PDF source region, and implement a per-string audit trail written to Amazon S3 / DynamoDB.
- Integrate with the existing eCOA ingest / screenshot-review pipeline: accept English instrument JSON and translated PDFs as inputs, and emit a draft JSON in the required target schema.
- Run validation across sample instruments and priority locales (including at least one non-Latin-script locale), and track/report placement accuracy and flag-rate metrics (target ≥90% correct placement, zero invented translations).
- Participate in daily scrum and weekly governance calls with Sonata and Medidata stakeholders; track progress in Jira.
Required Skills & Experience
- Hands-on experience with AWS Bedrock — invoking foundation models (Claude), including vision-capable models for document/image understanding.
- Experience building agentic AI systems using AWS Agent Core / MCP Server (Model Context Protocol) and AWS Strands, or comparable agent orchestration frameworks (e.g., LangGraph, LangChain agents).
- Strong Python development skills, including experience with PDF parsing libraries (PyMuPDF, pdfplumber, or similar).
- Experience with embeddings-based retrieval / semantic search, including cross-lingual or multilingual embeddings.
- Working experience with AWS Lambda, Amazon S3, DynamoDB, and CloudWatch.
- Solid understanding of prompt engineering and LLM verification/guardrail patterns — confidence thresholds, hallucination avoidance, fallback chains.
- Comfortable working with structured JSON schemas and schema validation.
- Experience delivering in Agile/Scrum teams using Jira.
Preferred / Nice to Have
- Experience in healthcare, life sciences, or clinical trials technology (eCOA, ePRO, EDC systems).
- Exposure to localization/translation-adjacent workflows or NLP/linguistics-adjacent projects.
- Experience with OCR or vision-based document understanding at scale.
- Prior experience integrating AI/automation solutions with third-party enterprise platforms via APIs.
📌 Generative AI Developer (Bengaluru)
🏢 Sonata Software
📍 Bengaluru