23 Aug
|
Care Health Insurance
|
Gurugram
23 Aug
Care Health Insurance
Gurugram
AI/ML Engineer
(GenAI, OCR & ML)
Experience: 1-5 Years
Job Summary
We are looking for a Backend AI Engineer to design and scale GenAI- and CV-powered production systems. You will own end-to-end backend architecturefrom raw data in S3 through OCR and object detection, to high-concurrency FastAPI services that call LLMs and return reliable, structured outputs for business workflows.
Key Stack: FastAPI, Django (optional), AWS (S3/EC2/ECS), LLMs, Vector DBs, YOLO / CV, OCR, Redis/Celery
Key Responsibilities
- Production APIs: Build and maintain scalable async backends with FastAPI (high-throughput AI inference) and, where needed, Django for application logic and state.
- LLM systems: Integrate and optimize LLMs (API or self-hosted)prompt engineering, structured output parsing, fallbacks, and cost/latency trade-offs.
- RAG: Design Retrieval-Augmented Generation using vector DBs (Pinecone, Milvus, or pgvector) for context-aware responses over policies, tickets, and documents.
- OCR at scale: Architect OCR pipelines for multi-page PDFs and noisy scans; combine OCR with layout understanding and LLM post-processing for field extraction and validation.
- Computer vision (YOLO & beyond): Lead detection/classification models for document regions, seals, signatures, ID cards, damage/claims images, or quality gates; manage training data, evaluation metrics (mAP, precision/recall), and model versioning in S3.
- AWS & data orchestration: Deploy on EC2/ECS/Fargate; manage large datasets and weights in S3; build Boto3 pipelines for ingest, cleaning, and model-ready datasets.
- Performance: Optimize latency with Redis caching, Celery async jobs, batching, and efficient GPU/CPU inference paths.
- Production hardening: Monitoring, retries, idempotency, PII-aware handling, and clear SLAs for AI services.
Required Skills
- 1-5 years Python with FastAPI and/or Django; robust async programming and ORM/SQL optimization (PostgreSQL).
- Proven GenAI delivery: LangChain or LlamaIndex (or equivalent), embeddings, tokenization awareness, production LLM integration.
- Solid OCR experience in production (Textract, Tesseract, PaddleOCR, or similar) including accuracy/error analysis.
- Hands-on YOLO / CV detection or document AI (PyTorch/Ultralytics/OpenCV); ability to improve models with data, not only run pretrained weights.
- AWS: S3 (versioning/lifecycle), EC2/ECS, IAM; Docker for reproducible AI environments.
- Vector DB experience and Redis/Celery (or equivalent) for async workloads.
Nice-to-Have
- WebSockets / SSE for real-time AI streaming
- Model serving (TorchServe, Triton, vLLM, or similar)
- MLOps basics: experiment tracking, dataset versioning, CI for models
- Domain experience: insurance documents, claims, KYC, email/ticketing automate
📌 Ai Ml Engineer (Gurugram)
🏢 Care Health Insurance
📍 Gurugram