13 Aug
|
GravEiens Eduservices
|
India
13 Aug
GravEiens Eduservices
India
Looking for PDF Data Collection Partners — Bengali, Telugu, Odia, Gujarati, Malayalam
We're sourcing PDF datasets in the above languages for document-model (AI/ML) training, and are looking for vendors, publishers, and collection partners who can support this.
What we need:
1.PDFs with diverse layouts — single-column, double-column, three-column, and complex pages, with tables, charts, figures, formulas, and headers/footers
2.A mix of categories — books, newspapers, magazines, academic papers, textbooks, exam papers, business/financial reports, notes, and slides
3. Both digital and scanned PDFs, readable and complete pages
4.Diverse sources — different publishers, institutions, and websites
Significant: We cannot accept large volumes of near-duplicate layouts or repeated templates from a single source. Layout diversity is the key requirement.
Licensing is mandatory. The data must be legally shareable and usable for AI/ML training — public domain, openly licensed, directly licensed, or owned by your organisation.
When reaching out, please share: language, document category, source, collection method, licensing status, available volume, delivery timeline, and pricing.
Work Location: Remote
📌 Indic PDF Collector (India)
🏢 GravEiens Eduservices
📍 India