Company: Shuru Tech
Website: https://shurutech.com/
Experience: 6+ Years
Location: Remote
Employment Type: Full-time
Role Overview
Shuru Tech is looking for a strong Senior Java Engineer / Java Lead with 6+ years of backend engineering experience and hands-on expertise in building document parsing and document processing systems .
This is a fully remote role where you will work on scalable backend systems responsible for ingesting, parsing, extracting, transforming, and processing data from multiple document formats including PDF, DOC/DOCX, XLS/XLSX, PPT/PPTX, HTML, XML, CSV, and scanned documents .
Key Responsibilities
- Design and develop scalable backend applications and microservices using Java and Spring Boot .
- Build and maintain document ingestion, parsing, extraction, and transformation pipelines .
- Extract structured and unstructured information including text, tables, metadata, images, and document properties .
- Handle formats including PDF, Word, Excel, PowerPoint, HTML, XML, and scanned documents .
- Work with document-processing libraries such as Apache Tika, Apache PDFBox, Apache POI, iText , or equivalent technologies.
- Integrate OCR solutions for scanned PDFs and image-based documents.
- Design systems capable of efficiently processing large documents and high volumes of files .
- Build asynchronous and event-driven workflows using Kafka, RabbitMQ, or similar messaging technologies .
- Develop REST APIs for document upload, processing, extraction, and retrieval.
- Handle corrupted files, unsupported formats, parsing failures, and partial processing scenarios.
- Optimize performance, concurrency, memory utilization, and processing latency .
- Write clean,
maintainable, well-tested production-grade code.
- Participate in architecture discussions, code reviews, debugging, and production issue resolution.
- For Lead-level candidates, mentor engineers and drive technical design and engineering best practices.
Required Skills
- 6+ years of backend/software engineering experience .
- Strong hands-on experience with Java .
- Strong experience with Spring Boot, REST APIs, and microservices .
- Hands-on experience with document parsing, document ingestion, content extraction, or document processing systems .
- Experience with libraries such as Apache Tika, Apache PDFBox, Apache POI, iText , or similar.
- Robust understanding of multithreading, concurrency, asynchronous processing, and JVM performance .
- Experience with relational and/or NoSQL databases.
- Strong understanding of data structures, system design, and API design .
- Experience with Docker, CI/CD, Git, and cloud environments .
- Strong debugging and problem-solving skills.
- Comfortable working effectively in a remote engineering environment .
Good to Have
- Experience with Tesseract, AWS Textract, Azure Document Intelligence, Google Document AI , or similar OCR/document intelligence platforms.
- Experience extracting tables, layouts, forms, key-value pairs, and structured data from complex documents.
- Experience with LLM/AI-powered document extraction or Intelligent Document Processing (IDP) .
- Experience with Elasticsearch/OpenSearch .
- Experience with AWS, Azure, or GCP.
- Experience building distributed systems capable of processing large volumes of documents.
- Experience working in a startup or high-ownership engineering environment.
📌 Lead Java Engineer (India)
🏢 Shuru
📍 India