Company: Shuru Tech
Website: https://shurutech.com/
Experience: 6+ Years
Location: Remote
Employment Type: Full-time
Role Overview
Shuru Tech is looking for a robust Senior Java Engineer / Java Lead with 6+ years of backend engineering experience and hands-on expertise in building document parsing and document processing systems.
This is a fully remote role where you will work on scalable backend systems responsible for ingesting, parsing, extracting, transforming, and processing data from multiple document formats including PDF, DOC/DOCX, XLS/XLSX, PPT/PPTX, HTML, XML, CSV, and scanned documents.
Key Responsibilities
- Design and develop scalable backend applications and microservices using Java and Spring Boot.
- Build and maintain document ingestion, parsing, extraction, and transformation pipelines.
- Extract structured and unstructured information including text, tables, metadata, images, and document properties.
- Handle formats including PDF, Word, Excel, PowerPoint, HTML, XML, and scanned documents.
- Work with document-processing libraries such as Apache Tika, Apache PDFBox, Apache POI, iText, or equivalent technologies.
- Integrate OCR solutions for scanned PDFs and image-based documents.
- Design systems capable of efficiently processing large documents and high volumes of files.
- Build asynchronous and event-driven workflows using Kafka, RabbitMQ, or similar messaging technologies.
- Develop REST APIs for document upload, processing, extraction, and retrieval.
- Handle corrupted files, unsupported formats, parsing failures, and partial processing scenarios.
- Optimize performance, concurrency, memory utilization, and processing latency.
- Write clean,
maintainable, well-tested production-grade code.
- Participate in architecture discussions, code reviews, debugging, and production issue resolution.
- For Lead-level candidates, mentor engineers and drive technical design and engineering best practices.
Required Skills
- 6+ years of backend/software engineering experience.
- Strong hands-on experience with Java.
- Strong experience with Spring Boot, REST APIs, and microservices.
- Hands-on experience with document parsing, document ingestion, content extraction, or document processing systems.
- Experience with libraries such as Apache Tika, Apache PDFBox, Apache POI, iText, or similar.
- Strong understanding of multithreading, concurrency, asynchronous processing, and JVM performance.
- Experience with relational and/or NoSQL databases.
- Strong understanding of data structures, system design, and API design.
- Experience with Docker, CI/CD, Git, and cloud environments.
- Strong debugging and problem-solving skills.
- Comfortable working effectively in a remote engineering environment.
Good to Have
- Experience with Tesseract, AWS Textract, Azure Document Intelligence, Google Document AI, or similar OCR/document intelligence platforms.
- Experience extracting tables, layouts, forms, key-value pairs, and structured data from complex documents.
- Experience with LLM/AI-powered document extraction or Intelligent Document Processing (IDP).
- Experience with Elasticsearch/OpenSearch.
- Experience with AWS, Azure, or GCP.
- Experience building distributed systems capable of processing large volumes of documents.
- Experience working in a startup or high-ownership engineering environment.
📌 Lead Java Engineer (India)
🏢 Shuru
📍 India