17 Sep
|
PVKL Tech Services
|
Pune
17 Sep
PVKL Tech Services
Pune
Role summary
You will work on Provakils backend and data infrastructure building the systems that retrieve, process, store, and serve large volumes of data that power the platform. This is not a frontend or full-stack role: your work lives behind the API layer, in the data pipelines, the database schemas, and the processing services that keep the platform running accurately at scale.
You’ll write Python every day — building data retrieval services, processing pipelines, extraction scripts, and backend APIs. You’ll work in a Linux environment, interact with production databases, and be expected to own the solutions you build. You’ll work under the guidance of senior engineers who will help you ramp up, review your code, and push you to write software that’s reliable, efficient, and maintainable.
What you’ll do
Data retrieval & processing
- Build and maintain data retrieval services that pull structured and unstructured data from multiple sources into the platform
- Write data processing scripts and pipelines in Python that clean, transform, validate, and load data reliably and at scale
- Develop metadata extraction logic to identify and capture key data points from incoming documents and records
- Monitor data pipeline health: track processing volumes, identify failures, investigate data quality issues, and fix them
Backend development
- Write and maintain backend APIs in Python that serve processed data to the application layer
- Design and work with database schemas — understand how data is structured, indexed, and queried across MySQL, MongoDB, or PostgreSQL
- Optimise existing queries and processing jobs for performance as data volumes grow
- Debug and fix backend issues: read error logs, trace through data flows, identify root causes, and ship fixes
Quality & ownership
- Write clean, readable, and well-documented Python code that follows the team’s coding standards
- Adheres to the SOPs established by the team to ensure optimal operational efficiency.
- Write tests for your data pipelines and backend services — data processing code must be tested before it runs in production
- Participate in code reviews: submit your code for review and provide feedback on peers’ pull requests
- Own the solutions you build — if a pipeline breaks or data quality drops,
you investigate and fix it, not just report it
Collaboration
- Work with your team lead and senior engineers to understand requirements, break down tasks, and plan your work
- Participate in sprint planning, daily stand-ups, and retrospectives — communicate progress, blockers, and dependencies clearly
- Coordinate with other engineering teams when your data services feed into their modules or vice versa
- Work with product and business stakeholders to understand what data is needed, in what format, and by when
What success looks like in your first 60 days
First 30 days
- Complete onboarding and set up your development environment — you should be comfortable working in the Linux environment, accessing repositories, and running code locally
- Understand the data architecture: where data comes from, how it flows through the system, where it’s stored, and how it’s served to the application
- Ship your first bug fix or small improvement to an existing data pipeline or backend service
- Get comfortable with the team’s development workflow: Git, PR process, CI/CD, and deployment procedures
- Familiarise yourself with the database layer: schema structure, common queries, and how data is indexed
Days 31–60
- Pick up and deliver small-to-medium data engineering tasks with decreasing guidance — new extraction scripts, pipeline enhancements, or schema updates
- Write code that passes review with fewer iterations — your code quality and Python proficiency should visibly improve
- Independently debug and resolve at least one data quality or pipeline issue
- Start contributing meaningful feedback in code reviews for your peers
Must-have skills
- 0–2 years of experience in software development with a focus on Python (internships and strong personal projects count)
- Hands-on proficiency in Python — you can write scripts, build services, handle file I/O, work with APIs,
and process data programmatically
- Working experience in a GNU/Linux environment — you’re comfortable with the command line, shell scripting basics, file systems, and process management
- Experience with at least one modern database: MySQL, MongoDB, or PostgreSQL — you can write queries, understand schema design basics, and work with data
- Hands-on experience with data extraction and processing — you’ve worked with raw data, parsed it, cleaned it, and loaded it into a structured format
- Familiarity with Git for version control — branching, committing, pull requests, and resolving merge conflicts
- Problem-solving mindset — you enjoy debugging, tracing data through systems, and figuring out why something isn’t working
- Good communication skills — you can explain what you’re building, ask clear questions, and document your work
Good-to-have skills
- Experience with data pipeline tools
- Familiarity with web scraping, document parsing, or text extraction techniques
- Understanding of REST API design and development
- Experience with data quality monitoring or validation frameworks
- Bachelor’s degree in Computer Science, IT, or a related field
What you’ll learn in this role
- How to build and maintain data infrastructure at scale for an enterprise SaaS platform — real production systems processing high-volume data, not toy projects
- Backend and data engineering depth in Python: pipelines, extraction, schema design, API development, and performance optimisation — a robust foundation for a career in data or backend engineering
- How to work with complex, messy, real-world data sources and turn them into clean, reliable, structured data that the rest of the platform depends on
About Provakil Provakil is one of India’s leading legaltech platforms, helping enterprises manage litigation, contracts, compliance, and intellectual property through a single integrated product. We serve 300+ enterprise clients including Fortune 500 companies.
We’re a product-first company with a strong engineering culture. The data infrastructure team works with massive volumes of structured and unstructured data, building the systems that make the platform intelligent, accurate, and fast.
📌 Walk-in || Python Software Developer (Immediate Joiners) (Pune)
🏢 PVKL Tech Services
📍 Pune