09 Sep
|
InfoVision
|
Pune
Job Description
Senior Data Engineer – Data Pipelines & Cloud Databases
n
Location: Pune | Experience: 4+ years | Type: Full-Time
n
Role Overview
n
Design, build, manage enterprise data pipelines on Azure and Databricks. Own schema design, API development, and data infrastructure for analytics and intelligence products.
n
Key Responsibilities
n
n
Data Pipeline Development
n
n
- Implement robust, scalable data pipelines using Microsoft Azure and Databricks stack
n
- Build reusable data pipeline components and frameworks
n
- Design and optimize data workflows for performance and reliability
n
- Monitor pipeline health and implement automated alerting
n
n
Database Architecture
n
n
- Design relational and non-relational database schemas
n
- MongoDB: schema design, aggregation pipelines, indexing, sharding, replica sets, performance tuning
n
- Azure Cosmos DB: multi-model access, partitioning, consistency levels, throughput management (RU/s), global distribution
n
- Azure Cosmos DB Gremlin API: graph data modeling, traversals, vertex/edge design, relationship analytics
n
n
API Development
n
n
- Build RESTful APIs using FastAPI with async endpoint design
n
- Implement Pydantic models, middleware, dependency injection, background tasks
n
- Generate and maintain auto-generated OpenAPI/Swagger documentation
n
- Configure Azure API Management (APIM) for API gateway and security
n
n
Project & Stakeholder Management
n
n
- Collect progress updates from squads regularly
n
- Consolidate updates into weekly project status reports
n
- Develop L3-level project plans with task details, milestones, dependencies
n
- Participate in early-stage design and feature definition
n
- Communicate complex data insights to non-technical stakeholders
n
n
Collaboration & Integration
n
n
- Work across multiple engineering teams on prototype integration
n
- Support integration of proven prototypes into core intelligence products
n
- Solid team collaboration and cross-functional communication
n
- Knowledge sharing and documentation
n
n
Required Experience:
n
Data Engineering & Databases
n
n
- 4+ years data engineering or data platform experience
n
- Relational and non-relational database expertise
n
- MongoDB: advanced schema design, aggregation, indexing, sharding, performance optimization
n
- Azure Cosmos DB: multi-model APIs, partitioning, consistency, throughput management
n
- Graph databases: Cosmos DB Gremlin API, graph modeling, traversals, relationship analytics
n
n
API & Backend Development
n
n
- FastAPI proficiency: async endpoints, Pydantic models, middleware, dependency injection
n
- Background tasks and job scheduling
n
- RESTful API design and best practices
n
- OpenAPI/Swagger documentation
n
n
Cloud Platform
n
n
- Azure fundamentals and hands-on experience
n
- Azure Data Factory or Databricks for ETL/ELT
n
- Azure Cosmos DB multi-region setup
n
- Azure API Management (APIM) configuration
n
n
Data Pipeline Skills
n
n
- ETL/ELT pipeline design and development
n
- Data quality validation and monitoring
n
- Schema design for analytics and reporting
n
- Performance optimization and scalability
n
n
n
Preferred Experience
n
n
- Databricks Delta Lake experience
n
- Azure Synapse Analytics
n
- Python for data engineering
n
- Spark SQL optimization
n
- Real-time data streaming
n
- Data governance and metadata management
n
- Agile/Scrum development model
n
- CI/CD pipeline experience
n
n
📌 Data Engineer (Pune)
🏢 InfoVision
📍 Pune