06 Sep
|
Arminus
|
Hyderabad
- Experience: 4–6 years (Mid-level)
- Location: Hyderabad (Onsite on Wednesdays and Thursdays)
- Required Skills: Python, hands-on AWS (EC2, Lambda, and mandatory EMR), and EMR performance and cost optimization.
- Budget: We can go a little higher on the budget for this role.
Interview Process
- L1: Technical Engineer (Virtual)
- L2: SME (Virtual)
- L3: Engineering Manager (In-person)
Overview
Location: Hybrid · Hyderabad
Employment Type: full time
Experience: 5 - 6 years
Compensation: INR 1,153,440 - 1,441,800
Experience Required
• Extracted: 5+ years of hands-on experience developing enterprise-scale applications, data platforms, or distributed systems. 5+ years of experience developing, and operating Big Data platforms and cloud-based infrastructure services, preferably using AWS EMR and the Hadoop ecosystem.
Overview
The role focuses on building scalable data engineering and backend services for large-scale batch Big Data applications. You will develop and maintain data transformation pipelines using Spark/SQL/Hive and build cloud-native solutions on AWS. The position also involves workflow orchestration with Apache Airflow, production troubleshooting, and collaboration in an Agile/Scrum environment.
Key Responsibilities
• Develop API-driven systems to govern, manage, and monitor large-scale batch Big Data applications
- Build scalable backend services and data engineering solutions supporting data processing and operational workflows
- Develop and maintain data transformation processes using Spark, SQL, Hive, Python, Scala, and related technologies
- Build cloud-native data and backend solutions using AWS services (S3, EC2, EMR, Lambda, DynamoDB, API Gateway)
- Build and enhance workflow orchestration using Apache Airflow (advanced DAG design, dependency management, scheduling, monitoring,
failure handling)
- Define technical scope/objectives and implementation approaches via requirements gathering, technical research, and process definition
- Participate in architecture reviews, code reviews, performance tuning, and operational readiness activities
- Contribute to test planning and validation for integrations, functional areas, and deliverables
- Partner with cross-functional teams (product owners, data engineers, backend engineers, QA, DevOps) in an Agile/Scrum setting
- Troubleshoot production issues, improve reliability, and drive automation across data and backend workflows
Required Qualifications
• 5+ years of hands-on experience developing enterprise-scale applications, data platforms, or distributed systems
- 5+ years of experience developing and operating Big Data platforms and cloud-based infrastructure services (preferably AWS EMR and Hadoop ecosystem)
- Hands-on experience with AWS, Spark, Python and/or Scala, Airflow, SQL, Hive, and related data processing technologies
- Proficiency in either Python or Scala to build production-grade data pipelines and backend services
- Experience with AWS services: S3, EC2, EMR, Lambda, DynamoDB, API Gateway
- Strong experience with Apache Airflow orchestration (complex workflows, scheduling, dependency handling, monitoring, operational support)
- Strong understanding of data engineering concepts (data pipelines, data transformation,
batch processing, real-time processing, data quality, performance optimization)
- Excellent analytical and problem-solving skills
- Strong communication and presentation skills for technical and non-technical audiences
- Experience with version control and CI/CD tools including Git and Jenkins
- Ability to work effectively in cross-functional Agile/Scrum teams
Technical Skills
• AWS
- Apache Spark
- Python
- Scala
- Apache Airflow
- SQL
- Hive
- AWS EMR
- Hadoop ecosystem
- Git
- Jenkins
- CI/CD
- Data pipelines
- Data transformation
- Batch processing
- Real-time processing
- Data quality
- Performance optimization
Nice-to-have
• LLMs
- Generative AI
- Agentic AI
- AI-assisted engineering workflows
- API design
- Microservices
- Event-driven architecture
- Serverless backend development
- Infrastructure as code
- Automated testing
- Data governance
- Metadata management
- Lineage
- Data observability
- Operational controls
Tools/Platforms
• AWS S3
- AWS EC2
- AWS EMR
- AWS Lambda
- Amazon DynamoDB
- Amazon API Gateway
- Apache Airflow
- Apache Spark
- Hive
- Git
- Jenkins
Preferred Qualifications
• Bachelor's degree in Computer Science, Engineering, Mathematics, Information Systems, or a related technical field, or equivalent practical experience
- Experience with LLMs, Generative AI, Agentic AI, or AI-assisted engineering workflows
- Experience with API design, microservices, event-driven architecture, or serverless backend development
- Experience with CI/CD pipelines, infrastructure as code, automated testing, and production deployment practices
- Experience with data governance, metadata management, lineage, data observability, or operational controls for enterprise data platforms
📌 Mid-level Data Engineer - 5015_4-6yrs_Hyderabad (Hyderabad)
🏢 Arminus
📍 Hyderabad