15 Aug
|
Soothsayer Analytics
|
Hyderabad
15 Aug
Soothsayer Analytics
Hyderabad
We seek a Data Engineer (Mid-level) with 46 years of hands-on experience in designing, building, and optimizing data pipelines. You will work closely with AI/ML teams to ensure data availability, quality, and performance for analytics and GenAI use cases.Key Responsibilities
1. Data Pipeline Development
Build and maintain scalable ETL/ELT pipelines for structured and unstructured data.
Ingest data from diverse sources (APIs, streaming, batch systems).
2. Data Modeling & Warehousing
Design effective data models to support analytics and AI workloads.
Develop and optimize data warehouses/lakes using Redshift, BigQuery, Snowflake, or Delta Lake.
3. Big Data & Streaming
Work with distributed systems like Apache Spark, Kafka, or Flink for real-time/large-scale data processing.
Manage feature stores for ML pipelines.
4. Collaboration & Best Practices
Work closely with Data Scientists and ML Engineers to ensure high-quality training data.
Implement data quality checks, observability, and governance frameworks.
o Programming: Python/Scala/Java (Python preferred).
o Big Data & Processing: Apache Spark, Kafka, Hadoop.
o Databases: SQL/NoSQL (Postgres, MongoDB, Cassandra).
o Data Warehousing: Snowflake, Redshift, BigQuery, or similar.
o Orchestration: Airflow, Luigi, or similar.
o Cloud Platforms: AWS, Azure, or GCP (data services).
o Version Control & CI/CD: Git, Jenkins, GitHub Actions.
o MLOps/GenAI pipelines (feature engineering, embeddings, vector DBs).
📌 Data Engineer (Hyderabad)
🏢 Soothsayer Analytics
📍 Hyderabad