08 Oct
|
HCLTech
|
Hyderabad
Technical Specialist
Hyderabad, Telangana
Job Summary
AI, Machine Learning, Kafka, Big Query
- Design and develop batch and real-time data pipelines for analytics, AI, and machine learning use cases.
- Build Kafka producers, consumers, topics, connectors, and stream-processing workflows with appropriate partitioning, ordering, replay, retry, and dead-letter handling.
- Create and optimize BigQuery datasets, tables, views, materialized views, partitioning, clustering, and SQL workloads for performance and cost efficiency.
- Integrate structured and unstructured data from applications, APIs, databases, files, and event streams.
- Prepare curated, reusable datasets and feature pipelines for model training, validation, deployment, and monitoring.
- Implement data-quality checks, schema validation, lineage, metadata, retention, and access controls.
- Develop reliable orchestration, CI/CD, automated testing, observability, and incident-response practices for data workloads.
- Troubleshoot pipeline failures, data delays, duplication, schema changes, and performance bottlenecks.
- Collaborate with data scientists to operationalize machine learning workflows and ensure reproducible data inputs.
- Document data models, interfaces, operational procedures, and architectural decisions.
Mandatory Technical Skills
- Strong programming skills in Python and advanced SQL.
- Hands-on experience with Apache Kafka or a managed Kafka service in production environments.
- Strong expertise in Google BigQuery, including data modeling, query optimization, partitioning, clustering, and cost management.
- Practical knowledge of ETL/ELT patterns, dimensional modeling, data lakes, data warehouses,
and distributed processing.
- Experience building data pipelines for AI and machine learning workloads.
- Knowledge of machine learning lifecycle concepts, feature engineering, model inputs and outputs, and production monitoring.
- Experience with Google Cloud data services and cloud-native security practices.
- Proficiency with workflow orchestration, version control, automated testing, and CI/CD.
- Understanding of data governance, privacy, access management, encryption, and audit requirements.
- Strong diagnostic and performance-tuning skills across streaming and analytical workloads
Key Responsibilities
AI, Machine Learning, Kafka, Big Query
- Design and develop batch and real-time data pipelines for analytics, AI, and machine learning use cases.
- Build Kafka producers, consumers, topics, connectors, and stream-processing workflows with appropriate partitioning, ordering, replay, retry, and dead-letter handling.
- Create and optimize BigQuery datasets, tables, views, materialized views, partitioning, clustering, and SQL workloads for performance and cost efficiency.
- Integrate structured and unstructured data from applications, APIs, databases, files, and event streams.
- Prepare curated, reusable datasets and feature pipelines for model training, validation, deployment, and monitoring.
- Implement data-quality checks, schema validation, lineage, metadata, retention, and access controls.
- Develop reliable orchestration, CI/CD, automated testing, observability, and incident-response practices for data workloads.
- Troubleshoot pipeline failures, data delays, duplication, schema changes, and performance bottlenecks.
- Collaborate with data scientists to operationalize machine learning workflows and ensure reproducible data inputs.
- Document data models, interfaces, operational procedures, and architectural decisions.
Mandatory Technical Skills
- Strong programming skills in Python and advanced SQL.
- Hands-on experience with Apache Kafka or a managed Kafka service in production environments.
- Solid expertise in Google BigQuery, including data modeling, query optimization, partitioning, clustering, and cost management.
- Practical knowledge of ETL/ELT patterns, dimensional modeling, data lakes, data warehouses, and distributed processing.
- Experience building data pipelines for AI and machine learning workloads.
- Knowledge of machine learning lifecycle concepts, feature engineering, model inputs and outputs, and production monitoring.
- Experience with Google Cloud data services and cloud-native security practices.
- Proficiency with workflow orchestration, version control, automated testing, and CI/CD.
- Understanding of data governance, privacy, access management, encryption, and audit requirements.
- Strong diagnostic and performance-tuning skills across streaming and analytical workloads.
Preferred Skills
Skill Requirements
Other Requirements
📌 Technical Specialist (Hyderabad)
🏢 HCLTech
📍 Hyderabad