Introduction At IBM Software, we transform client challenges into solutions. Building the world’s leading AI-powered, cloud-native products that shape the future of business and society.
Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You’ll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation.
With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM’s product and technology landscape. Here, you’ll have the tools and opportunities to advance your career while creating software that changes the world.
Your Role And Responsibilities About the Role
As a Data Engineer II of Confluent, an IBM company, Data team you will own meaningful slices of our data platform end-to-end -- from ingestion through transformation to the data products that power decision-making across the company. You will be part of Confluent R&D; team and work closely with various business stakeholder like Prod&Eng;, Marketing, FieldsOps, Sales etc to address their analytical needs. You would be working closely with the overall IBM CDO and CIO team to enable the wider organisation with data needs from the Confluent ecosystem.
You'll design and operate batch and real-time pipelines, partner closely with Data Scientists, Analysts, and business stakeholders, and raise the bar on craftsmanship, reliability, and developer productivity for the team.
You'll work in a modern, AI-augmented engineering workplace: shipping faster with Claude Code and other AI coding assistants, building AI-powered internal tools (natural-language-to-SQL, automated data quality, lineage, anomaly detection), and bringing GenAI thoughtfully into data workflows. We expect you to take significant projects from ambiguous problem statement to delivered outcome, mentor newer engineers, and represent the team confidently with cross-functional partners.
What You Will Do
- Design, build, and operate efficient,
reliable, and well-documented data pipelines across batch and streaming systems -- from source ingestion through the warehouse to consumer-facing data products.
- Own the data models, SLAs, and quality contracts for the domains you cover; treat documentation, testing, and observability as first-class deliverables.
- Improve the data platform itself: introduce or extend tooling, templates, CI/CD, testing frameworks, and monitoring that make the entire team faster and more reliable.
- Adopt and champion AI-assisted engineering -- use Claude Code and similar tools as part of your daily workflow, build internal AI-powered utilities (text-to-SQL on the semantic layer, automated PR review, log triage, data discovery), and evaluate where LLMs add durable value vs. hype.
- Partner directly with Data Scientists, Analysts, and business stakeholders to translate ambiguous requirements into well-scoped deliverables; communicate trade-offs, anticipate push-back, and drive alignment without needing escalation.
- Mentor interns, new hires, and L2 engineers -- through onboarding, code review, design feedback, and pairing.
- Contribute to cloud infrastructure and governance: IAM, service accounts, secrets management, cost optimization, and Terraform-managed GCP / Confluent Cloud resources.
- Drive incident response and root-cause analysis for the pipelines you own; close the loop with durable fixes, runbooks, and prevention.
Preferred Education Master's Degree Required Technical And Professional Expertise 6 to 9 years of experience in Data Engineering at a technology company, including production ownership of non-trivial pipelines and data models.
- Strong command of advanced SQL and Python -- you can write production-grade code,
review others' code, and pick the right tool for the job.
- Have end to end ownership of Data pipeline and a clear grasp on requirement gathering, data modeling, data quality and access management
- Solid grounding in dimensional modeling, data warehousing, and ELT/ETL fundamentals; you can design a clean star schema and explain the trade-offs.
- Hands-on experience operating orchestration frameworks in production (Airflow / Cloud Composer, Dagster, or similar).
- Hands-on experience building streaming and event-driven pipelines -- Kafka, Flink, or equivalent -- and reasoning about exactly-once, schema evolution, and back-pressure.
- Experience on a major cloud platform (GCP preferred) and with infrastructure-as-code (Terraform).
- Demonstrated ability to operate with moderate ambiguity: scope your own work, identify what to clarify, recommend an approach, get buy-in, and ship.
- Fluency with AI coding assistants (Claude Code, Cursor, Copilot) as a daily tool -- and judgment about where they help vs. where they hurt.
- Clear written and verbal communication; comfortable presenting designs, post-mortems, and trade-offs to technical and non-technical audiences.
Preferred Technical And Professional Experience Experience with open table formats (Apache Iceberg, Delta Lake) and modern lakehouse patterns.
- Familiarity with data quality / observability tooling (Great Expectations, Soda, Elementary, Monte Carlo) and catalog / lineage systems (DataHub, Atlan, Unity Catalog).
- Experience building or integrating with AI/ML infrastructure: feature stores (Feast, Tecton), vector databases (pgvector, Pinecone, Weaviate), or RAG pipelines for internal knowledge systems.
- Experience contributing to or operating a semantic layer that powers self-serve analytics and natural-language interfaces.
- Familiarity with Confluent Cloud, Kafka Connect, and Flink SQL.
- Open-source contributions or technical writing in the data engineering space.
Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience.
📌 Data Scientist , Engineering - Confluent (Bengaluru)
🏢 IBM
📍 Bengaluru