IBM

Foundations of AI Native Data Engineering

IBM

Foundations of AI Native Data Engineering

Antonio Cangiano
Ruslan Podgaets

Instructors: Antonio Cangiano

Included with Coursera PlusLearn more

Gain insight into a topic and learn the fundamentals.
Intermediate level

Recommended experience

2 weeks to complete
at 10 hours a week
Flexible schedule
Learn at your own pace
Gain insight into a topic and learn the fundamentals.
Intermediate level

Recommended experience

2 weeks to complete
at 10 hours a week
Flexible schedule
Learn at your own pace

What you'll learn

  • 1.Explain how AI-native data engineering differs from traditional pipeline design.

  • 2.Describe LLMs, embeddings, vector databases, and RAG from a data engineering perspective.

  • 3.Map AI workload requirements to data platform components and lifecycle responsibilities.

  • 4.Build a small semantic search or RAG-oriented workflow.

Details to know

Shareable certificate

Add to your LinkedIn profile

Recently updated!

July 2026

Assessments

32 assignments

Taught in English

See how employees at top companies are mastering in-demand skills

 logos of Petrobras, TATA, Danone, Capgemini, P&G and L'Oreal

Build your Machine Learning expertise

This course is part of the IBM AI-Native Data Engineering Professional Certificate
When you enroll in this course, you'll also be enrolled in this Professional Certificate.
  • Learn new concepts from industry experts
  • Gain a foundational understanding of a subject or tool
  • Develop job-relevant skills with hands-on projects
  • Earn a shareable career certificate from IBM

There are 9 modules in this course

Get oriented to Foundations of AI Native Data Engineering and learn how the course is structured, what to expect, and how to prepare for success. You’ll review the learning journey, assessments, final project arc, prerequisites, and responsible use of AI assistants before starting the technical modules.

What's included

1 video2 plugins

Learn how AI workloads are changing data engineering and why traditional BI-focused pipelines often need redesign to support semantic search, RAG, LLM applications, embeddings, and agentic consumers. You’ll compare ETL and AI-native pipelines, identify architecture and governance gaps, and begin a course project by defining an AI use case, mapping requirements to platform components, and clarifying ownership boundaries.

What's included

5 videos5 assignments2 app items5 plugins

This module introduces core AI concepts from a data engineering perspective, helping you distinguish AI, machine learning, deep learning, generative AI, and LLM-based systems. You will also learn how key AI stack components such as data preparation, embeddings, vector databases, retrieval systems, RAG, and serving interfaces work together in real applications.

What's included

4 videos5 assignments2 app items4 plugins

Learn how the AI data lifecycle extends data engineering beyond traditional analytics into feature engineering, embeddings, retrieval, inference, evaluation, monitoring, and governance. You’ll examine how AI data engineers design AI-ready pipelines, identify drift and leakage risks, define monitoring signals, and coordinate ownership across teams to support reliable, explainable, and governed AI systems.

What's included

4 videos5 assignments2 app items4 plugins

Learn how large language models shape data engineering decisions around prompting, context management, hosting, cost, latency, scaling, fine-tuning, and retrieval-augmented generation. You’ll focus on practical architecture, governance, and production tradeoffs rather than model research, and create project artifacts to support LLM integration decisions.

What's included

4 videos5 assignments2 app items4 plugins

Learn how retrieval-augmented generation (RAG) pipelines work and how data engineers design, support, and troubleshoot them in AI-native systems. You will examine the full RAG workflow—from source preparation and chunking to retrieval, prompt assembly, evaluation, governance, and failure diagnosis—and compare when RAG, fine-tuning, or both are the right choice.

What's included

4 videos5 assignments2 app items4 plugins

Apply the concepts from Course 1 to design or build a semantic search workflow using AI-native data architecture patterns. You will work with corpus preparation, embeddings, vector retrieval, metadata, governance, and optional RAG-style context assembly, then package your work into a final project-ready semantic search or RAG design.

What's included

4 videos5 assignments4 plugins

Wrap up Course 1 by connecting what you built to portfolio-ready evidence, professional practice, and the next step in the certificate. In this short bookend module, you’ll review your final project artifacts, reinforce the AI-native data engineering mindset, and create a practical plan for refining your work and preparing for Course 2.

What's included

1 video1 plugin

What's included

2 assignments1 plugin

Earn a career certificate

Add this credential to your LinkedIn profile, resume, or CV. Share it on social media and in your performance review.

Instructors

Antonio Cangiano
IBM
10 Courses750,442 learners
Ruslan Podgaets
IBM
0 Courses0 learners

Offered by

IBM

Why people choose Coursera for their career

Felipe M.

Learner since 2018
"To be able to take courses at my own pace and rhythm has been an amazing experience. I can learn whenever it fits my schedule and mood."

Jennifer J.

Learner since 2020
"I directly applied the concepts and skills I learned from my courses to an exciting new project at work."

Larry W.

Learner since 2021
"When I need courses on topics that my university doesn't offer, Coursera is one of the best places to go."

Chaitanya A.

"Learning isn't just about being better at your job: it's so much more than that. Coursera allows me to learn without limits."

Frequently asked questions