talk-data.com talk-data.com

Filter by Source

Select conferences and events

People (320 results)

See all 320 →

Companies (1 result)

Michaels 1 speaker
Chief Information Security Officer
Showing 1 result

Activities & events

Title & Speakers Event
Michael Collado – Staff Software Engineer at Datakin

Data within today’s organizations has become increasingly distributed and heterogeneous. It can’t be contained within a single brain, a single team, or a single platform…but it still needs to be comprehensible, especially when something unexpected happens. Data lineage can help by tracing the relationships between datasets and providing a cohesive graph that places them in context. OpenLineage provides a standard for lineage collection that spans multiple platforms, including Apache Airflow and Apache Spark. In this session, Michael Collado from Datakin will show how to trace data lineage and useful operational metadata in Apache Spark and Airflow pipelines, and talk about how OpenLineage fits in the context of data pipeline operations and provides insight into the larger data ecosystem.

Airflow Spark
Airflow Summit 2022
Showing 1 result