talk-data.com talk-data.com

Filter by Source

Select conferences and events

Showing 2 results

Activities & events

Title & Speakers Event
Git for Data 2025-10-15 · 18:00

Talk on distributed version control and how data projects can leverage Git and open formats like Apache Iceberg to enable multi-user data pipelines with snapshotting, time-travel, and branching.

Git apache iceberg
Git for Data: How Table Formats Unify Software and Data Development
Git for Data 2025-10-15 · 18:00

Distributed version control systems - such as Git - unlock software development in multi-player mode: devs can safely work over the same code base, with standard (albeit perhaps not user-friendly!) abstractions for snapshotting, time-travel, and branching. Data folks have rarely been so lucky, as their projects crucially depend on data, whose life-cycle management is often cumbersome and custom. In this talk, we present open formats - such as Apache Iceberg - to practitioners with limited exposure to modern cloud infrastructure. In particular, we show how moving from datasets to tables unlocks a similar multi-player mode when building data pipelines, with equivalent abstractions for snapshotting, time-travel, branching, and a unified backbone for pipelines, data science, and AI use cases.

Git apache iceberg
Git for Data: How Table Formats Unify Software and Data Development
Showing 2 results