Hands-on workshop on cleaning and preparing high-quality datasets using Data Prep Kit. Topics include extracting content from PDFs and HTML, cleaning up markup, detecting and removing SPAM content, scoring and removing low-quality documents, identifying and removing PII data, and detecting and removing HAP (Hate Abuse Profanity) speech. More about Data Prep Kit: https://github.com/IBM/data-prep-kit
talk-data.com
Topic
1
tagged
Activity Trend
3
peak/qtr
2020-Q1
2026-Q1
Filtering by:
[AI Alliance] Workshop: Preparing High Quality Datasets with Data Prep Kit
×