Training Data Provenance
Structures where the origin and use-scope attributes of AI training data flow downstream without independent verification — public-API scraping of chat platforms, dataset distribution outside the licensed scope, gaps in AI training-data audit layers.
7.6 petabytes of Hugging Face training data held 221,303 live secrets
detected and notified, never revoked (Truffle Security)
Over 40,000 unauthorized likeness and voice posts across major platforms
and a 100% takedown rate did not stop the same person's models from reappearing (JAPRO FY2025 survey)
Figma
AI content training defaulted on for individuals and small teams, and off for enterprise
Common Crawl: about 12,000 live credentials embedded in a public corpus used to train LLMs
training-data provenance not verified before ingestion (Truffle Security)
Bright Data SDK: your living-room TV became a relay node for AI-scraping
the origin and consent of collected data and relayed traffic not independently verified (Include Security)
Generated Until the Rightsholder Said No
The Consent-and-Provenance Gap Behind OpenAI Sora 2
12.8 Billion Training Images Contained Passports, Résumés, and Faces
The Provenance and Consent of Training Data Were Never Verified at Collection
Discord 2.05 Billion Message Scraping via Public API
How Public Channel Data Gets Redistributed as AI Training Datasets