tiago.zip
these datasets may contain personal or sensitive information. by downloading or using them, you agree to the dataset license. access can be revoked at any time, and all data is provided as-is. we are not responsible for any consequences arising from its use.
collection of over 5 billion tweets and 3B profiles actively being scraped from Twitter since 2025, for twitter.cat. scrape ongoing
2.30+ TB
millions of posts and profiles from across mastodon and the fediverse, deduplicated across instances and exported monthly.
50M+ posts
hundreds of millions of posts, profiles and follows scraped from bluesky's firehose since 2026. scrape ongoing
200M+ posts
collection of over 15M gifs and stickers from tenor in video and database form, harvested after the 2026 api shutdown. scrape ongoing
8+ TB
scraped typeahead twitter data during mid 2025, includes 186k topics and 78M users
40 GB
over 14M linktree profiles, including links, bio, and social media pages
176 GB
data dumps from manifold markets's internal Supabase API, a prediction markets platform, in the form of JSONL files.
5.0 GB