NOT a furry tiago.zip

datasets

these datasets may contain personal or sensitive information. by downloading or using them, you agree to the dataset license. access can be revoked at any time, and all data is provided as-is. we are not responsible for any consequences arising from its use.

read the full license

twitter.cat

collection of over 5 billion tweets and 3B profiles actively being scraped from Twitter since 2025, for twitter.cat. scrape ongoing

2.30+ TB

fedi.parquet

millions of posts and profiles from across mastodon and the fediverse, deduplicated across instances and exported monthly.

50M+ posts

bsky.parquet

hundreds of millions of posts, profiles and follows scraped from bluesky's firehose since 2026. scrape ongoing

200M+ posts

tenor archive

collection of over 15M gifs and stickers from tenor in video and database form, harvested after the 2026 api shutdown. scrape ongoing

8+ TB

twitter-typeahead.db

scraped typeahead twitter data during mid 2025, includes 186k topics and 78M users

40 GB

linktree.db

over 14M linktree profiles, including links, bio, and social media pages

176 GB

manifold.zip

data dumps from manifold markets's internal Supabase API, a prediction markets platform, in the form of JSONL files.

5.0 GB