← All posts / Research

4.5 Billion TikTok Records, 289GB, Three Weeks: The Private-API Scrape That Just Landed on Hugging Face

An anonymous researcher scraped 4.5 billion TikTok video records through the app's private Android API in three weeks and published all 289GB on Hugging Face — no login, no account, just forged devices and reverse-engineered signatures.

4.5 Billion TikTok Records, 289GB, Three Weeks: The Private-API Scrape That Just Landed on Hugging Face

On September 2–3, 2026, one of the largest social-media datasets ever assembled appeared on Hugging Face with no press release and no institutional backing. The repository kuben-developer/tiktok-videos-4b contains 4.5 billion TikTok video records — captions, view counts, likes, comments, shares, saves, sound identifiers, country codes, and posting times — packed into 27 zstd-compressed Parquet files totaling roughly 289 GB. The author, an independent researcher, says the entire corpus was collected in about three weeks.

The scale deserves a moment of honesty: this is metadata, not video. At 289 GB for 4.5 billion rows, each record costs about 60 bytes — enough for an ID, a caption, and a column of engagement counters, and nothing more. The videos themselves are not in the dataset. What is in it is the engagement skeleton of a large fraction of the platform, one row per content_id, deduplicated from a raw collection that ran to 5.94 billion video records, 3.23 billion creator profiles, and 2.8 billion comments.

How it was done: the app’s own API, not the web

The collection method is the technically interesting part. Almost every public TikTok scraper drives a headless browser or hits the web endpoints — slow, fragile, and missing most fields. This system instead spoke directly to the private HTTP+JSON API that TikTok’s Android app itself uses, the same interface com.zhiliaoapp.musically hits on every scroll. It is faster than the web tier and returns considerably more.

Getting in required four unrelated things to be right simultaneously:

  1. A device credential TikTok issued — device_id and iid values are granted by TikTok’s device-registration endpoint in exchange for a plausible handset profile, not invented by the client.
  2. A valid request signature — the signing stack spans X-Argus, X-Ladon (Speck-128/256), the legacy X-Gorgon digest, and X-Khronos, built on Simon, Speck, and SM3 primitives, plus TikTok’s own TTEncrypt body cipher. A single creator-timeline request carries 38 device-identity parameters, in a fixed order, because the signature hashes the query string literally.
  3. The correct regional host — TikTok partitions traffic across regional API clusters, and talking to the wrong one gets you nothing.
  4. A TLS handshake that looks like a phone — a standard Go or Python HTTP client presents a server-like JA3 fingerprint and is silently dropped.

The remarkable defensive detail is what failure looks like. Get any one of the four wrong and the API returns a clean HTTP 200 with a zero-byte body. No error, no status code, no challenge page. The researcher describes it as the single most important thing to understand about this API: an unsophisticated scraper will store empty JSON for hours while every dashboard stays green, because all four failure modes are indistinguishable from success. There is no feedback loop to debug against — a wrong rotation constant, a wrong byte order, a wrong host, and a wrong cipher suite all produce the same well-formed request and the same nothing.

Everything ran on anonymous device registrations. No login, no account, no session cookie exists anywhere in the pipeline — which also bounds the blast radius: genuinely account-gated content (DMs, private videos, per-user like graphs) was out of reach and stayed out of reach.

What researchers actually get

The dataset card is unusually candid about limitations, which is more than most corporate releases manage:

  • It’s a sample, not a census. The export covers 27 of 32 storage partitions, split on a hash of the creator ID — an unbiased random subset of what was collected, which was itself not all of TikTok.
  • Engagement counts are a snapshot, not a time series. Every number is whatever it was at collection moment, somewhere in a three-week window; comparing raw counts across distant posting dates without normalizing for age is a methodological error.
  • Rows are grouped by creator, not shuffled. Sequential reads produce highly correlated training batches. If you train on it, shuffle.
  • Creator identity is deliberately excluded — no author IDs, usernames, or profile data. Videos can be grouped by sound, mention, or caption, but not by who posted them.
  • Media URLs are excluded because TikTok’s CDN links carry signed expiry parameters and die within days.
  • country and language are TikTok’s own inferred labels, wrong often enough that the card warns against treating them as ground truth.

For anyone who wants to poke at it without a 289 GB download, the card leads with a DuckDB one-liner that queries the Parquet files in place, streaming from the Hub.

The card’s “responsible use” section is where the story sharpens. The author concedes that collection was contrary to TikTok’s Terms of Service, that the dataset is unaffiliated with TikTok or ByteDance, and — more consequentially — that captions written by real people constitute personal data under GDPR, the UK GDPR, and CCPA regardless of having been publicly posted. The obligations transfer to the downloader the moment the data lands on their machine. The card prohibits identification, profiling, targeting, or contacting individuals, and offers a removal path via the repository discussion board.

That framing is not paranoid. X and Meta have both won (and lost) scraping-related litigation, and platform ToS has repeatedly been treated as binding contract in civil suits. A 4.5-billion-row corpus of human-authored text, collected in explicit violation of those terms, now sits on the industry’s default open-model hosting platform — days after Nvidia closed its $12.93 billion acquisition of Hugging Face, making the chipmaker the steward of the platform on which the largest informal TikTok corpus now lives.

The Hacker News thread (48 points, 58 comments) split along predictable lines: gratitude for the data, irritation at the heavily LLM-flavored prose of the write-up, and a crisp summary from Simon Willison — “I scraped the metadata for 4.5 billion TikTok videos using the same API as their Android app… The video content itself is not included. I’ll sell you my Go scraping code.”

Because yes — there is a commercial engine behind the generosity. The full method writeup at seeksocial.io sells the complete Go implementation for $699 (signing stack with test vectors, device registration pipeline, all 24 endpoints with measured success rates) and offers a done-for-you version at $1,899. The free dataset is best understood as a demonstration of capability: proof that the machinery works at billions-of-rows scale, aimed at an audience of data vendors, hedge funds, and research groups.

Why it matters

Three overlapping storylines make this more than a weekend curiosity.

For AI training data, engagement signal at this scale is scarce. Recommender research, virality prediction, caption–success correlation, cross-lingual content analysis — all of it normally requires partnerships or years of collection. A free, deduplicated, well-documented 4.5-billion-row table changes the entry price from “corporate data deal” to “289 GB of bandwidth and a DuckDB binary.”

For platform defense, the empty-200 soft block is a genuinely well-designed anti-bot mechanism — and it still lost. When device registration, signature forging, and TLS mimicry can be solved by one motivated individual with a decompiler, the practical enforceability of API-boundary ToS against determined collectors looks increasingly theoretical.

For data governance, the dataset is a stress test arriving at an awkward moment: regulators in both the EU and US are actively litigating what obligations attach to “publicly available” personal data, and the industry’s central repository — now owned by Nvidia — is hosting the answer at 289 GB. Whether it survives a takedown request, a GDPR complaint, or quiet removal is a signal worth watching in its own right.

The dataset is live at this writing. The code is for sale. The corpus, once downloaded, is anyone’s to keep — which is precisely the problem, and precisely the point.