← All posts / Meta

The Cookie That Crosses Sites: Inside OpenAI's __obi Ad Pixel and What It Knows About You

A reverse-engineering deep dive reveals how OpenAI's ad-measurement pixel at bzr.openai.com mints a JWT-bound __obi cookie that silently links your browsing on advertiser sites to your ChatGPT account — scraping hashed emails, phones, and location data along the way.

The Cookie That Crosses Sites: Inside OpenAI's __obi Ad Pixel and What It Knows About You

When OpenAI launched advertising inside ChatGPT earlier this year, the pitch to marketers was straightforward: a new kind of inventory where people are actively thinking, comparing, and deciding. The pitch to users was equally simple — the ads would respect your privacy, and OpenAI’s developer documentation still describes its Measurement Pixel as using “a privacy-preserving identifier.” This week, an independent reverse-engineering investigation pulled that claim apart, thread by thread, and what emerged is one of the most consequential privacy stories of the year: a tracking mechanism that silently links what you do on ordinary websites to your ChatGPT account.

What the researcher found

On September 20, an independent threat-intelligence researcher publishing under the Buchodi banner released a detailed technical write-up of OpenAI’s ad infrastructure. The subject is an endpoint called bzr.openai.com — “bzr” standing for “bazaar,” OpenAI’s internal name for its ads platform — and the cookie it mints: __obi.

The mechanism works in three steps, each individually mundane, together powerful.

Step one: ChatGPT creates an identifier and signs it. While you are logged into chatgpt.com, the client generates 16 random bytes and calls a backend endpoint (/backend-api/bazaar/obi/sync-token). The server returns an RS256-signed JWT whose payload binds together your account subject (a 64-hex identifier), a 22-character obi identifier, and a consent decision. The token expires in 60 seconds.

Step two: the identifier becomes a cross-site cookie. The client POSTs that JWT to bzr.openai.com/v1/obi/sync, and the response sets:

Set-Cookie: __obi=…; Domain=.openai.com; HttpOnly;
Max-Age=31536000; Path=/; SameSite=none; Secure

That SameSite=None is the critical detail. It is the one configuration that allows a cookie to ride along on cross-site requests. Every other OpenAI cookie observed in the study was blocked by the browser at the advertiser’s page — oai-did and oaicom-stable-id are SameSite=Lax, session cookies are domain-scoped. __obi is the only OpenAI identifier configured to cross sites. And Max-Age=31536000 means it lives for a full year.

Step three: advertiser sites send it back. Any company that buys ads on ChatGPT can install OpenAI’s Measurement Pixel — a small JavaScript tag loaded from bzrcdn.openai.com/sdk/oaiq.min.js, exactly the way retailers have installed Meta and Google tags for a decade. On the researcher’s own phone, one __obi value was transmitted to OpenAI from twelve commercial websites under thirteen distinct pixel IDs, including Chewy, Wayfair, ThriftBooks, Eventbrite, HelloFresh, Coursera, and SeatGeek. Every request was accepted with a 202 status.

The pixel does not merely relay the identifier. The SDK also harvests identity material from the advertiser’s page, sorting it into four channels that OpenAI’s own payloads label: in for values the advertiser deliberately passes, and fm, ht, js for values scraped from form fields, rendered page text, and the tag-manager bus.

The numbers are stark. In the observed traffic, scraped identity outnumbered advertiser-supplied identity 685 events to 255. The tag-manager bus was the single largest source of email addresses: the SDK replaces window.dataLayer.push with its own function, reads adobeDataLayer, and locates renamed Google Tag Manager layers by parsing the l= parameter off the gtm.js script tag.

Email addresses, phone numbers, and names are SHA-256 hashed before transmission — a nod to privacy, though hash-based identifiers are trivially reversible for anyone holding a dictionary of emails. Country, region, city, and postal code are sent in the clear. Postal code was the most-harvested form field: 100 events across 28 sites.

One subtle detail deserves emphasis: even the act of loading the SDK discloses the identifier. The pixel’s code contains a path that omits credentials, but it doesn’t help — the browser attaches cookies to the <script src> request that loads the SDK before any of OpenAI’s code runs. By virtue of embedding the tag, the site has already disclosed who is browsing.

The scale of the researcher’s dataset lends weight to the findings: several months of observed traffic covering 936 distinct advertiser pixels across 1,029 hostnames, with 932 sync tokens decoded. Of those, 736 carried an account_user subject and 196 carried an anonymous subject — and the anonymous identifier is just as persistent, one per device, surviving at least 27 days even when you are logged out.

Here is where the story moves from engineering to policy. OpenAI’s cookie policy lists __obi under Analytics cookies, with a one-year lifetime — the only entry in that section. OpenAI runs analytics and marketing as two separate consent choices (oai_consent_analytics and oai_consent_marketing), and every sync token the researcher decoded carried consent_decision: analytics_allowed.

In other words: a user who grants analytics consent but explicitly refuses marketing consent still receives an advertising identifier that follows them across the web and resolves back to their ChatGPT account. The researcher sent two questions to OpenAI’s press and privacy addresses on September 14 — why __obi is classified as an analytics cookie, and whether an analytics-only user still gets it. OpenAI Support acknowledged the inquiry, said it would be shared internally for review, and did not answer either question.

The path data that survives is not always benign either. URLs are reduced to origin plus path before sending, but paths reaching the collector included a medical condition, a debt-solutions funnel, and a litigation intake form. Automatic matching was enabled for 638 of 881 pixels with a known setting — including every credit and lending advertiser observed. A denylist does exclude passwords, one-time codes, card numbers, SSNs, dates of birth, medical history, diagnosis, and court fields, which suggests OpenAI anticipated the sensitivity — and shipped it anyway.

Standard adtech, unprecedented product

The researcher is careful about what the evidence does and does not show: the account-level join happens server-side by design, but was not directly observed. And the honest counterpoint is included: Meta built the structural equivalent years ago — a logged-in account, third-party cookies on pixel fires, off-site conversions resolved to a profile. The mechanism is textbook adtech.

What has no precedent is running it on an AI chat product. People tell these products things they would never put on a social network — health symptoms, legal troubles, career doubts, relationship conflicts — and these products increasingly act on their behalf, browsing, booking, and purchasing. Wiring a year-long cross-site identifier into that context collapses the wall between “anonymous web browsing” and “conversation history” in a way no social platform ever could.

There are real technical limits. The mechanism was observed on Chrome for Android; Safari’s Intelligent Tracking Prevention blocks all third-party cookies, and every iOS browser runs on WebKit, so the mechanism does not operate on iOS at all. Roughly one ChatGPT session in five produced a sync token, and the mobile web client serves ads without syncing. Chrome’s own third-party-cookie deprecation saga — repeatedly delayed, now reportedly abandoned in favor of a user-choice prompt — is the reason this entire class of tracking still functions at scale.

Why it matters

Advertisers themselves are in the dark. __obi belongs to a domain their scripts cannot read; a merchant that installed a humble conversion pixel has no way to know its visitors are being resolved to ChatGPT identities. And the advertisers’ own cookie, __obref, is set on each advertiser’s domain individually — 2,828 of 2,860 observed values appeared under exactly one advertiser, meaning sites cannot cross-observe each other through it.

The disclosure landed on Hacker News with 150+ points within three hours and is already rippling through the privacy community. For OpenAI, the timing is delicate: the company is scaling an ads business to support a reported $100B+ revenue ambition while simultaneously asking the world to trust it with agent-style products that act autonomously on users’ behalf. A cookie that quietly follows ChatGPT users across the web — classified as “analytics,” enabled by granular consent defaults, and unanswered by the company when asked directly — is exactly the kind of trust leak that agent products cannot afford.

For users, the practical takeaways are limited but real: iOS browsers are unaffected; Safari and Firefox block the mechanism; and in Chrome, clearing .openai.com cookies or refusing analytics consent in ChatGPT’s data controls removes the identifier. For everyone else, the quiet 202 responses continue, one browsing session at a time.