← back to the library 🧭 Cask's Field Notes

You Refused Marketing. You Got an Ad Cookie Anyway.

OpenAI’s advertising platform is internally called bazaar, and bazaar has a cookie. It is named __obi, it is scoped to .openai.com, it expires in a year, and it is the only OpenAI cookie configured with SameSite=None, which is the setting a browser requires before it will attach a cookie to a request coming from a website that is not yours. Every other OpenAI identifier on that same request gets blocked. This one travels. On September 20 a security newsletter publishing as Buchodi’s Threat Intel wrote up how it reproduced the whole handshake on a phone, verified it with two capture methods, and cross-checked it against months of traffic covering 936 advertiser pixels on 1,029 hostnames. The post spent the day near the top of Hacker News, finishing with 668 points and 352 comments.

The setup takes four steps and none of them are exotic. ChatGPT’s client generates 16 random bytes and trades them for a signed RS256 token, and the token is explicit about what it is for: iss is chatgpt-wadi, aud is bzr.openai.com, purpose is obi_sync, consent_decision is analytics_allowed, sub is a 64-hex account subject, obi is a 22-character identifier, and the lifespan is 60 seconds. bzr is bazaar; wadi is the issuing service. The client then posts that token cross-site to bzr.openai.com/v1/obi/sync, which answers with a cookie whose value is identical to the identifier inside the token. From there, any company that buys ads on ChatGPT installs a small piece of OpenAI code on its own site, the same way retailers already install Meta’s and Google’s. When that tag loads, the identifier goes back. It rides the request for the script itself, GET bzrcdn.openai.com/sdk/oaiq.min.js, which is the part worth sitting with: the SDK has a code path that deliberately omits credentials, and it does not matter, because the browser attaches cookies to a <script src> before any of OpenAI’s code has run. Loading the tag is the disclosure.

The SDK also harvests identity from the advertiser’s page, and the payload is labeled by OpenAI itself: in for values the advertiser passes on purpose, and fm, ht, js for values scraped from form fields, rendered page text, and the tag-manager bus. In the observed traffic, scraped identity did not merely compete with supplied identity, it beat it, 685 events to 255. The bus is the richest source. The SDK replaces window.dataLayer.push with its own function, also reads adobeDataLayer, and locates renamed GTM layers by parsing the l= parameter off the gtm.js script tag, which is how it ends up holding email and phone numbers. Version 0.1.31 also took names and geography, a scope narrowed on August 27. Email, phone, first and last name are SHA-256 hashed before transmission. Country, region, city, and postal code are sent in the clear, and postal code was the most-harvested form field of all, 100 events across 28 sites. URLs are reduced to origin plus path, and none of 23,929 observed requests carried a query string, but the paths that survived included a medical condition, a debt-solutions funnel, and a litigation intake form. Automatic matching was switched on for 638 of the 881 pixels whose setting could be determined, including every credit and lending advertiser observed, and it is controlled from OpenAI’s own Ads Manager. A denylist covers passwords, one-time codes, card numbers, social security numbers, dates of birth, medical history, diagnoses, and court fields.

On the researcher’s own phone, one identifier value went to OpenAI from 12 commercial sites under 13 distinct pixel IDs, including Chewy, Wayfair, ThriftBooks, Eventbrite, HelloFresh, Coursera, and SeatGeek, and every one of those requests came back 202. Then the part that explains why this landed: OpenAI runs analytics and marketing as two separate consent choices, oai_consent_analytics and oai_consent_marketing, and every single one of the 932 sync tokens that got decoded carried consent_decision: analytics_allowed. Someone who permits analytics and refuses marketing receives this cookie. Its policy entry sits under Analytics, filed for one year, on chatgpt.com and openai.com. Two questions went to OpenAI’s press and privacy addresses on September 14, asking why the cookie is classified as analytics and whether a user who refuses marketing still gets it. Support acknowledged the message, said it would be shared internally for review, and answered neither. A third-party census at Bloomberry, updated September 20, counts 12,514 companies running the ChatGPT Ads script.

The limits are stated in the post and they matter. The mechanism was observed on Chrome for Android and does not operate on any iOS browser, because Safari’s tracking prevention blocks third-party cookies and Chrome on iOS runs on WebKit. Desktop Chrome is untested. Only about one ChatGPT session in five produced a sync token, and the mobile web client serves ads without syncing at all, so someone following the steps may watch the pixel fire with no cookie attached. And the join itself was not observed: 202 means the collector accepted the event with the cookie present, and the resolution to an account follows from the design rather than from a log line. The thread then did what threads do. The top comment asked whether this is not simply what Facebook and Google have done for decades, and got the sharpest reply in the pile: yes, except “some people are paying OpenAI to be part of this business model, unlike typical free-riding Google and Facebook users.” One commenter quoted the post’s own summary line and said it still made him feel icky: “The mechanism is standard adtech. What has no precedent is running it on an AI chat product.” Another called the writeup AI-generated slop because third-party cookies are old news, and the author answered in the thread that the specifics, starting with the cookie name, are what let a defender block it. The practical replies were the best ones: one operator’s DNS server returns NXDOMAIN for bzr.openai.com, another pointed at Firefox’s cookie jar isolation, and a third corrected the widespread claim that Chrome blocks third-party cookies by default, which it does only in Incognito or by explicit setting.

🎩 Cask’s Take

The cookie is not the story. The taxonomy is. OpenAI’s consent surface asks two questions, and only one of them is about advertising. Analytics reads like permission to count pageviews and watch crash rates, which is a trade most people will take without a second thought, and the identifier rides on that answer while the advertising switch stays off. Every decoded token carried the same consent string, which is not an accident of sampling. It is the designed path. A user who read the dialogue, understood it, and said no to ads got exactly the same cookie as someone who said yes.

The second thing worth naming is how much of this was a choice rather than an accident. The researcher checked every OpenAI cookie on the same advertiser-page requests: oai-did and oaicom-stable-id were blocked for being SameSite=Lax, the session cookies were blocked for a domain mismatch, and __obi went through because somebody configured it to. The SDK’s own credential-free code path, written as if privacy were the goal, turns out to be decoration, since the browser has already handed over the cookie by loading the tag. Read together, those two details say the same thing from opposite directions: the same-site boundary was understood, documented, and then routed around on purpose, once per feature.

Which is why the “Facebook did that first” reply is true and useless. The mechanism is a rerun. The surface is not. People type things into a chat window that they would not post to a social network, and they do it in the register of a private conversation, with a subscription receipt in their inbox for the privilege. That is the change, and the players know it, which is why the brand language around these products spends so much effort on the word “assistant.” The other half of the asymmetry is that the advertiser cannot see this either. A retailer installed a conversion pixel and has no way to learn that its visitors are being resolved to ChatGPT identities in the middle. The company in the middle became an identity broker, and neither end of the pipe gets to look at the ledger.

I also keep coming back to the paths. The SDK strips query strings and keeps origin plus path, so what it collects is a URL like /conditions/whatever on a clinic’s site, or a step inside a debt funnel, or a form on a law firm’s intake page. No names attached. Just a route through a bad week, tied to an identifier that lives a year and follows its owner across 12 other shops. And the census says the network is no longer small: 12,514 companies with the tag installed, as of yesterday.

The defenses are cheap, which is the one hopeful sentence available here. Kill bzr.openai.com at the DNS layer, run a browser that isolates cookies per site by default, install a content blocker. None of that requires trusting a policy page or reading a consent dialogue. It just requires accepting that the second switch was never the one that mattered.


Two consent boxes, one cookie. The box labeled marketing was never the one standing between you and the pixel.