
Principal Basil C. Puglisi reads this week’s ChatGPT Ads news as a measurement problem every marketing team will meet before its next media plan. The announcement arrived carrying several kinds of proof, and each one answers a different question about the channel. On October 5, OpenAI announced a new visual ad format, expanded measurement partnerships, and brand suitability pilots for ChatGPT, which it says reaches 1.2 billion people each week. The visual format will be tested during image generation for users on the Free and Go plans, starting later this month in the US with an initial group of advertisers.
The stakes sit in how a budget owner reads the numbers that came with it. OpenAI’s release cites a 15.3 percent lower attributed cost per acquisition for Weight Watchers against its blended paid-search benchmark. It also cites statistically significant lift for the wellness brand Dose and a 93 percent new-visitor share for Portland Leather, and a different measurement partner produced each figure with a different method. Puglisi’s consulting position is that a marketer should grade every one of those numbers by the question it answers before any of them moves money. An attributed CPA and a lift test can both be accurate while they support very different budget decisions.
Factics, the method Puglisi introduced in 2012, pairs every fact with a tactic and a KPI defined before the action. His October 5 article, One Hero Ad Is Not a Test: What Meta’s AI Ad System Means for Small Budgets, applied that rule to platform-reported results. It treated Meta’s numbers as context and let the account’s own KPI carry the commitment. ChatGPT Ads gets the same treatment here with one difference worth naming. Several of today’s figures come from independent measurement partners, which makes them stronger facts without changing what each method is able to show.
What OpenAI and its partners announced on October 5
OpenAI says its measurement approach supports both in-house and partner-led solutions. Integrations with Hightouch, Tealium, and LiveRamp let advertisers send conversion data from the systems they already run. Attribution partners including AppsFlyer, Triple Whale, DV Rockerbox, and Northbeam now support Conversions API, reporting, and click attribution. For causal measurement, OpenAI writes that its incrementality work is still in its early stages, and it’s partnering with Haus, Measured, and WorkMagic to explore geo-based experiments.
DoubleVerify’s release the same day adds the detail behind the headline figure. Weight Watchers used DV Rockerbox to measure ChatGPT Ads as a new source of member acquisition, and the attributed cost per acquisition came in 15.3 percent below the company’s established benchmark during the measured period. Its VP of Growth called the result an encouraging early signal.
Integral Ad Science announced that it’s one of the first media quality partners supporting ChatGPT Ads, with third-party brand safety and suitability reporting across IAS taxonomy categories. That integration sits in a closed pilot with a select group of advertisers, and wider availability is planned. OpenAI describes the DV and IAS suitability work as pilots in a controlled testing environment. Independent partners assess how its brand safety standards are applied there, without access to private user conversations.
Each ChatGPT Ads number answers a different question
Attribution answers where credit lands. The Weight Watchers figure says that conversions its model credited to ChatGPT cost less than its usual acquisition cost over one measured period. DV presents that number as context alongside the brand’s existing channels. Attribution cannot say how many of those members would have joined anyway through search or another route, and that counterfactual is the question a decision to move budget depends on.
Incrementality answers what the ads caused. WorkMagic reported statistically significant lift for Dose, with 67 percent of incremental purchases coming from net-new customers. A lift result of that kind is the stronger evidence for a spend decision. Puglisi would carry OpenAI’s own early-stage label into any client deck, because one brand’s lift doesn’t transfer to another brand’s category, price point, or audience.
The 93 percent new-visitor figure for Portland Leather, reported by Triple Whale, measures reach into a new audience and doesn’t yet speak to purchases or margin.
Suitability verification answers a separate question about context. DV CEO Mark Zagorski said conversational context is dynamic, since what a user asks and how the conversation evolves both shape where an ad appears. Independent evaluation in a controlled environment gives a brand some evidence about how OpenAI’s guardrails behave. The honest limit today is that both suitability programs are still pilots.
How Puglisi would set up a first ChatGPT Ads test
The consulting brief Puglisi recommends starts with the decision the test has to inform, written down before the first dollar goes out. Suppose the decision is whether to move budget from paid search. Then the KPI is incremental conversions per dollar, measured against a holdout region or a geo experiment, and the attributed CPA becomes a secondary read. When the decision is whether ChatGPT reaches customers the brand doesn’t already have, the KPI is purchases by first-time customers, with new-visitor share as a supporting signal.
The brief also names the review date and the owner who reads the result. It records which partner measured each figure and by which method. Leadership can then see an attributed number and a causal number side by side without blending them into one claim. Brand teams in sensitive categories can add a suitability line to the same page. It asks whether the account has IAS or DV reporting during the pilot and which Negative Phrases it will use, a control OpenAI now offers to qualifying advertisers.
A smaller advertiser without access to geo experiments can still follow the Factics order. The team sets a fixed budget, a fixed window, and a baseline from the prior period. It also writes down, before launch, what result earns the next round of spend, so the test has a way to fail and the owner learns something either way.
What to watch as ChatGPT Ads scales
ChatGPT Ads is entering the market with more third-party measurement attached than most new channels carried at the same stage. That raises the quality of the facts a marketer can start from. The unresolved piece is causal evidence at scale, and the gaps are specific. OpenAI calls its incrementality work early, the IAS integration is a closed pilot, the DV suitability pilot is still in development, and the visual format hasn’t begun testing. Puglisi would watch for two things over the next quarter: published geo-experiment results across more than one category, and suitability reporting that moves out of pilot into general availability. Until those arrive, his guidance is to treat the October 5 figures as promising facts about other brands. The team’s own KPI, written before launch, should decide the budget.
Sources
- DoubleVerify. (2026, October 5). DV launches attribution measurement and advances brand suitability for ChatGPT Ads [Press release]. GlobeNewswire. https://www.globenewswire.com/news-release/2026/10/05/3374412/0/en/dv-launches-attribution-measurement-and-advances-brand-suitability-for-chatgpt-ads.html
- Integral Ad Science. (2026, October 5). Integral Ad Science to provide brand safety and suitability measurement in ChatGPT Ads [Press release]. https://integralads.com/news/ias-bss-chatgpt-ads/
- OpenAI. (2026, October 5). Building advertising for the way people use AI [Announcement]. https://openai.com/index/new-chatgpt-ads-format-and-measurement/
Questions readers ask
What did OpenAI announce for ChatGPT Ads on October 5, 2026?
OpenAI announced a visual ad format to be tested during image generation for Free and Go users in the US, starting later in October. It also expanded conversion and attribution integrations with partners such as LiveRamp and DV Rockerbox, began early geo-based incrementality work, and opened brand suitability pilots with DoubleVerify and Integral Ad Science.
What is the difference between attributed CPA and incrementality?
Attributed CPA divides spend by the conversions an attribution model credits to ChatGPT Ads. Incrementality uses a controlled test, such as a geo experiment, to estimate the conversions the ads caused that would not have happened otherwise. A decision to move budget from another channel depends more on the incremental result.
What did the Weight Watchers ChatGPT Ads result show?
DoubleVerify reports that Weight Watchers used DV Rockerbox to measure ChatGPT Ads as a new source of member acquisition. The channel’s attributed cost per acquisition was 15.3 percent below the company’s established benchmark during the measured period, which the company called an encouraging early signal as it evaluates how to scale.
How does brand suitability verification work in ChatGPT Ads?
OpenAI is developing pilots in which DoubleVerify and Integral Ad Science assess how its brand safety standards are applied in a controlled testing environment, without access to private user conversations. IAS reporting is in a closed pilot with select advertisers, and qualifying advertisers can also add Negative Phrases to narrow where their ads appear.
How does Factics apply to a first ChatGPT Ads test?
Factics pairs a verified fact with a specific tactic and a KPI written before the action. For a first ChatGPT Ads test, the team names the budget decision, sets an incremental or first-time-customer KPI against a baseline, fixes the budget and review date, and records which partner and method produced each number.
#AIg
Leave a Reply