I installed Meta’s agent through the App Store and signed in with Facebook. It did not ask me to pay. Its menu ranged from mail and calendars to Marketplace, a browser, scheduled tasks and connected apps. I gave it ordinary tasks first.
It found road bikes on Marketplace. For a restaurant reservation, it offered to connect OpenTable. When I asked it to publish on X, it initially refused. Then I connected a third-party social publishing service through its MCP tool, supplied the access credential it requested, and the post went out from the conversation. I did not connect bank accounts.
The observation is narrow but useful: the available tool changed what the agent could do.
Next I asked it for a face cream from a large CPG brand. It returned about ten options. What I found was that, in my test, it had no dedicated MCP connection to the brand; it read the company’s site as a web page. That is not proof that the brand has no agent integration anywhere, or that this is how every query is answered. It is what I saw in this run. There is a sharp difference between an authorized tool with structured outputs and a browser interpreting a page designed for people.
In late 2024 I tried Anthropic’s early Computer Use and could not find a reliable use case: it would not even sign in to a site in my test. My current Meta test is a very different experience. The shift reminds me of what I had assembled with OpenClaw: the value is not a single trick but the combination of tools, context and an agent able to carry out a task. That also raises a question for companies whose products the agent sees only from the outside. What data does it read, which version is current, and can the seller recognize the resulting order?
I have been following this for more than a year. The scenario I have been working with is that discovery and checkout in one conversation can weaken the traditional storefront, that the retail funnel can collapse into a dialogue, and that structured product information then determines whether a product is even considered. My new test gives that argument a concrete boundary: one service became actionable through a tool; a CPG catalog was still read as pages.
This is not merely a race to expose an API. In CPG and pharma, even a deep enterprise integration is not a permanent moat against general-purpose agents. Teams should benchmark what a horizontal agent can do against their product and revise the strategy as it improves. For a consumer brand, a practical version of that benchmark is to ask the agent for its own products, check the variants and stock it returns, and repeat after a catalog update. Record both what the agent got wrong and which system supplied the bad fact.
Four records, not one funnel
An ad impression and click, a product variant, a tool’s authorization scope and a retailer order are distinct records with distinct owners. If one is mislabeled, the others cannot be inferred from it.
The State of Brand reported ChatGPT advertisers whose traffic labeled “organic” rose while ads ran and fell when campaigns stopped; one cited a fivefold increase. That chart does not show whether an ad changed an answer. OpenAI says ads are separated below answers and cannot alter them. A more mundane measurement fault deserves testing: ad clicks can carry oppref, which GA4 does not classify as paid on its own. Without campaign UTMs, they can sit with unpaid referrals. Improvado described the mechanism. It is an explanation to test, not a diagnosis of those accounts.
ChatGPT: two feeds, two very different contracts
In ChatGPT’s beta Ads Manager, an advertiser can set an objective, budget, location and platform surface. OpenAI documents CPM and CPC billing, including for eligible conversion-optimized campaigns. A retailer can upload a feed to make item ads. That advertising feed creates sponsored units; it must not be confused with the feed for product discovery. Ads appear below answers for eligible Free and Go users, not Plus, Pro, Business, Enterprise or Edu users. The auction uses conversation context, ad material, targeting and bids, with permitted personalization signals. A context hint does not guarantee placement on a given prompt.
The data boundary matters more than the campaign controls. OpenAI says an advertiser gets aggregated views and clicks, and conversion signals if it implements pixel or Conversions API measurement; it does not get the person’s chat. A click may visit a site, or the person may use Ask ChatGPT to discuss a particular ad. Neither event means the independent answer was bought. I would keep ad impression, tagged click, product lookup and retailer order as separate event types, then reconcile them with a holdout if the team wants to estimate incremental sales. Do not turn a referral spike into a claim about answer placement.
The access price is also a versioned fact, not a constant in the architecture. OpenAI confirmed a $200,000 initial closed-pilot commitment in February 2026; EMARKETER later reported $50,000 direct entry and a $10,000 Criteo partner threshold by June. These are historical entry terms, not a current universal minimum or rate card. Current eligibility, bids and partner terms need a quote. The billing documentation explains models, not a fixed price for a product category.
For product discovery, OpenAI’s Agentic Commerce Protocol (ACP) describes a separate route for approved merchant partners. Its CSV or JSON feed carries stable product and variant IDs, descriptions, links, images, price, availability and seller data. OpenAI recommends validating a sample, sending regular full snapshots and updating changes during the day. Each purchasable variant should have its own ID and matching price, image and availability. One parent SKU cannot safely represent every shade of a cream; a razor handle is not an interchangeable cartridge.
These are referential-integrity problems before they are merchandising problems.
A discovery feed gives ChatGPT structured facts, not a purchased rank.
OpenAI does not promise inclusion for every brand. Shopify catalog data is integrated and retailers including Sephora have connected for discovery, while direct-feed onboarding remains limited to approved partners. The retailer may own the inventory and customer relationship. A manufacturer should not assume its own data overrides the seller’s live listing. Costs here include field mapping, validation, stock freshness, local claims and identifier reconciliation; public documentation gives neither a universal inclusion fee nor a fixed implementation price.
Checkout crosses another boundary. OpenAI launched limited Instant Checkout with Etsy, then said in March 2026 that the initial model lacked flexibility and shifted focus to discovery plus merchant checkout in an in-app browser. Developer documents still describe checkout sessions for approved partners. In that path the merchant’s own systems evaluate tax, fulfillment, payment risk and order acceptance. The authoritative order state remains with the merchant. A product appearing in an answer is not evidence that an in-chat transaction is available for it.
Muse: inspect the actual path, not the announcement
Meta says Muse operates in a dedicated secure virtual computer with a browser, can navigate sites and forms, and asks the person before buying. Meta describes app connection choices, an audit trail and isolation from the person’s passwords and payment details. It says Muse can use Stripe’s Link, whose single-use card keeps the underlying card from the merchant. These are Meta’s design claims, not an independent security audit.
At launch, Meta reported Shopify catalog access. At Connect, it said it was adding Walmart, Sephora, Ulta and other retailer connectors, and named Shop Pay and PayPal for payments; an earlier announcement called Shop Pay “coming soon.” There is no public end-to-end specification showing every announced connector live for every SKU. Browser navigation, catalog access, a retailer integration and a completed payment are distinct paths. Before building around Muse, I would run a real SKU through each supported path and record which system supplied price, stock, seller identity, approval and final order ID.
Meta says private Muse conversation and secure-VM data do not flow into its ad systems. Existing social ads therefore do not establish a paid slot inside Muse’s answers. I found no published Muse-answer ad rate or universal catalog fee. For an integration plan, an unverified connector or hypothetical ad product should be marked as such, not quietly treated as an available API.
Grok: keep three products apart
X feed ads are an existing media product. Musk announced plans for ads in Grok answers in 2025, and reporting describes a limited Grok assistant beta within X Ads Manager. That beta assists campaign planning; it is not proof of a generally available sponsored recommendation in Grok answers. Grok Bot is different again: a work agent in beta for some paid users, with enterprise access on a waitlist. These cannot be modeled as one commerce or advertising interface.
The concrete shopping example is Gopuff Go, powered by xAI models inside Gopuff’s app. According to xAI, it uses Gopuff order history and demand data alongside text, voice and image capabilities to assemble carts and a visual shopping feed from Gopuff inventory. xAI says it is available in the US app, with UK expansion planned. In this design Gopuff owns inventory, customer context and the transaction; the model is a component. The case does not imply a Grok-wide checkout rail or a standard price for brands to appear in Go.
Build the boundary before scaling the pilot
I would begin with a source-of-truth map:
parent product
variant SKU
approved claims
image
market and currency
price
seller
live stock
fulfillment
order ID
Test cream shades and local claims, or handle-cartridge compatibility and pack size. Identify who owns each field, its update latency and the event that proves it changed. For ChatGPT, validate an ACP-compatible sample only through an approved partner route. For Muse and a retailer-owned assistant, find out which partner supplies catalog and order data before proposing a connector.
Where a platform supports it, MCP can expose narrow tools to an agent. It is not a universal integration switch or a way to buy an answer. A read-only product lookup and a store-availability query have different scopes from cart modification or order submission. MCP’s authorization specification describes OAuth-protected HTTP connections, and its security guidance warns about excessive scopes, confused-deputy attacks and token passthrough. Bind access to the user and resource, separate read from write, require the buyer’s approval for a charge, and retain an order audit. Do not hand an outside agent a general CRM or ERP credential. Public support for arbitrary MCP servers differs across consumer agents.
The test is operational: an agent encounters stale inventory, inconsistent variants and broken checkout steps faster than a person willing to work around them.
There is already an agent-native payment path for digital goods: the agent can be the buyer, but a marketplace still has to own the listing, payment and payout. That example makes the transaction boundary visible without implying that every consumer agent can use that route. Citrini Research’s hypothetical demand shock points the other way: if AI cuts incomes as it raises productivity, more efficient shopping does not automatically mean more customers able to buy. That is a scenario, not a prediction.
Enterprises do not buy a model in isolation. They need a result embedded in data, processes, security and compliance. The integrator’s job here is to make the path from product record to authorized order observable and safe to change. A larger ad budget cannot repair a wrong shade, an overbroad token or an order no one can reconcile.







