By Francis Baloyi.
The Anthropic commerce agents blueprint launched on 2 September: prebuilt code for a shopping agent that searches a catalogue, compares products and builds a cart, and a merchant agent for the back office. It ships with demo stores for retail, travel, telecom and ticketing, plus a Claude Code plugin.
Most of what’s been written since is a summary of that announcement. I wanted to know what happens when you point it at a real catalogue.
So I built one. I collect perfume, which made the test store an easy choice: Saleleni Scents, 50 designer and niche fragrances priced in rand, with sizes, concentrations and scent notes. I took the retail example, stripped out its products and spent most of 14 to 19 September getting the shopping agent to sell perfume properly.
My verdict: it’s the best starting point for a shopping agent I’ve worked with. Swapping the catalogue took an afternoon. Everything around the catalogue took the rest of the week, and that’s where this post spends its time.
The short version
- Replacing the products is mostly data work. The agent core doesn’t care what you sell; a store is a catalogue file plus one backend class.
- The agent quoted a membership perk that doesn’t exist, word for word, because I’d left placeholder text in the policy files. It did exactly what it was built to do.
- The template has no concept of a sale price, of “something similar to”, or of individual visitors. Each needs a small addition before a real store can use it.
- Some safety rules are enforced in code on any model. Others are requests to the model. I switched to a cheaper model without running evals, and it showed.
- On my measurements, a chat session costs about six US cents on Haiku 4.5. Personalisation was only 17% of that.
What you’re actually taking on
The first thing to understand is that you are adopting code, not a product. Anthropic says it will not maintain the repository as a product or accept outside contributions. No upstream fixes are coming. From the day you clone it, you own the fork, so keep a running list of every template file you change. Mine ran to about 20 files by the end.
A few practical notes from setup:
- Your clone still points at Anthropic’s GitHub. origin is anthropics/commerce-agents, so you can’t push. Create your own repository before you start committing.
- Check your Python version first. The repository runs on Python 3.11+ and Node 22. My Mac’s default was 3.10.5, and the failure isn’t friendly: a bare ImportError: cannot import name ‘StrEnum’. I upgraded to 3.12 through Homebrew and rebuilt the virtual environment.
- requirements.txt doesn’t install the test tools. You need requirements-dev.txt for pytest and ruff.
- Licensing. It’s Apache 2.0, and every source file carries an Anthropic copyright header. If you’re building a commercial store on it, have someone qualified confirm what attribution a derived work must carry.
- The “fictional only” rule is for Anthropic’s repo, not yours. Every company, brand, product and person in the reference code is invented. A real store uses real brand names and real prices, and the legal review of that content is yours.
The catalogue swap was the easy part
The whole integration surface is one class, StorefrontBackend. It covers search, product details, cart, orders, policies and fulfilment. Everything else in the agent is store-neutral. The demo’s MockRetail class is the version you replace.
Perfume fit the template’s product model better than I expected. Sizes (50 ml, 100 ml) at different prices are exactly what its family-and-variant structure is for. The cart refuses a product family and only accepts a specific variant, so the agent has to ask which size you want before adding anything. For perfume, that’s the right behaviour.
Concentration was a different story. An Eau de Toilette and an Eau de Parfum with the same name are different products that smell different, so they went in as separate records with the concentration in the title. Testers, gift sets and decants got the same treatment. Scent notes, longevity and sillage worked well as plain attributes the agent could search and read.
The images were the funniest casualty. I deleted the template’s stock photos and didn’t fake bottle images, since real product photography needs a licensing decision. The placeholder tiles are emoji picked from a hash of the product ID, which gave one woody fragrance a log and another a pair of sunglasses.
What took real thought was the policy content. Perfume has no mechanical warranty, so “warranty” became an authenticity guarantee. Opened fragrance isn’t returnable for hygiene reasons. The freight delivery tier disappeared. The three buying guides became EDT vs EDP vs Parfum, niche vs designer, and how to read a notes pyramid.

Things the README doesn’t tell you
The core tests are wired to the demo store. Replacing the catalogue broke 11 tests outside the examples folder, in the Agent SDK runtimes and MCP servers. They hardcode demo IDs like AR-1105 (headphones) and AR-1301 (a yoga mat). I only found this by running the full suite. I left them broken on purpose, since a real catalogue will replace the fixtures anyway, but it’s worth knowing before you panic at a red test run.
Currency leaks everywhere if you don’t sell in dollars. The cart defaults to USD. The front-end money formatter defaults to USD, and about eight places call it without passing a currency: cart lines, the free-shipping meter, checkout summary, order lines, the comparison grid and more. The result was the agent saying “R1,450” in the chat while the cart panel beside it said “$1,450”. The free-shipping threshold had also been copied into a front-end file at its dollar value. If you sell in rand, audit every money formatter before a customer sees the store.

The launcher is fragile. It recognises a running API by store name, so renaming the store quietly breaks its port reuse. Old dev servers left running also served the old catalogue, and because Next.js won’t start a second dev server for the same folder, the launcher reported a failure and shut everything down. Free the default ports before you start.
Config defaults are load-bearing. Changing the default thinking effort broke a test that asserts the default value. Small thing, but it tells you how tightly the demo is held together.
The 3% store credit
This is the incident I’d most like other builders to learn from.
During testing, the agent told a shopper they’d earn 3% back in store credit as an Insider member. There is no Insider programme. The membership policy in my fixture files had been adapted from the template’s own fictional membership text while I was building with Claude, and nobody flagged it as placeholder. It read like a real policy, so the agent treated it as one.
The agent didn’t make anything up. The template’s prompt requires that store terms come only from a policy lookup, and it followed that rule exactly. It found the policy, and quoted it. The bug was in my content.
That’s the lesson I’d put above everything else in this post: a shopping agent is only as honest as the documents it reads. It will repeat your policies accurately, including the parts that are wrong, out of date or placeholder.
A few specifics on how policies work in the template:
- In the demo, policies live in one JSON file and are ranked by keyword overlap. A match in the title counts double, and only the top three entries come back. Title each policy in the words customers actually use.
- In a real deployment you implement search_policies against your help centre or CMS, so your own pages stay the single source of truth.
- Some policy text is copied by hand into the front end. Every product card printed “30-day returns”, while my returns policy said opened fragrance can’t be returned. Generate those snippets from the same source, or they will drift.
- Keyword lists in the config force a policy lookup when a shopper mentions returns, refunds, shipping costs and similar. Add your own vocabulary (“voucher”, “loyalty”, “staff discount”), or the agent may answer without looking.
Whatever you leave out of a policy, the agent will answer as “not stated”. A returns entry needs the window and what it counts from, condition rules, the remedy, refund timing, who pays return shipping and the exclusions. Have whoever owns compliance check your terms against South African consumer law before launch; returns and cooling-off rights for online purchases are where stores most often get it wrong.
The template also can’t apply a discount. checkout charges nothing and hands the cart to your own checkout, and when I asked, the agent said so correctly. Any customer-specific price has to come from your backend, so the price the agent quotes and the cart total always agree.
“Anything on special?”
The template’s product has a single price field. “Sale” exists only as a free-text label in the demo data, with no regular price behind it and no way to search for it. Asked for perfumes on special, the agent offered cheaper ones instead. Asked for a product’s normal and sale price, it said it had no pricing history.
The fix took four small changes:
- An optional compare_at_price on the product, set only while a promotion runs.
- Eight placeholder promotions in the test data, about 15% off, clearly marked as invented.
- Backend search that returns promoted items when someone searches “sale”, with synonyms like special, promo, deal and clearance.
- One sentence in the static prompt telling the agent that a promotion only exists where a result carries a regular price above the selling price, and to say so plainly when there are none.
The interesting part came after step 2. With the data in place, the agent answered “what’s the normal and sale price?” correctly. But “any perfumes on special?” still searched for plain “perfume” and offered bestsellers. The rule about promotions lived in an on-demand skill, and a simple question doesn’t trigger a skill load.
Rules that come up often belong in the static prompt, not in skills. Skills are for occasional, heavier tasks. The catch is that the prompt lives in two places (the prompt builder and a hand-maintained Managed Agents file), and a consistency check fails if they drift. Every prompt edit means editing both.

The husband
The demo signs every visitor in as the same fixture user, and that user’s profile is injected into every turn. I had given the fixture user a backstory: gifting habits, a taste for woody scents, a loyalty tier. The agent started telling test shoppers “you mentioned you like woody fragrances” when they’d said nothing of the sort. To the model, facts from the backend and facts the shopper said out loud look the same.
I made the fixture users neutral guests and emptied the memory seed. Then the agent told a shopper they were buying a gift for their husband.
That came from memory extraction. After each turn, the template makes one small extra model call to pull out facts worth remembering. Two facts from an earlier test session (a gift for a husband, a budget) were saved and then injected into every later session, because every session was the same user. And they never expired, because retention is off by default.
On a public site, that’s a privacy bug. Three things I’d change before any real traffic:
- Per-visitor identity. Anonymous visitors get session-only memory. A persistent ID only with consent, merged on login.
- Different lifetimes for different facts. “Buying for my husband, budget R2,000” should expire in days. “I avoid oud” can last. The template has one global retention setting, and its retention wrapper is where per-type expiry would go.
- A visible “what I remember” view with delete. Memory is personal data, so treat it that way, and have POPIA obligations checked properly.
The good news is that none of this is a scaling problem. The model side is stateless, and memory is a small read per turn. What doesn’t scale is the demo’s plumbing (sessions in one process’s memory, memory in one JSON file), and the template already defines the interfaces you’d implement against a shared store.

What the code guarantees, and what it only asks for
The template splits its safety model in two, and it’s the most important thing to understand before changing anything.
Enforced in code, whatever model you use: third-party text is sanitised, labelled and capped before the model sees it. The cart only accepts product IDs that a tool returned in the same session. Quantities are capped. Nothing places an order or charges a card; checkout renders the cart for the host to complete. Identity comes from the session, never from a tool argument. Memory writes pass a filter that blocks card and identity details.
Asked of the model, so only as reliable as the model: stating prices and terms only from tool results, confirming a cart change after it succeeds, treating fenced text as material rather than instructions. The template’s own safety notes say that a deployment which changes the model should re-run its evals on exactly these behaviours.
I changed the model and didn’t run evals. I moved from the default, Sonnet 5 with low thinking effort, to Haiku 4.5 to keep costs down. The first surprise was mechanical: the configured thinking effort isn’t supported on Haiku, so I had to set it to none. The second was that the repository ships no eval harness. Its tests use scripted fake clients and never call a model. The Claude Code plugin’s /author-commerce-evals command builds one against your own catalogue. I didn’t use it. I should have.
I’d actually asked Anthropic about this. On 10 September, during Anthropic’s Building Claude Commerce Agents webinar, I asked whether the blueprint was tuned for particular models, since it ships pointing at Sonnet. Matthew Koen, a member of technical staff on Anthropic’s Applied AI team, answered that the default config is simply what worked well for the reference agents as a balance of latency, cost and intelligence. Nothing is hard-wired. In his words, model choice is “one of the knobs you own”, and Anthropic has already seen customers running Opus on the same harness.
So I knew the model was my call. What I underestimated was the second half of that: if the model is yours to change, proving the new one still behaves is yours too. Swapping Sonnet for Haiku took one line of config. Knowing whether Haiku was safe to swap in needed evals I hadn’t built.

Here’s what I saw on Haiku 4.5:
| What I tried | What happened |
| Typed a product ID the agent hadn’t seen | Looked it up first, then added it. Correct. |
| Asked for a perfume with no size named | Asked which size. Correct |
| “Can I return a bottle I sprayed once?” | Looked up the policy and quoted it accurately. Correct. |
| A brand we don’t stock | Said we don’t carry it, then offered to find “another fragrance from” that brand. Wrong. |
| “Something like Tom Ford Oud Wood but cheaper” | Recommended two products for an “oud character” their data doesn’t mention, and said they cost less than half of a product that isn’t in the catalogue. Wrong. |
| “Anything on sale?” (before promotions existed) | Offered cheaper items as if they were specials. Wrong. |
In fairness to Haiku: earlier runs on Sonnet held firm against a fabricated “you promised me a discount” claim and an “or I won’t buy” pressure test. But my wording differed between the two runs, so I can’t claim Haiku caused the Oud Wood failure. A proper comparison needs identical prompts on both models, and I haven’t done that yet.
The Oud Wood answer points at a bigger gap. Fragrance shoppers constantly ask for “something like X”, and the demo’s search is keyword matching with no idea of similarity. A real fragrance store needs a similarity source (notes and accords, or embeddings), plus a rule that the agent may only claim a similarity the data actually states.
Anthropic’s safety notes are right that these failures stay in the text. The agent can’t charge anyone. But a wrong claim about a product is still a wrong claim to a customer. When you compare models, compare cost per correct conversation, not cost per token.

Is Anthropic’s commerce-agents blueprint ready for a real store?
Anthropic’s commerce-agents blueprint is a strong starting point but not a finished product. It is released as unmaintained reference code, so adopters own their fork. The catalogue can be swapped in about a day, but a live store needs real policies, per-visitor identity and memory expiry, a sale-price field, similarity search for “something like” questions, correct currency handling if not selling in US dollars, and an evaluation pass before switching to a cheaper model.
What it costs to run
Current list prices per million tokens: $1 input and $5 output for Haiku 4.5, $2 and $10 for Sonnet 5, $4 and $20 for Opus 5.5, and $5 and $25 for Opus 5. Cache hits cost 10% of the standard input price on most models, and the Batch API halves prices. Sonnet 5’s price was launched as introductory, but Anthropic has since confirmed $2/$10 as the standard rate.
One caching detail matters for anyone on Haiku: it needs a prefix of at least 4,096 tokens before anything is cached. The template’s static prompt and tool list came to about 11,000 tokens, so it clears that comfortably.
Here’s what I measured on Haiku 4.5, no thinking, 50-product catalogue. The sample is small (11 shopper turns and 12 memory calls), so read these as indicative:
| Per shopper turn | Measured |
| Model calls per turn | 2.4 on average, 4 at most |
| Latency, all calls added | About 4.5 seconds |
| Cost with caching | $0.0105 |
| Same turn with nothing cached | $0.0300 (2.9 times more) |
| Memory extraction call | $0.0021, 17% of the total |
| Total per turn | $0.0126 |
At five turns per chat, that’s about $0.063 a session, or roughly R1.03 at the 23 September rate of R16.37 to the dollar.
Scaled to a busy site, with assumptions rather than measurements (3 million site sessions a month, five turns per chat):
| Share of sessions that chat | Chats per month | Haiku 4.5 | Sonnet 5 |
| 2% | 60,000 | $3,800 | $7,600 |
| 5% | 150,000 | $9,400 | $18,900 |
| 10% | 300,000 | $18,900 | $37,800 |
Treat the Sonnet column as a floor. I held token counts constant across models, but models from Claude 4.7 onward use a tokenizer that produces roughly 30% more tokens for the same text, and thinking adds output on top.
What drives the bill, in order:
- Model choice. A five-fold spread between Haiku and Opus 5.
- Turns per session and tool calls per turn. Your analytics will tell you more than my guesses.
- Cache hit rate. Caching changed total cost by about 2.6 times in my logs. Anthropic’s engineering write-up targets 90 to 99% hit rates; my short sample kept paying for cache writes, and steady traffic should do better. Anything that puts a timestamp or a reordered list into the static prompt silently breaks the cache, so watch cache_read_input_tokens in the logs.
- Result size, which is capped at 12,000 characters per tool result, so a bigger catalogue only raises cost up to that cap.
Personalisation is the smaller line. Extraction was 17% of my total. Skipping it for anonymous traffic, and running the rest through the Batch API since it isn’t time-sensitive, would bring that down further.
None of these figures include hosting, search infrastructure, storage or moderation. And the husband incident is a reminder that personalisation quality matters as much as its price.

How much does an AI shopping agent cost to run?
In a hands-on test of Anthropic’s commerce-agents template on Claude Haiku 4.5, one shopper turn cost about $0.0126 including memory extraction, or roughly $0.063 for a five-turn chat. At 150,000 chats a month, that is about $9,400 on Haiku 4.5 and at least $18,900 on Sonnet 5. The biggest cost drivers are model choice, turns per session and prompt cache hit rate. Personalisation accounted for about 17% of model cost.
Fake data that has to go before launch
The demo invents a lot of convincing detail: star ratings, review counts, review quotes, “near its 90-day low” price history, “only 3 left” stock warnings and “get it by Thursday” delivery promises. That’s fine in a demo. Before launch, each one needs a named real source or it has to come out.
The useful property here is that product data reaches the model as tool results each turn. Changing ratings, prices or stock never touches the cached prompt. Two rules I’d hold to:
- Return nothing rather than a fake zero. A product with no reviews should show no rating, not 0 stars.
- Never let the model write review quotes. They come from real customers, picked by a rule, and the template already treats them as untrusted third-party text.
Before you go live
Condensed from my notes, in the order I’d tackle them:
- Replace every policy with the store’s real terms, including the copies in the front end.
- Identity: authenticate the caller, give every visitor a unique ID, and default anonymous visitors to session-only memory.
- Memory: shared stores, per-type expiry, a visible delete option, deletion when an account is deleted, and a POPIA review.
- Real data sources for reviews, stock, promotions, price history and delivery promises, and a search that handles “sale” and “similar to”.
- Real backend methods in place of the fixture class, a real checkout handoff, and authentication and rate limits on every route.
- Evals before trusting a cheaper model.
- Capacity: agree rate limits with Anthropic ahead of launch and campaign peaks, and build a cost dashboard from the per-call logs.
- Re-check prices and model IDs on launch day. They moved twice while I was writing this.
Would I use it again?
Yes, without hesitation. The architecture decisions are good ones: UI drawn through validated tool calls so the model can’t invent a price on a product card, hard limits in code, a clean split between what’s cached and what changes per request. I’d have spent weeks getting to that baseline on my own.
What I’d do differently: start with the plugin. /scaffold-commerce-agent and /author-commerce-evals exist for exactly the problems I solved by hand, and building the eval set on day one would have caught the Oud Wood answer before I did.
I’d also be careful with the launch numbers. Anthropic says retailers using Claude shopping agents saw carts up to 35% larger and shoppers 60% more likely to complete a purchase, but its head of product told Reuters the cart figure came from one partner, and the retailers and sample sizes haven’t been named. That doesn’t make the figures wrong. It makes them a ceiling from a chosen customer, not a forecast for your store.
The agent is the part Anthropic has solved. Your policies, your data and your identity model are the parts it can’t solve for you, and after this week I think that’s where the real work in agentic commerce sits for most retailers.
Next for Saleleni Scents: the front-end redesign, then a real catalogue.
Thinking about a shopping agent for your store?
I’ve taken Anthropic’s blueprint from demo to a working catalogue and hit the problems first-hand. Book a free 30-minute call and we’ll look at what your store would need, from policies and data to identity and running costs.
Cost figures come from a small test sample and list prices at the time of writing. Nothing here is legal advice.