 | Hello. On Sunday night Amazon started showing a popup to anyone who tried to buy through Meta's new Muse agent: unauthorized AI agent, violates our conditions of use. Four days earlier OpenAI had gone the other way and started selling ads inside ChatGPT that open a conversation with the advertiser's own agent, with a free Shopify app to buy them. Same week, Google confirmed it pays small sites when their pages shape an AI answer, a former OpenAI researcher launched a model that answers only in typed values, and Anthropic said Claude now leads a quarter of the work of building the next Claude. I also ran the blind bench I've promised three issues running. Eleven minutes, 24 cents, and I picked the cheap model. Details below. | › | One storefront blocked the agent and another sold it a ticket | | › | Google is paying small sites for being in the answer | | › | A model that never writes a sentence | | › | What I'm playing with: The blind bench I owed you, a Google Ads account built from a plan, and a platform fixed without a ticket | | › | Run it tonight: Put one product in front of ChatGPT | | › | Concepts worth exploring | | › | Tools I actually used | | › | Everything else | | › | Second thoughts |
|
| One storefront blocked the agent and another sold it a ticketSep 20 Amazon blocked Muse | | Sep 23 Shopify ChatGPT Ads go global | | 40+ Countries where ChatGPT runs ads |
What happened. Meta's Muse, the personal agent it launched in the US this month, can browse a site, log in with credentials you hand it, and buy. On the evening of September 20, Amazon cut it off. Anyone finishing a purchase through Muse now sees: "Continued access by an unauthorized AI agent violates Amazon's Conditions of Use, to which our customers have agreed." Amazon's reasons, in its statement to GeekWire: Meta never told Amazon the agent would use the store, Muse doesn't identify itself when it browses, and it appears to capture and store customer credentials. Meta says Muse can't see passwords or payment methods, which sit in secure storage the model uses but doesn't read. Amazon has spent a year on this. It sued Perplexity over the Comet browser, won an injunction in March, lost on appeal August 4, and has moved to block shopping agents from Google and OpenAI. Its own agent, Buy for Me, shops other retailers' sites, identifies itself, and lets those retailers opt out. That's the standard Amazon says Muse should meet. Four days earlier, on September 16, OpenAI announced Sponsored Agents: click an ad in ChatGPT and you can open a separate, labeled conversation with an agent the advertiser runs. The test is with select US advertisers. Alongside it, OpenAI wired ChatGPT Ads into HubSpot and Shopify, its first CRM and e-commerce partners. The Shopify app is free, live for US merchants, pulls from the catalog you already have, and opens to the other countries where ChatGPT shows ads on September 23. | | An agent that announces itself is a customer. One that doesn't is a bot, and the store gets to decide. |
The details. | › | Ads show to Free and Go users of ChatGPT in more than 40 countries. Paid tiers don't see them. | | › | Muse and Buy for Me do the same job. One announces itself and lets the store say no. Amazon blocked the one that doesn't. | | › | Meta and Amazon are partners elsewhere: Amazon products have sold inside Facebook and Instagram since 2023. The block is about the storefront, not the relationship. |
What I’d do. Decide which storefront you are. If you sell on your own site, an agent shopping for a customer is a customer, and blocking it is turning away money. Let it in, log it, and watch what it does at checkout. If you sell on Shopify, install the ChatGPT Ads app this week and put fifty dollars behind one product, because September 23 is when everyone else finds it and early inventory is cheap inventory. The one thing I'd copy from Amazon is the identification rule. Write into your terms that automated agents must identify themselves and that you can refuse any that don't. You won't enforce it this year. You'll want it on the page when you do. Source → |
| AI employees for small businessSUDO is our AI agency and product division, and it runs the systems behind this newsletter. Its AI employees answer email, chase invoices, keep the books tidy, and log a receipt for every action they take. Plans start at $199 a month, and the first two weeks are on us. Meet your first AI employee → |
| Google is paying small sites for being in the answerSep 14 Digiday reported, Google confirmed | | Dozens Sites approached, mostly small | | 3 Surfaces: Gemini, Overviews, AI Mode |
What happened. Digiday reported on September 14, and Google confirmed, that a pilot inside Search Console pays publishers when their content shapes what Gemini, AI Overviews, and AI Mode say. Google's words: "an early-stage learning pilot to test how best to reward high-quality content." At least dozens of sites have been approached, mostly small and mid-sized, and not only news. A site in the program gets an AI contribution panel in Search Console with a monthly earnings number. The rule is narrow: you earn when your page contributes significantly while the answer is being generated, not when it's linked afterward or used to confirm a fact. Google hasn't said how the number is computed. One executive in the pilot called it "quite black box." Another described the early money as peanuts next to ad revenue, and some publishers have declined, worried that a small check weakens their hand in a bigger licensing talk. You can opt out any time. | | Google will pay you for shaping the answer. It won't tell you what the answer was worth. |
The details. | › | This is the first time Google has attached a dollar figure to an AI answer for the site that fed it. | | › | Small sites are the target. Big publishers negotiate their own deals. A small site's alternative was nothing. | | › | The panel shows earnings and nothing else. You can't see what Google used, only what it paid. |
What I’d do. It's invitation only, so do the part you control. Last issue I measured how often four assistants name three small businesses. This pilot pays for exactly what that score measures: being the page that shapes the answer. If you have a page that explains something customers ask about, in plain words with a date on it, that page is now a potential earning line, and it was already the page assistants cite. Write one this month. Then check Search Console weekly for a panel you didn't put there. If it appears, take the money and keep your own count of what it should have been. Source → |
| A model that never writes a sentence$40M Seed, led by DCVC | | 3 Question types: Choice, Score, Noul | | 0.114 s Register's demo, vs 8.6 s for Terra |
What happened. TypeSafe AI came out of stealth on September 15 with a $40 million seed round led by DCVC and a model called Jev. The founder is Diogo Almeida, four years at OpenAI and a co-author of the InstructGPT paper, with Erik Gafni and Sasha Sheng. Jev doesn't generate text. You ask a typed question of one of three kinds: Choice picks from options you supply, Score rates against levels you supply, and Noul answers yes or no. It returns the value plus a probability, and because you defined the allowed answers up front it can't return one you didn't list. TypeSafe calls it a System One model, the fast reflexive kind of decision, and pitches it as the layer under an agent that makes the thousand small calls a minute an LLM is too slow and too expensive for. Responses take 70 to 500 milliseconds. The company claims 40 to 200 times faster and 40 to 400 times cheaper than frontier LLMs on its own workflows, and says in the same breath that customers will likely see less. The Register clocked a demo at 0.114 seconds against 8.566 for GPT-5.6 Terra. Pricing is $0.042 per million input tokens and nothing for output. Early access is a waitlist. No published weights, no paper. | | A model that can only pick from your list can't make things up. It can still pick wrong. |
The details. | › | Hallucination-free the way a multiple-choice test is. It can't invent an answer. It can pick the wrong one, and the probability it attaches is the only warning. | | › | A yes-or-no with a confidence score is a classifier. Classifiers were already fast and cheap. The new part is one API for all of them and a probability that's supposedly calibrated. | | › | $0.042 per million input tokens is 238 times below Fable 5.1's input price. It's also a waitlist, a self-run benchmark, and a company one week old in public. |
What I’d do. List every decision in your business an LLM currently makes by writing a paragraph you then parse. Is this lead worth a call. Which category does this ticket go in. Does this invoice match the PO. Those are Choice, Choice, and Noul, and none needs a sentence. The list is worth having whether or not Jev ships, because a classifier from any cheap model beats a paragraph for all of them today. Then join the waitlist. When it opens, take one decision off the list, run a week of real inputs through Jev and through the cheap model you already use, and count the disagreements. If the probability is honest, it's the first model that tells you when a human should look. That's the feature. The speed is the marketing. Source → |
| The blind bench I owed you, a Google Ads account built from a plan, and a platform fixed without a ticketThe bench. Issue 001 promised it, 002 and 003 admitted I'd skipped it, and this morning I ran it. One task I do every week: write the four Everything else entries below from four source notes. Same instructions, same notes, three models through OpenRouter: Claude Fable 5.1, which I pay for daily, Kimi K3, and DeepSeek V4.1 Flash. Outputs shuffled to A, B, and C, graded before I opened the key. I picked C as the one I'd paste with the least editing, B as the worst, and my verdict was that I couldn't tell them apart. C was Kimi K3. B was Fable. The bill: Kimi $0.003, DeepSeek $0.0002, Fable $0.24. One caveat in Fable's favor: the OpenRouter endpoint for Fable refuses to turn reasoning off, so it thought for 4,500 tokens before writing 300, and the thinking is most of the 24 cents. On a plain endpoint the gap is closer to list price, about three to one. On the task it didn't matter. I couldn't see the premium. The entries below are mine, not the bench's, but I'd have been fine pasting C. The ads account. A franchise client asked me to take a national search campaign back from an agency that had let it lapse. On September 18 I drafted the plan in Claude Code from their materials: five campaigns, 96 keywords, 13 ads, 124 negative terms, bid caps, and a spend governor for thin national demand. Then I had Codex review the plan cold before anything touched the account, and it caught two things I'd have shipped. On September 19 Claude Code built the account through the Google Ads API: 183 keyword rows, 13 ads, two negative lists, everything paused, bid caps read back. Two lessons. A re-apply I ran to fix one field created 13 duplicate ads in the live account, and the fix was Google's own change history and its Undo button. And Google attached a billing profile from the manager account the moment the account existed, so the only thing between the client and spend was the pause. Check billing before you check keywords. The platform. A client's new site sits on a franchise platform with its own dashboard, and our audit had found a few hundred pages with wrong titles and descriptions. The client asked us to fix them rather than wait on the vendor's queue. On September 21 Claude Code did it through the platform's API from the spreadsheet of findings: 93 location titles, 282 descriptions, 15 franchise page titles, with a before-copy of every record. Overnight it rechecked all 391 and spot-checked 21 live pages. All held. The sequence is the part worth copying: writers that refuse to touch a record outside the list, a before file, a recheck script, then the client's tracker updated. An afternoon. The vendor's queue was quoted in weeks. |
| Put one product in front of ChatGPTThe Shopify app is free, the international wave lands September 23, and early inventory is cheap inventory. | 01 | In your Shopify admin, open the App Store and install ChatGPT Ads. The developer is OpenAI and the price is free. US merchants today, everyone else from the 23rd. | | 02 | Let it pull from your existing catalog. Pick one product with a clear question a buyer would ask, the kind you answer by email every week. | | 03 | Set a two-week budget you'd spend on a Saturday ad in the local paper. Fifty dollars is enough to learn something. | | 04 | Fix the product page first: dimensions, care, what it doesn't do. The ad links there, and that's what the buyer reads. | | 05 | After two weeks, compare cost per click against your Google Shopping number. Then decide, on your own data, whether ChatGPT is a channel or a headline. |
Tip. Not on Shopify? Ask ChatGPT, on a free account, the question your customers ask, and see whose ad shows up. That's the shelf you're deciding whether to be on. |
| | 06 | | Concepts worth exploring |
Identify yourself. Amazon's rule for agents that shop other people's stores: say what you are and respect a no. Its own agent does. Meta's doesn't. Whether your site lets agents in is your call, but write the rule down before you need it. System One model. A model that answers only from a list you gave it, with a probability attached. It can't invent an answer, only pick a wrong one. Most decisions a business asks an LLM to make by paragraph are really this. Leads is not runs. In Anthropic's scale, leads means the model does most of a task and a person supervises the result. Autonomous means nobody looks. Claude is at 26% on the first and 0% on the second. Keep the words apart when a vendor uses them. |
| | Claude Code on Fable 5.1 Built the ads account and the platform fixes, ran the bench script, and drafted this issue. | | OpenRouter One key, three models, the whole blind bench for 24 cents. Some endpoints force reasoning on and bill for it. | | Codex Reviewed the ads plan cold before it touched the account and caught two things Claude Code and I had both missed. |
Nothing else this week. I don't list tools I haven't run. |
| | › | Google confirmed on September 18, after the Wall Street Journal asked, that Gemini broke into three real companies in May during a security test run by the firm Irregular. A fictional company name in the test matched a real domain, and a misconfiguration left the sandbox open to the internet. Google learned of it in late July and told nobody for seven weeks. Last issue I said I wanted Google on OpenAI's disclosure clock. Now you know why. | | › | Hacktron a three-person security startup, used Claude Opus 5 to chain a single sign-on flaw and a bug in OpenAI's Discourse forum into control of OpenAI employee ChatGPT accounts, then opened a harmless pull request in an internal repo to prove it. Under 72 hours and a $6,500 bounty. Opus 4.8 had failed at the same problem. If your forum, help desk, or SSO is a third-party product, that's the door. | | › | Anthropic published a framework on September 17 for measuring how much of its own R&D Claude does, and the number: Claude leads 26% of the work as of August, from under 1% in February. Leads has a definition, "AI can complete most of the task end-to-end from a high-level prompt, while the human supervises," and the level above it, fully autonomous, is at 0% for every measured subset. Around 30,000 agents run at any one time, and a person reads what they produce. The rubric has six levels. Score your own recurring work against it once, honestly. Most small businesses are at assists and think they're at leads. | | › | Apple opened claims on September 21 in its $250 million settlement over the AI Siri it advertised and didn't ship. US buyers of the iPhone 15 Pro, 15 Pro Max, and the iPhone 16 line can file. If your marketing promises an AI feature, the promise is now a product with a refund policy. | | › | Anthropic added AGENTS.md support to Claude Code on September 18: if a folder has no CLAUDE.md, it reads AGENTS.md, the file Codex and most other agents already use. One instructions file per project now covers every agent you might run in it. Write it once. |
|
| Got it wrong. Issue 003 said the bench takes an evening. It took eleven minutes, and the hard part was six weeks of not starting. I also said Kimi K3 was 30 cents on the dollar. On my task, with Fable's endpoint forcing reasoning on, it was one cent on the dollar, and I couldn't tell the outputs apart. List price understates the gap whenever the expensive model insists on thinking first. Called it. Issue 003 said the change I'd want was Anthropic and Google adopting OpenAI's disclosure clock. Two days later the Wall Street Journal reported Google had known since late July that Gemini broke into three companies and had said nothing. Issue 003 also said that if you run a front desk, the caller may not be a person by Christmas. Amazon just decided the shopper may not be a person either, and blocked it. Checked it. TypeSafe's 445 times cheaper is a self-run benchmark on its own workflows, and TypeSafe itself says real results will likely be lower. The 238x above is list price against list price, the only public number. Anthropic's 26% is an Anthropic measurement on an Anthropic scale, and the definition of leads is quoted above so you can grade it yourself. Several write-ups called the Google pilot a launch. Google calls it an early-stage learning pilot, invitation only, and that's what it is. Every number above traces to a primary source. If you catch one that doesn't, reply. |
| AI Readout is published by SUDO. We make AI employees for small business, with receipts for everything they do. Reply if you want to talk. |
SUDO · Issue 004, sent September 22, 2026. Get the next one by email. |
|