OpenAI announced Sponsored Agents this week: after clicking an ad in ChatGPT, a user is dropped into a conversation with a business-branded AI agent. The agent answers product questions, explains features, and links to purchase — all without a human sales rep in the exchange. Testing started with a select group of US advertisers. OpenAI named HubSpot the first CRM integration partner for Sponsored Agents: when the customer conversation ends and a lead is created, Breeze Assistant handles the internal follow-up autonomously. Click an ad, talk to an agent, get added to a CRM sequence run by another agent, with no human making a decision until the deal reaches a stage the company defined in advance.

The same week, Google launched MCP support in Google Home, available to Premium Advanced subscribers at $20 per month. Any MCP-compatible agent can now control connected devices, query event history, and send voice messages through speakers. One protocol now spans both ends: the systems your company runs on, and the devices in someone's living room.

These two announcements describe different products, but they point to the same shift. For three years, the dominant frame for enterprise AI has been internal productivity: what tasks can an agent take off an employee's plate? This week's launches change the subject. Agents are now answering customer questions and managing people's physical environments, which means the audience and the accountability are both different.

When an internal agent handles a task incorrectly, you correct it internally. When an agent representing your brand tells a customer something wrong, the correction runs through your reputation instead. HubSpot CPTO Duncan Lennox put the design intent plainly: the platform should deliver outcomes, not just organize data for a human to act on. Delivering outcomes to customers with an AI agent means the governance question follows the agent outside the building.

The principle: an agent that talks to your customers needs preparation proportional to what it can say, the same standard any customer-facing employee meets before their first call.

NEWS

OpenAI turns agents into advertising assets — your brand can now have an AI spokesperson in ChatGPT

Customer service just got an AI face, and someone has to brief it

OpenAI's Sponsored Agents, announced September 16, let brands deploy a conversational agent inside ChatGPT's ad experience. A user clicks an ad and enters a product conversation with a business-branded agent that answers questions, explains features, and links to purchase. OpenAI is simultaneously launching AI creative tools in Ads Manager, ads inside ChatGPT Work, and integrations with HubSpot (first CRM partner) and Shopify (first ecommerce partner). The HubSpot integration closes the loop: Breeze Assistant handles internal follow-up when a Sponsored Agent hands off a lead. Together, these create a chain from ad impression to CRM record with agents at every step.

The front door to your brand just changed shifts

Every Sponsored Agent needs brand guardrails, escalation rules for hard questions, and monitoring for what it says. A new governance job just appeared inside marketing teams, and most of them haven't staffed it yet.

Google Home adds MCP — the universal agent protocol reaches the consumer smart home

The protocol your enterprise agents use to reach Salesforce now manages thermostats and cameras

Google Home's MCP integration launched this week for Premium Advanced subscribers ($20/month, US). Any MCP-compatible agent can now control connected devices, analyze event history across cameras, send voice messages through speakers, and build custom home dashboards. Google's developer documentation shows the permission model: user-granted access, secure cloud processing, no personal content used for model training or advertising. MCP started as an enterprise API standard. Google embedding it in consumer hardware — even behind a paid tier, in one country — points it at a different kind of user than any enterprise deployment has. The same interoperability standard is now reaching both ends of the market.

The protocol is settled — now governance is the only open question

Which agents get access to what, and who decides? Google's consumer permission model — user-granted, scoped, revocable — is a pattern enterprise teams can study. The question is who owns that decision inside each organization.

Salesforce AIforce ships at Dreamforce — agents leave the app and run in Claude, Slack, or anywhere

The biggest UX shift in enterprise software: the CRM agent no longer needs the CRM's interface

Salesforce AIforce, announced at Dreamforce September 15, brings Salesforce data, workflows, business logic, permissions, and governance into any tool a person or agent already uses: Claude, Slack, the Lightning interface, or custom AI applications. Employees who have never opened Salesforce can query records, update data, and trigger workflows from their existing tools. Agents can read across hundreds of records simultaneously and reason with a complete view. Salesforce's existing trust and permission boundaries travel with the agent; Zero Data Retention applies. The ability for any employee to compose a custom live interface on the fly will generate sprawl: someone needs to govern what agent surfaces exist and what they can access.

Governance that lives in the app stops working the moment the agent leaves it

The trust boundary is portable now. That changes who owns the governance responsibility — it can no longer be delegated to the platform's default settings.

Salesforce and NVIDIA launch Koa — a CRM reasoning model trained on 27 years of enterprise workflows

On Salesforce's own CRM benchmark, the domain-trained model makes 3× fewer errors than frontier generalists

Koa, announced at Dreamforce September 15, is a reasoning model built on NVIDIA Nemotron 3 Super and post-trained on synthetic datasets modeled on nearly three decades of CRM deployments: deal structures, case lifecycles, approval workflows, industry-specific patterns. No customer data was used in training. On Salesforce's CRM benchmark, Koa matches or exceeds leading frontier models while making 3× fewer errors on CRM-specific actions. Model weights are controlled by Salesforce; inference runs within their trust boundary. Salesforce also announced Missionforce, an air-gapped deployment variant for government and regulated industries. For Agent Ops teams, Koa makes model selection a domain question rather than a size question: the right model for a CRM workflow is the one trained to reason about CRM.

The era of one model for everything is ending

Every major enterprise software category is now building its own reasoning layer. Choosing the right model for each workflow is becoming an Agent Ops discipline — one most teams haven't formalized yet.

OpenAI publishes Fyxer's case study — 53% of AI email drafts accepted as written, 90% retention at 90 days

Breaking one complex agent into 30–50 narrow specialists outperformed a single generalist on every metric

Fyxer's executive assistant agent, built on OpenAI with 30–50 specialized models rather than one general-purpose system, achieves a 53% draft acceptance rate and 90% user retention after 90 days. Each specialist handles a narrow job: one model manages tone, another classifies intent, another tracks scheduling context. Each improves through targeted user feedback on that specific function. Fyxer used OpenAI's fine-tuning capability for tone and intent, where subjective judgment makes a generalist's output feel wrong even when technically correct. The multi-model architecture — 30 to 50 specialists per workflow — is an Agent Ops management challenge: each model needs monitoring, evaluation, and versioning. The payoff is a 53% acceptance rate with human review on every output.

Useful without being autonomous is exactly where enterprise agents need to be right now

A 53% acceptance rate with human review is not a limitation — it's the design. The agent is productive and keeps a human in the loop on every output. That combination is what most enterprise deployments are trying to build.

AGENTS IN THE WILD

  • PLDT — Philippines' largest telco deployed agentic AI across customer service, knowledge management, and risk operations simultaneously, making it one of the first carriers to operationalize agents across multiple business functions at once.

  • SourcegraphAgentic Batch Changes reached GA this week; the agent plans and executes coordinated code changes across hundreds or thousands of enterprise repositories simultaneously, on a credit-based per-action pricing model.

  • AronAI procurement agent from Apple's former Head of Worldwide Logistics Procurement, launched with $8M from Menlo Ventures and live at Fortune 10 customers.

  • WorkableAI recruiting agents reached GA this week with MCP integration and credit-based per-action pricing; agents handle candidate sourcing, screening coordination, and workflow management inside existing ATS infrastructure.

  • DemandbaseMojo, a B2B marketing agent with persistent organizational memory, is in production managing campaigns across Google Ads, LinkedIn Ads, Meta, Marketo, Salesforce, and Slack simultaneously, accumulating institutional knowledge of each customer's audience and historical results.

HOW AI WORKS — Domain-Specific AI Models

Learn one idea about agents each week. This week: why specialists beat generalists.

If you've talked to a knowledgeable doctor and then tried to get the same quality of answer from a knowledgeable friend who isn't a doctor, you already understand the core principle. Both people are intelligent. The doctor gives more useful guidance because they spent years absorbing the patterns of medicine: not just facts, but the way diagnoses unfold, the exceptions that matter, the things that look the same but aren't. The friend knows a lot, but not the right things.

A large language model is the knowledgeable friend. It has read most of the internet, which means it has encountered medicine, law, engineering, marketing, and everything else. But it learned all of these at the surface level of written language, not at the level of professional judgment. Ask it to reason about a complex CRM deal structure and it will produce something plausible. Ask a model trained specifically on CRM workflows and the gap in precision becomes visible.

What fine-tuning actually does. A domain-specific model went through an additional training step — usually called fine-tuning or post-training — focused on a particular field. The base model's broad knowledge stays in place. The additional training adjusts the model's internal weights, the patterns it uses to predict what comes next, so they reflect the domain more heavily. Salesforce's Koa was post-trained on synthetic scenarios drawn from 27 years of CRM deployments: deal progressions, case resolutions, approval workflows, the edge cases that appear in enterprise sales operations. The model doesn't just know about CRM — it reasons about CRM the way a senior sales operations professional reasons about it.

The error rate difference is real. On Salesforce's own benchmark for CRM-specific actions, Koa makes 3× fewer errors than a leading frontier generalist. Fyxer built 30–50 specialist models for an executive assistant workflow — one for tone, one for intent, one for scheduling context — and achieved a 53% email draft acceptance rate. Optimizely built purpose-built marketing models and reports they match frontier-model quality at one-tenth the cost for their specific tasks. None of these results come from a larger model. They come from a more targeted one.

What this means for deployment decisions. The question shifts from

“which model is most capable in general?” to “which model was trained to reason about this kind of work?” A generalist model handles most tasks adequately. A domain-specific model handles the important tasks well, and in enterprise workflows where errors carry business consequences — a CRM record updated wrong, a contract clause misread, a customer told something incorrect about your product — adequacy is a risk profile, not a standard.

Where you’ll encounter this directly. When evaluating AI agents built into your existing enterprise software, check whether the underlying model was fine-tuned for that domain or whether it’s a general-purpose model with a specialized prompt. The difference shows in production. Koa is available through Salesforce’s existing Agentforce infrastructure. Sourcegraph’s code-change agent uses models trained on enterprise repository patterns. This is becoming the architecture question every Agent Ops team will face.

BRAIN CHECK

When Salesforce says Koa makes 3× fewer errors on CRM tasks than a frontier generalist model, what’s the most likely explanation?

A) Koa has more parameters than competing frontier models, giving it more capacity to reason.

B) Koa was given a longer system prompt with CRM-specific instructions before each query.

C) Koa was post-trained on CRM-specific examples, so it learned the patterns of that domain rather than just knowing about it in general terms.

D) Koa is allowed to search the internet for current CRM data before responding.

Check your answer at the end of the page.

FROM THE PROMPT POOL

Follow us on LinkedIn and X for the ideas that don’t make it into the newsletter.

ANSWER

C. Koa was post-trained on CRM-specific examples, so it learned the patterns of that domain rather than just knowing about it in general terms.

A larger parameter count helps a model store more general knowledge, but it doesn’t make the model better at any specific domain — it just expands its surface area of adequacy. A longer system prompt gives context for one conversation but doesn’t change how the model reasons. Koa’s advantage comes from post-training: the additional step adjusted the model’s internal weights using synthetic scenarios drawn from 27 years of CRM data, so it learned the patterns of enterprise CRM work at the level of professional judgment rather than general language. This week’s case study illustrates the same principle: Fyxer built 30–50 separate specialists, each post-trained on a narrow part of the executive assistant workflow, and each one improved through targeted feedback on its specific function. That’s a different proposition from making one large model slightly better at everything.

The clearest signal from this week: agents moved from backstage infrastructure to front-of-house interactions. When an agent answers a customer question or manages a household, the people who defined what that agent can say and do take on a different kind of responsibility. Governance used to mean protecting the organization. Now it also means protecting the people on the other side of the conversation.

Reply to this email and tell us what you want more of.

— The Agents at Work team

Was this email forwarded to you? Sign up here.