On Tuesday, Dataiku published a risk framework for agentic AI, built on a survey of 800+ global data leaders conducted with Harris Poll. The headline finding: 75% of those leaders say trust in their agent deployments is a concern. The more useful output was what Dataiku did with that uncertainty — they named the five specific ways it materializes.

Privileged access inheritance: an agent inherits the permissions of whoever deployed it, including access it has no business using. Multi-agent drift: a network of coordinating agents gradually diverges from intended behavior through accumulated small deviations, with no single failure to point to. Data poisoning: corrupted or manipulated data changes what an agent retrieves or learns from. Compliance misreporting: an agent logs what it did incorrectly, producing audit trails that don't reflect actual decisions. Goal misalignment: the agent pursues a subtly different objective than intended, often optimizing for something measurable while missing the actual goal.

Any team running agents in production will recognize at least three of those. The value of naming them is that they can now be owned. This failure mode belongs to this team, gets caught by this monitor, gets escalated to this person. Unnamed failures are everyone's problem and no one's responsibility.

The same week, Gartner released its inaugural Emerging Market Quadrant for AI Agents for Marketing. Writer landed in the upper right, named a Market Shaper. The more significant part of that release is the creation of the category itself. When Gartner makes a quadrant, procurement teams write RFPs against it. Marketing AI agents crossed from "early adopter experiment" into "analyst-tracked procurement option" this week. Agent Ops teams will be asked to govern, audit, and monitor whatever those procurement teams buy.

And on Wednesday, ForceEquals launched a product explicitly called Agent Operations. The company's framing: "Control and operations are easy with one agent. They stop being easy the moment there's more than a few." Business teams are currently managing agent complexity with spreadsheets, Slack threads, and tribal knowledge. ForceEquals built the structured alternative. When a startup names its flagship product after your job title, the category has reached a particular kind of maturity.

These three events are sequential. Gartner formalizing a category means adoption is moving from pilots to procurement. Dataiku naming the failure modes means production deployments are failing in documentable, repeatable ways. A startup naming its product "Agent Operations" means there is enough demand to build a company around the discipline of running agents reliably.

The principle: before a technology becomes infrastructure, the failure modes get names. This week, three organizations named them at the same time.

NEWS

Gartner publishes the first Emerging Market Quadrant for AI Agents for Marketing — Writer named a Market Shaper

When Gartner makes a category, the RFPs follow

Gartner released its inaugural Emerging Market Quadrant™ for AI Agents for Marketing — Startup Vendors (July 2026). Writer was named a Market Shaper, with Gartner describing the quadrant leaders as "ushering in a new era of industrialized, autonomous revenue engines, combining enterprise-grade governance, compliance and data security with frontier AI innovation." Writer CEO May Habib: "Forward-thinking marketing leaders aren't asking for another AI tool that helps one person create one more asset." The existence of the quadrant matters more than any individual placement — it means marketing procurement teams now have an analyst-backed framework for evaluating, comparing, and budgeting for marketing AI agents. The Agent Ops question follows immediately: who governs the agents the procurement team selects?

Source: Writer Blog

Dataiku publishes the first structured enterprise risk framework for agentic AI — five failure modes, three governance pillars, 30/90/365-day roadmap

A failure taxonomy changes who owns each failure mode

Dataiku's framework, built on a Harris Poll survey of 800+ global data leaders, names five agentic-specific failure categories: privileged access inheritance, multi-agent drift, data poisoning, compliance misreporting, and goal misalignment. The governance response runs three pillars — visibility, defense, governance — implemented across a phased timeline: 30 days for internal controls, 90 days expanding to third-party risk, 12 months for enterprise-wide rollout. The framework maps to NIST AI RMF and the EU AI Act. The phased structure is the most practically useful part: it gives Agent Ops teams a timeline for building governance incrementally rather than as a single large initiative. Trying to govern everything at once is how governance projects get deprioritized. Starting with 30-day internal controls means something ships.

Source: Dataiku Blog

ForceEquals launches Agent Momentum — a product built around two components called Agent Control and Agent Operations

When a startup names its product after your job title, the category is real

ForceEquals announced GA of Agent Momentum, where Agent Control monitors agent behavior against five trigger types (guardrail breaches, temporal anomalies, semantic ambiguity, strategic signals, and HITL checkpoints) and escalates to the right human. Agent Operations captures those decisions and feeds learnings back into the agent portfolio. The company's diagnosis: "Control and operations are easy with one agent. They stop being easy the moment there's more than a few." The five trigger types map directly to the daily concerns of anyone running agents at scale — something behaved unexpectedly, something took longer than expected, the agent hit an ambiguous instruction, a strategic signal emerged, or a human checkpoint was reached. Agent Momentum formalizes the response workflow for all five.

Source: PR Newswire

Amazon Bedrock AgentCore Payments reaches general availability — agents can now transact autonomously at scale

When agents can spend money, the governance requirements jump an order of magnitude

AWS announced GA of AgentCore Payments, enabling AI agents to "transact safely and autonomously at scale." AWS's own characterization of the trajectory: agents have "evolved from simple chat applications to autonomous, long-running systems that dynamically discover and compose dozens of tools per task without human oversight." The payments capability adds a new dimension to Agent Ops risk. Watching what agents say and write is one governance problem. Watching what they buy — spending limits, approval workflows, audit trails for autonomous financial decisions — is a structurally different one. Agent Ops teams that have built observability for agent actions but not for agent transactions need to extend their governance scope. The window between "agents can spend money" and "someone notices a problem" is where the risk lives.

TRM Labs builds Clara — an internal agent that handles all pre-call and post-call work for customer-facing teams

The pre/post pattern: agent handles the surrounding work, human handles the actual relationship

TRM Labs (crypto compliance infrastructure) detailed how their deployment strategy team built Clara, an agent that prepares customer call briefs, drafts post-call summaries, generates action items, writes follow-up emails, and updates CRM records in minutes. Before Clara, that preparation required digging through notes, reviewing past meetings, and coordinating across a portfolio spanning DeFi startups to traditional financial institutions. Clara was built internally by the team, not purchased as a product, which signals that the tools and skills for internal agent development have reached enough maturity that non-specialist teams can build production-grade agents for their own workflows. The pattern — agent handles everything before and after the human interaction, human handles the actual conversation — is directly replicable across customer success, sales, and consulting.

AGENTS IN THE WILD

  • JetBlue — uses Sprout Social Trellis to generate executive-ready social performance summaries; a report that previously took weeks now takes minutes.

  • Ipsy — runs Trellis for real-time launch monitoring across social channels, with agents surfacing signals as campaigns go live.

  • TRM Labs — Clara prepares every customer call brief and drafts every post-call summary; the human's focus stays on the conversation itself.

  • Whistic — four specialized agents coordinate to handle roughly 90% of vendor risk assessment workflow administration; every risk decision escalates to a human.

  • ABC Legal — moved from scattered individual AI experiments to a governed fleet of specialized agents using Claude Managed Agents; the governance layer was built because ungoverned proliferation made it necessary.

HOW AI WORKS — Agent Identity

Imagine you work at a company with 200 employees. Each person has an email address, a badge that opens certain doors, a job title, and a manager. When something goes wrong — a document is deleted, an email goes to the wrong person — the company can trace who did it, what access they had, and who was accountable for their actions.

Now add another 200 workers who have no email addresses, no badges, no titles, and no managers. They show up, take actions, and disappear. When something goes wrong, there is nothing to trace.

That is the agent identity problem in its simplest form.

An AI agent — a software program that takes actions on behalf of a user or system — can search data, send messages, book meetings, move files, trigger workflows, and now initiate financial transactions. In most organizations today, agents do all of these things without the basic identity infrastructure that every human employee takes for granted.

Giving agents identity means providing each one with four things:

A unique identifier. Every agent gets a name or ID that follows it across every system it touches. When the agent submits a document, the submission record says "Agent X," not "system" or "automated process" or nothing at all.

Scoped permissions. An agent's access is limited to exactly what it needs for its specific tasks — the same principle as a badge that opens three doors, not all of them. This directly addresses "privileged access inheritance," the first failure mode in Dataiku's framework: agents inherit too much access because no one assigned the minimum required.

An audit trail. Every action the agent takes is logged — what it did, when, what data it accessed, and what decision it made. This is how compliance misreporting gets detected. A log that says "marketing campaign executed" is not the same as a log that says "Agent Y ran campaign variant 3 at 14:23 UTC against segment 445, excluded segment 221."

An owner. Every agent has a human or team accountable for its behavior. That owner receives the alert when behavior is unexpected, and is responsible for the response.

Adobe demonstrated this model with the launch of AI Collaborators in Workfront: agents are assigned work exactly like human team members, with explicit permissions, task records, and output posted back for review. Every step is logged. CollectivIQ's Digital Direct Reports go further — agents have names, org-chart positions, and reporting relationships, treated as permanent members of the team.

The right parallel here is HR onboarding. A contractor who shows up with no badge, no contract, and no manager might do excellent work. But when something goes wrong, there is nothing to trace, no one to call, and no clear path to remediation. Agent identity is the layer that makes everything else — monitoring, governance, accountability — possible. You cannot govern what you cannot identify.

BRAIN CHECK

Dataiku's agentic AI risk framework names five specific failure modes. Which of the following is one of them?

A) Model version mismatch — an agent running on a deprecated model produces inconsistent outputs across requests

B) Multi-agent drift — a network of coordinating agents gradually diverges from intended behavior through accumulated small deviations

C) Prompt injection overload — too many simultaneous instructions cause the agent to drop key constraints

D) Orchestration latency — delays between agents in a multi-step workflow cause failures under production load

FROM THE PROMPT POOL

Follow us on LinkedIn and X for the ideas that don't make it into the newsletter.

ANSWER

B. Multi-agent drift.

The other four from Dataiku's framework: privileged access inheritance (an agent inherits excessive permissions from whoever deployed it), data poisoning (corrupted input changes what an agent retrieves or learns), compliance misreporting (an agent logs its own actions incorrectly), and goal misalignment (the agent optimizes for a measurable proxy while missing the actual objective). Multi-agent drift is the hardest to detect because it happens at the fleet level, not the individual agent level. Any single agent's output looks reasonable at a point in time — the problem only surfaces when comparing the fleet's behavior now against its behavior 90 days ago. This is why the Gartner forecast matters: 150,000 agents per Fortune 500 company by 2028, with only 13% of organizations currently governance-ready. At that scale, multi-agent drift stops being an edge case and becomes a baseline operational risk.

The categories are filling in. The failure modes have names. Gartner is tracking the market. A startup named its product after the job. Agent governance moved from "best practice" to "documented discipline" in a single week — not through a single announcement, but through three separate organizations arriving at the same conclusion from different directions.

Reply to this email and tell us what you want more of.

— The Agents at Work team