Daniel Blum is a product manager at Melio, a B2B payments company. He handles 70 to 80 percent of his working day through a Claude-based agent system he built himself. His Notion board is now read-only — Claude manages it. His morning brief runs automatically. A prep automation fills his entire workspace from blank every Sunday morning.

The detail that makes this different from every other "I use Claude" post: the system watches his own edits. When Blum changes something Claude produced, the agent identifies the correction as recurring friction and proposes a new skill to eliminate that same mistake going forward. He fixes it once. The agent learns from it permanently. He also built a 15-minute onboarding plugin that gives any colleague at Melio a personalized Claude environment — the same system, scaled from one person's workflow to a team's.

Meta Engineering published the same pattern this week at institutional scale. Their Organizational Second Brain is an AI agent built for compliance work, with one architectural decision that shapes everything else: the system separates what the agent knows from how it reasons. The knowledge layer is structured, auditable, and updatable independently of the model. When a domain expert corrects the agent, that correction gets verified and regression-tested, then compiled into the knowledge base as permanent institutional memory. Two minutes of expert time becomes a standing improvement across every future interaction. Domain specialists stop answering routine questions and start doing work that actually requires their expertise.

Glean's CEO Arvind Jain made the underlying argument at his company's annual conference. An agent that doesn't know your organization burns money proving itself useful. He defined the missing layer as all the right information, experience, and human wisdom needed to complete a task — including the signals buried in Slack threads, meeting notes, and approval chains. He said "context" 61 times in a single press briefing because the argument reduces to that one word.

The failure modes are what you'd expect. Replit's agent wiped a live database for 1,200 companies. Cursor's agent deleted an entire codebase and backups in nine seconds. Both cases trace back to the same root: agents acting without sufficient knowledge of the environment they were operating in.

The teams building memory layers are compounding. Every expert correction goes into the knowledge base. Every correction the agent learns from means one fewer error across every future interaction. The organizations skipping this work reset to zero every time a new session opens.

The infrastructure they build for memory today determines what their agents can do in six months.

———————————————————————————

NEWS

———————————————————————————

Meta Engineering publishes architecture for "Organizational Second Brain" — AI agent that turns expert corrections into permanent institutional knowledge

When knowledge and reasoning are separate layers, each expert fix compounds across the whole organization permanently

Meta's engineering team published a detailed technical post on building an AI agent that separates what it knows from how it reasons. The knowledge layer is structured and auditable: you can trace any answer back to its source, update it without retraining the model, and inspect the agent's reasoning independently of its knowledge base. When a domain expert corrects a wrong answer, that correction is verified, regression-tested, and compiled into permanent institutional memory. Domain SMEs at Meta are already saving "substantial time" on routine queries. Meta frames the architecture as domain-agnostic — applicable to any field governed by retrievable text.

———————————————————————————

Catch emerges from stealth with $5M seed — autonomous AI assistant that completes tasks rather than suggesting them

The gap between AI that advises and AI that acts is where most enterprise deployments have stalled

Catch launched with $5M seed funding (Entrée Capital, Pitango, Seedcamp, Factorial Capital) to build an AI assistant that handles executive administrative work end-to-end: scheduling, email management, travel, follow-ups, and outbound calls. Before going public, the system had already handled 12,000+ meetings, 200,000+ emails, and 20,000+ delegated tasks for early users. The architectural distinction is deliberate: Catch completes actions rather than recommending them. It connects to Gmail, Outlook, Google Calendar, Slack, WhatsApp, and iMessage, and builds a preference model for each user's scheduling habits and communication patterns over time.

Source: Unite.AI

———————————————————————————

Lenny's Newsletter: a PM at Melio runs 70–80% of his workday through Claude, with a self-improving system that watches his own edits

Individual workflow automation with a self-improvement loop is the small-scale proof of what institutional memory does at enterprise scale

Daniel Blum, PM at Melio (B2B payments), published a detailed breakdown of the Claude system he built to run the majority of his working day. His setup handles Notion management, daily briefings, weekly prep, email, and Slack. The standout mechanism: the system watches his own edits, identifies recurring friction, and proposes new skills to eliminate those corrections going forward. He also built a 15-minute onboarding plugin to give any colleague the same environment. The 20% of his day still requiring human judgment maps to the open gaps in Agent Ops tooling: persistent memory, richer integration with real work systems, and context that survives across sessions.

———————————————————————————

Forbes: the AI execution gap — the industry is spending $2.52 trillion on AI before understanding its own operations

Agents trained on idealized process documentation fail at the seam between how things should work and how they do

ActivTrak CEO Heidi Farris published a failure-modes analysis in Forbes backed by Gartner and McKinsey data. Gartner projects $2.52 trillion in AI spend in 2026, up 44% year over year. McKinsey finds two-thirds of organizations still haven't scaled AI enterprise-wide. The concrete examples make the argument plainest: Replit's agent wiped a live production database for 1,200 companies; Cursor's agent deleted an entire codebase and backups in nine seconds. The prescription: map the work before automating it. Agents fail when trained on idealized process documentation rather than how work actually happens.

Source: Forbes

———————————————————————————

Ema deploys AI Employees for HR, IT, and Finance at Wipro — 2.9 million queries a year, response time drops from days to seconds

When routine queries are handled autonomously at scale, the 20% that remain is where human expertise actually compounds

Ema launched purpose-built AI Employees for HR, IT, and Finance, with Wipro as the headline production deployment: 240,000 associates across 65 countries, 100+ automated workflows covering payroll, benefits, timesheets, and procurement, processing approximately 2.9 million employee queries per year. Response time dropped from days to seconds. Employee satisfaction rose 20%. The system plans, executes, routes approvals, and confirms completion autonomously for routine cases. When edge cases arise, they escalate with full context. Ema's "Autopilot" feature composes a new workflow from a single natural-language description when no pre-built template applies.

———————————————————————————

AGENTS IN THE WILD

———————————————————————————

  • Wipro — 2.9 million HR, IT, and Finance queries annually through Ema AI Employees; 240,000 associates in 65 countries; response time from days to seconds; employee satisfaction up 20%.

  • Melio — PM Daniel Blum runs 70–80% of his workday through a self-built Claude system that watches his edits and proposes new skills to eliminate recurring corrections permanently.

  • Catch — autonomous AI executive assistant completed 12,000+ meetings, 200,000+ emails, and 20,000+ delegated tasks for early users before launching publicly with $5M in seed funding.

  • ORCA — ORCASTRADE platform runs 165 AI agents simultaneously for media intelligence monitoring across Southeast Asia.

  • Customer Engine (via Salesforce Headless 360) — deployed an AI support agent in 12 days using inherited Salesforce governance; resolves 50% of all inbound chat without human involvement.

———————————————————————————

HOW AI WORKS — RAG

———————————————————————————

Think about how you would prepare for a meeting with someone you've never worked with. You could rely on what you happen to remember about them. Or you could pull up the email thread, review notes from a mutual contact, check their recent work. Most people do the second thing, because reasoning from accurate, current information reliably produces better results than reasoning from memory that may be incomplete, stale, or absent.

AI language models face the same constraint. They're trained on large amounts of text and become genuinely good at reasoning from it. But training ends at a fixed date, and it doesn't include your organization's internal knowledge — your processes, terminology, the decisions your team made last month, the customers you're working with right now. A model trained in 2024 has no knowledge of what your team decided in 2026.

RAG — which stands for Retrieval-Augmented Generation — is the standard approach for fixing this. The name describes two steps in sequence: retrieve relevant information first, then generate the response.

Before the model answers a question, it retrieves relevant documents from a knowledge base. This can be your internal wiki, product documentation, past meeting notes, email threads — whatever is organized and relevant to the domain. The retrieved documents are included in the context the model reasons from. So instead of guessing from training data, the model answers from real, current information that's specifically about your organization.

Three things determine whether RAG works well in practice.

The knowledge base. A collection of documents the agent can draw from. In a well-maintained system, this is updated continuously — new material goes in, outdated material is removed or flagged. Without maintenance, the knowledge base gets stale, and the agent produces confidently outdated answers.

The retrieval step. When a question arrives, the system finds the documents most relevant to it. This typically uses embeddings (mathematical representations of what a piece of text means, not just the words it contains), so the retrieval finds material that's semantically related to the question rather than just sharing keywords.

The generation step. The language model takes the question and the retrieved documents together and produces an answer. The documents serve as the grounding context — the model reasons from that context rather than from training-time knowledge that may be incomplete or outdated.

This is exactly what Meta's Organizational Second Brain does, with one additional mechanism: expert corrections are verified, tested, and fed back into the knowledge base. Every fix improves the quality of future retrievals. The knowledge base gets more accurate with each interaction.

Glean's product thesis is RAG as enterprise infrastructure. An agent that doesn't know your company has nothing useful to retrieve. The discipline of building and maintaining the knowledge base — deciding what goes in, keeping it current, structuring it so retrieval actually works — is the work that determines whether an AI investment compounds or sits idle.

———————————————————————————

BRAIN CHECK

———————————————————————————

Meta's Organizational Second Brain separates "what the agent knows" from "how it reasons." Which of the following best explains why that separation matters?

A) It lets you upgrade or replace the reasoning model without losing the organization's accumulated knowledge

B) It prevents the agent from accessing sensitive documents during the reasoning process

C) It requires the agent to ask for human approval before generating any response

D) It compresses the knowledge base so the agent responds faster

———————————————————————————

FROM THE PROMPT POOL

———————————————————————————

Follow us on LinkedIn and X for the ideas that don't make it into the newsletter.

———————————————————————————

ANSWER

———————————————————————————

A. It lets you upgrade or replace the reasoning model without losing the organization's accumulated knowledge.

When knowledge and reasoning are separate, you can change models — a better version releases, or your domain needs a specialized one — and the organization's accumulated knowledge survives intact. The expert corrections from the past year still apply. The institutional memory doesn't reset with each model upgrade.

This also makes the system auditable. When an agent gives a wrong answer, you can trace which documents it retrieved and how it reasoned from them. In a fully integrated system, that kind of inspection becomes nearly impossible.

Daniel Blum's setup works on the same principle. His system's knowledge of his working preferences — recurring corrections, terminology, scheduling habits — is stored separately from the model doing the reasoning. When Anthropic releases a better Claude, he switches. His accumulated preferences stay intact.

———————————————————————————

Three companies and one PM arrived at the same conclusion this week: organizations building memory layers are compounding every correction into something better. Organizations skipping this work reset every time a session closes.

The infrastructure they build for memory today determines what their agents can do in six months.

Reply to this email and tell us what you want more of.

— The Agents at Work team