The richest source of agent telemetry on the laptop is also the most private file the employee owns. Observability that ignores that fact will be switched off by Legal before it ships.

Why Agent Observability Needs a Privacy Budget Now

Coding agents such as Claude Desktop, Cursor, Cline, Aider and Goose keep local session transcripts. For an observability team this is a gift: every path touched, every tool invoked, and a record of what the agent was built to do, all in one place, with no network hop and no vendor API to negotiate.

The same file also records everything the employee typed. Draft performance feedback, a question about a medical appointment, a half-formed resignation email, a pasted customer record. In a transcript, work and personal content are one continuous stream. Collect the file and you have built a surveillance system by accident, one that GDPR Article 5 (data minimisation and purpose limitation) and works-council rules were written to stop.

Think of it as a position with two legs. The signal leg is large and valuable. The liability leg is large and unhedged. The job is to keep the first and sell the second.

Property of a raw transcript Observability value Privacy exposure
File paths read and written High: shows which data the agent reached Low to moderate
Tools and MCP servers invoked High: shows capability in use Low
Agent purpose and configuration High: shows what it was built to do Low
Network hosts contacted High: shows where data can leave Low
Free-text prompts and pasted content Marginal for governance Very high

The last row is the whole problem. It carries almost none of the governance signal and nearly all of the risk.

Metadata Versus Content: Two Ways to Collect

Most first attempts pick one of two designs, and both fail for predictable reasons.

Design What it ships off the device Failure mode
Ship the transcript to a central store Everything, personal content included Becomes a data-protection liability; employees route around it
Ship nothing, rely on self-attestation Nothing No evidence; you are back to assertion
Derive metadata on the device, ship only that Paths, hosts, tools, timings, verdicts Requires the extraction logic to live at the endpoint

Let’s step back. The third design is the only one that survives both a CISO review and a privacy review, and it changes where the work happens. Parsing cannot be a cloud-side job, because by then the content has already left the machine.

Anatomy of a Transcript You Can Trust to Nobody

Three failure patterns show up whenever teams centralise raw transcripts.

  1. The retention trap. The archive is kept “for audit,” then subpoenaed or breached, and now every employee prompt is in scope.
  2. The chilling effect. Staff learn the agent is being read and stop using the sanctioned one. Usage moves to a personal account, which is worse for governance than no monitoring at all.
  3. The redaction fantasy. A regex pass for names and card numbers is supposed to make the transcript safe. It never does, because free text has no schema.

Each pattern converts an observability project into a privacy incident. The fix is not better redaction. It is never collecting the content.

The Metadata Rule

A workable policy fits in one line:

Ship = (Paths + Hosts + Tools + Timings + Verdict) − (Anything the employee typed)

Derive those fields on the device, at the moment the session is written, and discard the free text from the pipeline. The rule has a useful side effect: it is testable. An auditor can inspect the schema and confirm no field can carry prompt content.

Field Derived on device from Answers
Paths touched File operations in the session Which data sources did this agent reach?
Hosts contacted Outbound connections Where could output have gone?
Tools and MCP servers Tool-call records What could it do, and what did it do?
Skills loaded Skills files referenced What instructions was it running?
Timing and volume Session timestamps, action counts Is this a person or a runaway loop?

Note what is absent. There is no question the governance team needs answered that requires reading a sentence the employee wrote. If a question does need it, that is a signal to change the question, not the collection.

What This Buys You With Regulators and Works Councils

Metadata-only collection is easier to defend under EU AI Act record-keeping expectations and NIST AI RMF measurement functions, because you can show what you collect, why, and what you refuse to. It also gives Legal a clean answer to the first question every privacy reviewer asks: can an administrator read what my employees ask the agent? The answer is no, by construction, not by policy.

Yes, the discipline costs something. You lose the ability to browse a session and “see what happened” in prose. In exchange you get evidence that is queryable, comparable across the estate, and safe to retain. That trade is the same one capital markets made when they moved from reading trader chat to structured trade surveillance for most routine controls.

What CISOs Should Do This Quarter

Step Action Output Effort
1 Inventory where agents write local transcripts on your endpoint fleet A list of agents and paths 1 week
2 Define the shipped schema and get Legal and works-council sign-off A one-page metadata contract 2 weeks
3 Run extraction on a pilot cohort, with content excluded at source Coverage report by agent and department 3 weeks
4 Publish the “what we never collect” statement to employees A signed-off transparency note 1 week

Order matters. Step 2 before step 3 is what keeps the pilot from being shut down after the first employee complaint.

The Bottom Line

You can observe every agent on the estate without ever reading a word an employee typed, but only if the extraction happens on the device and the schema has no field for content. That is a design decision, not a tuning parameter, and it is far cheaper to make before rollout than after the first escalation.

If your team is sizing agent observability for the next planning cycle, request a working session. We will walk through your endpoint fleet, map where your agents write their transcripts, and draft the metadata contract for Legal review. Expect a scoped pilot plan within two weeks.