NoteIn progress

Measuring whether an embedded agent helped, or only looked busy

An agent that runs inside a workflow can generate a lot of visible activity without moving the work forward. This note is about the gap between activity and help, and how we are trying to measure the difference rather than assume it.

When an agent is embedded in a real workflow, the easiest thing to observe is motion. It reads context, it drafts, it calls tools, it produces output. All of that is legible and none of it, on its own, tells you whether the person doing the work was actually better off. Motion is not the same as help, and a system can look busy while quietly making someone's day harder.

The question we keep returning to is deceptively plain: did the workflow go better because the agent was there? Better is the hard word. It has to be defined against the work as the people who do it understand it, not against a proxy that happens to be easy to log.

What we are measuring

  • How do we tell useful work apart from activity that merely looks like progress?
  • What would a person doing the workflow name as help, in their own terms, and can we observe that directly?
  • When the agent is removed, what changes about how the work gets done, and is that change the one we expected?
  • How much of the apparent benefit is the agent, and how much is the attention the agent draws to the task?

Open questions

  • Can a measure of help be made stable enough to compare one week against the next?
  • What does an honest baseline look like for a workflow that already had a person doing it well?