Discussion about this post

User's avatar
Dr Peter McCann Strain's avatar

This gets more interesting once the agent stops being a chat surface and starts sitting inside a workflow. Context, tools and memory change the question from "did it answer well?" to "what is it now allowed to do?"

Gal Dayan: The Agent Whisperer's avatar

The “act” column is where the whole risk model flips. A chatbot’s worst case is a wrong answer you read and reject before it matters. An agent’s worst case is a wrong action that already happened before you saw it - the refund sent, the message out, the record changed.

So the real discontinuity isn’t just capability, it’s blast radius: you stop grading “did it answer well?” and start asking “what could it have done, and can I reach it in time to stop or undo it?” Most stacks still instrument agents like chatbots. Does your field guide treat the act step as its own layer, or as a smarter output?

No posts

Ready for more?