The New User Your Dashboard Can't See
When a customer's AI negotiates for them, who designs for the human on the other end?
I first noticed it on LinkedIn.
Someone had posted a video of their own AI agent negotiating a bill with a company’s customer service agent while also filing a complaint with a government regulatory body. The poster gave instructions to the agent about the desired outcomes, and the AI agent simultaneously and autonomously interacted with a human agent while filing the complaint.
I am not surprised that the agent existed; I build AI agents myself. We have all been watching these get better for two years. What got me was that the AI agent picked a strategy. It made a call about how to achieve the defined goals, and then went and did just that.
Which means that at the receiving end of that conversation was a human being, talking to a customer’s patient, well-briefed AI agent that was running a plan and representing someone else’s interest, like an attorney. We have spent years designing for a human on the other side of that conversation, the one who is frustrated, or in a hurry, or genuinely wants to resolve a billing issue. That person is still there. But our experiences have not evolved to accommodate their (AI) representative.
It also doesn’t get tired or embarrassed, and it doesn’t take no for an answer. The training, the scripts, and the scorecards at the human end of that conversation were built around a counterpart who does all three.
Then there is a question about the disclosures that’s still unanswered. We’ve mostly settled that people should be told when they’re talking to an AI, or that it can make mistakes. The reverse isn’t settled at all. If a human agent is now talking to someone’s AI, is that human owed the same disclosure? I think so. But who would enforce it? You also can’t disclose what you can’t detect, which is where this stops being philosophical.
How would we even know? None of the usual channels would catch it. This won’t show up on a dashboard, in a test case, or a flag anyone thought to build. It surfaced because a person was pleased enough with what they’d built to show it off in public.
The human agents already know
In the early 2000s, I worked at a Hewlett Packard call center. Back then, if you called because you had an issue with your HP personal desktop computer or a printer, I may have taken that call. During one particular Thanksgiving, HP had launched a new line of PCs that turned out to be faulty. That weekend we were flooded with calls about new computer issues. The guidance we were given was to ask the customer to return it for a full refund. But we, the agents sitting on the floor taking calls, noticed the pattern early and knew about the issue with the computers before the company did.
The people taking the calls notice pattern shifts early, long before it’s a number on a dashboard. They notice the conversations that start going strange. We have all gotten good at spotting the AI slop tells by now. Someone handling chat all day would pick up on the tidier words faster than most of us.
So the knowledge tends to exist, in some form, on the floor. What’s usually missing is a feedback loop to carry it to the people who could act on it.
The gap in the feedback loop
The gap is that often talking to the frontline is treated as research. Something you do once, at the start, to inform a design. A study, a set of interviews, a ride-along that gets written up and then stops. Observation is a one-off by construction, and the advice to go “do more of” it is aimed at a person who may or may not be there next quarter, which, if we’re now designing for a class of user that arrives with a plan and doesn’t announce itself, feels like the thing to get right as part of product discovery.
Evals have become their own discipline in the last two years, with teams and even companies dedicated to watching what these systems produce. That discipline is getting better quickly. What it watches is the output. It doesn’t hear from the person who has been on the receiving end of forty strange conversations this week.
What’s the solution?
The shape of it seems to be:
A standing channel from the frontline. A scheduled conversation where the human agents share with the product team what’s gotten weird lately.
A named owner for the signals nobody is assigned to watch. Public posts, sub-reddits, community forums, the places customers brag about what they built.
Detection treated as a design input. Whether we can tell a bot conversation from a human one changes how we design the evaluations, and right now most of us are finding out by accident.
None of this is exotic. It’s mostly a decision to treat the loop as part of the product, with the same ownership and maintenance as anything else we ship.
The customer’s AI agent in that video was better prepared than most humans who chat in. It knew what it wanted, and it had a plan for getting there. Plenty of people are working on how to design for that now. I am curious to see how the design practice changes as a result.
If you work on these systems, what would it take for the frontline to be where a change like this gets noticed first?
