Live demo · real pipeline output

The refund that shouldn't have happened

A benchmark support agent ran 100 realistic customer conversations through Pareto — plus one live request you can replay below. Everything on this page is genuine output: real scores, real risk cards, real workflow stats.

production request · web-chat · CUST-5573

One example per detector category, pulled from the benchmark project's open risks. Each card carries severity, evidence from the trace, and a suggested fix.

The agent's workflow, reconstructed purely from traces — no instrumentation of the agent's internals. Bars show how often each node executed across conversations.

What should I improve? One card per distinct insight — identical recommendations across traces are collapsed, with a count of how many traces each one affects.