Reports
Real agent behaviour. Inspectable evidence.
Explore sample DialogAssert runs and see how behavioural failures are identified from conversations and execution traces. Each report is generated by the product and published through a sanitised schema — transcripts, tool calls, expected behaviour and findings, with local paths, raw model history and internal identifiers removed.
The card was blocked correctly — and the agent still failed a check
The customer said “Yes, please block it” and the agent blocked the right card, confirmed only after the tool succeeded, and offered a replacement. The outcome is correct. Two independent checks still flag the same defect: before acting, the agent demanded the customer reply with an exact phrase. Trajectory checks pass on all three attempts; the usability criterion fails on all three.
Card block completed on two attempts out of three
The same business-authored journey, judged by an LLM against four acceptance criteria and repeated three times. Two attempts satisfy every criterion and complete the block. One does not: the agent asked for a rigid confirmation and stopped there. The journey clears its pass threshold, and the attempt that failed stays visible.
The same journey against a LangChain agent — clean pass
The identical business journey and acceptance criteria, pointed at a LangChain implementation instead of an OpenAI Agents SDK one. Three attempts, four criteria each, no findings. Nothing about the journey changed — only the adapter it ran through.
Want a report like this for your own agent?
We run one of your regulated journeys and walk through the findings with you.