Evaluation sets should come from real failures

Failure-derived evals have commercial value.

Useful for: Vertical AI services, support agents, sales agents

LangSmith visual for agent evaluation, datasets, and experiment review
Image source: LangSmith.

Where the workflow shifted

LangSmith evaluation puts datasets, evaluators, and experiment results into product iteration.

Vertical agents should not rely on demo prompts. Wrong answers, timeouts, tool failures, and handoffs should become repeatable evaluation cases.

Tool names are not outcomes

The signal matters when it clarifies a real service task, deliverable, and acceptance rule, not when it only shows a demo.

Check permissions and failure

  • Turn five recent failures into input, expected result, scoring rule, and review date
  • Keep the test narrow: one service scenario with clear inputs, deliverables, acceptance rules, and human review

What still needs proof

Without eval sets, model upgrades do not prove quality improved. Keep the original source open so the announcement, the evidence, and this site's interpretation stay separate.

AI agent evaluationevalsvertical AI SaaS