Where the workflow shifted
LangSmith evaluation puts datasets, evaluators, and experiment results into product iteration.
Vertical agents should not rely on demo prompts. Wrong answers, timeouts, tool failures, and handoffs should become repeatable evaluation cases.
Tool names are not outcomes
The signal matters when it clarifies a real service task, deliverable, and acceptance rule, not when it only shows a demo.
Check permissions and failure
- Turn five recent failures into input, expected result, scoring rule, and review date
- Keep the test narrow: one service scenario with clear inputs, deliverables, acceptance rules, and human review
What still needs proof
Without eval sets, model upgrades do not prove quality improved. Keep the original source open so the announcement, the evidence, and this site's interpretation stay separate.