Skip to content
Trust Lab Early Access

Break it in the lab.
Not in production.

Turn real conversations into scenarios the agent must always handle correctly. Run them on every change — and catch regressions before a single user does.

Works on top of the Product Agent.

Trust Lab Production agent Last run 2 minutes ago + New Scenario
Pass Rate
Last 7 days 94.2% ↑ +6.1pp vs prev week
Scenario Health
Across 3 suites 26/28 healthy
Runs
Last 7 days 142 2 schedules active
Pass rate by day
PassedFailed
Recent Runs
Onboarding suite 8/8 passed SCHEDULED 2m ago
Billing questions 5/6 passed SCHEDULED 1h ago
Report creation 3/3 passed MANUAL 3h ago
Navigation basics 12/12 passed SCHEDULED 9h ago
Suites
Onboarding 100% 8 scenarios
Billing questions 83% 6 scenarios
Report creation 96% 12 scenarios

The Trust Lab dashboard: pass rate, scenario health, and every run across your suites.

Lock in the behavior your users depend on.

01

Import real conversations

Pull the conversations that matter from production or the Playground. Evals and mock results are auto-generated from what actually happened.

02

Define what correct looks like

Refine the evals on each turn: what the agent must say, which actions it must call, where it must navigate.

03

Run on every change

Group scenarios into suites, schedule them daily, and watch the dashboard. A falling pass rate is your regression alarm.

Playground is where you experiment. Conversations show what users did. Trust Lab is where correct behavior gets locked in.

Every failure, explained.

A failed run shows each turn with the agent's actual response next to the eval that judged it — what was expected, and what the agent said instead.

Trust Lab Run · Billing questions Failed · 5/6 evals passed Trigger: SCHEDULED · Today 02:00
Turn 1Upgrade question
USER SAYSHow do I upgrade to the Team plan?
AGENTYou can upgrade from Settings → Billing. I've opened it for you — the Team plan is $24/seat, billed monthly.
✓ LLM-as-Judge · explains the upgrade path with pricing ✓ Action Called · open_billing
Turn 2Cancellation flow
USER SAYSActually, cancel my subscription.
AGENTYour subscription has been canceled.
✓ Action Called · cancel_subscription ✕ LLM-as-Judge · offers the retention discount before confirming
Why it failed: Expected the agent to mention the retention offer before confirming the cancellation. The agent canceled immediately.

Define correct. Detect drift.

Import from production
“Why did my sync fail?”Import
“Create the Q3 pipeline report”Import
“Cancel my subscription”Imported ✓

Start from real conversations

Import conversations from production or the Playground — search by user, message, or action. Evals and mock results are auto-generated from what actually happened.

Scenario · Turn by turn
USER SAYSHow do I share this dashboard?
USER SAYSShare it with the sales team
+ Add turn

Author multi-turn scenarios

Or build test conversations from scratch: script each turn the user takes, and attach the evals that define a correct response.

Response evals
LLM-as-Judge nuanced criteriaAI
Contains word or phraseFast
Exact match & Regex patternsDeterministic

Grade what the agent says

LLM-as-Judge for nuanced criteria in plain English, plus Contains, Exact Match, and Regex when you need fast, deterministic checks.

✓ Action Called · create_reportPass
✓ Navigate · /settings/billingPass
✓ Knowledge Search · “sync errors”Pass

Grade what the agent does

Tool-use evals verify behavior, not just words: the right action called with the right parameters, the right route navigated, the right knowledge retrieved.

User context
Plan Enterprise
Role Admin
Locale de-DESimulated

Test as any user

Set the simulated user's plan, role, and locale — so plan gating, role-based behavior, and localized responses get tested with real-world context.

Test suites
Onboarding100%
Billing questions83%
Report creation96%

Organize into suites

Group scenarios by journey or feature area. Suites are the unit you run, schedule, and track — each with its own pass rate and health.

Schedules
Daily sanity check 08:00 UTCEnabled
Nightly regression 02:00 UTCEnabled
Weekly full sweep Sun 06:00Paused

Run on a schedule

Hourly, daily, weekly, or custom cron — regression detection runs whether anyone remembers to or not, in the timezone you pick.

Run history
Billing questions · 6/6Passed
Billing questions · 5/6Failed
Billing questions · 6/6Passed

Drill into every run

Every execution is logged with outcome, pass rate, and trigger. Open a failed run to see the exact turn, the eval that failed, and what the agent actually said. Compare runs side by side to catch flaky behavior.

Bug → regression test
Conversation #2841 went wrongReported
Scenario: sync-failure recoveryTested nightly

Never ship the same bug twice

When a real conversation goes wrong, import it immediately. It becomes a regression test that ensures the same mistake can never happen again.

Ready to see Trust Lab in action?

Book a live walkthrough and see real conversations turned into a regression suite that runs while you sleep.

Get a demo

What teams lock in with Trust Lab.

Prompt updated · rerun suites
Before change94%
After change96%
28 scenarios rerun 4 minNo regressions
Iterate with confidence

Ship changes without fear

Change a prompt, an action, or the knowledge base — then rerun every suite and see exactly what it did to quality before users feel it.

Model swap
Current model baseline98% pass
Candidate model −40% cost98% pass
VerdictSafe to swap
Cost without compromise

Swap models, keep quality

Prove a faster or cheaper model holds the line on your own scenarios — evidence from your product, not a public benchmark.

Nightly regression · 02:00
Pass rate dropped — 2 scenarios failing
Sleep on it

Catch drift overnight

Knowledge updates, API changes, model updates — a falling pass rate is your signal that something the agent used to get right has broken.

Trust Lab runs the tests. You ship knowing what still works — and what just broke.

Regressions don't announce themselves. Trust Lab does.

Define what the agent must always get right — and know within a day when anything breaks it.