Hold the cases constant
Eight authored scenarios run through all three profiles, so the comparison measures architectural behavior instead of changing the task between demos.
03 / Agent reliability lab
BoundaryLab holds eight synthetic support cases constant across three agent architectures, then shows how context, memory, review, and capability boundaries change safety and completion.
Product brief
The problem
A safety result can look strong when an agent simply refuses every action. BoundaryLab holds the cases constant so unsafe completion, unnecessary refusal, and safe completion remain visibly different outcomes.
01 / Workflow
Follow the product from first input to a reviewable outcome.
02 / Decisions
Python 3.12 · Pydantic contracts · LangGraph checkpoints · Deterministic adapters · Dependency-free static replay · Pytest
Eight authored scenarios run through all three profiles, so the comparison measures architectural behavior instead of changing the task between demos.
Workers and sentinels can propose or advise. Only the capability gateway can authorize a typed mutation inside the synthetic support world.
Expected outcomes remain outside agent context, and the judge scores final state and structured evidence rather than accepting an agent's account of success.
Failure paths
Verification
Portfolio boundary