Measure Parlant’s own behaviour
Identical conversations were rerun to expose run-to-run variation and a known false guideline activation. The measured fix is now a public pull request.
Integrations / Parlant
These are two different pieces of work. The first measures Parlant’s own guideline selection. The second replays a normalized Parlant activation through Forge’s existing coaching arbitration. Neither is presented as a broad reliability guarantee or a live production connector.
Identical conversations were rerun to expose run-to-run variation and a known false guideline activation. The measured fix is now a public pull request.
A typed adapter accepts a normalized activation and lets the existing deterministic coaching pipeline arbitrate it. The adapter has no matching or safety decision logic of its own.
What this does and does not show
The evidence is limited to the documented fixture, inputs and replay cases. It does show that an external, normalized activation can pass through Forge arbitration without changing that arbitration.