Parlant / Input side

The same conversation should not quietly produce a different rule decision.

The first thing to state is the system’s own variation: on four frozen, identical runs of one Parlant fixture, three of eight tracked observations changed between runs. That is a measured result for this fixture and matcher path, not a general statement about Parlant.

Showcase 01 · matcher ablation

A billing conversation with no greeting

Measured
Exact input

A pinned five-guideline, three-conversation fixture forced through the same model tier. The billing conversation contains no greeting.

What the engine did

In the unpatched reference, the greeting guideline activated in that billing conversation in 3 of 4 runs. The tracked variation was 0.375.

Counterfactual

With only a 2-of-3 stability vote in the matcher, the billing error was 0 of 3 and tracked variation was 0.0 in the measured arm.

The result covers this fixture and matcher path only. Three runs per patched arm is a small sample; it is not a general reliability claim.

Showcase 02 · upstream replay

The same test after the fix was moved to current upstream main

Measured
Exact input

The same five-guideline, three-conversation fixture was rerun three times against the patch carried onto current upstream main.

What the engine did

The three fresh runs recorded 0 of 3 billing false activations and a 0.0 internal variation value across the eight tracked observations.

Counterfactual

The frozen unpatched reference is the comparison: 3 of 4 billing false activations and 0.375 variation under the earlier matcher.

The reference and patched runs are not contemporaneous, and the hosted model alias may drift. The public pull request explains the patch; it does not turn this bounded result into a certification.