Week of July 4–10, 2026
A lighter week. No new features — just fixes to things that were off after last week's push. These are worth calling out because they affect the accuracy of results you're already seeing.
Scoring accuracy
Fixed a parser mismatch between the score and confidence values returned by the LLM evaluator. In some cases, a check's displayed score wasn't matching what the evaluator actually returned. Results now reflect the evaluator's judgment accurately.
Scenario field cleanup
When you delete a field from a Context definition, the data for that field in your existing Scenarios is now removed automatically. Previously, the data could persist invisibly and show up in unexpected places. The Context and its Scenarios now stay in sync.
Evaluation prompt fix
Fixed how scoring criteria were being included in the prompt sent to the evaluator. In some configurations, criteria weren't being passed correctly, which could cause qualitative checks to evaluate against the wrong standard. This is corrected.
