DolphinBench

Test 151

Jan 1, 2028 / 1 fact

YAML

Request

Create a historical evidence scorecard for the CS auto-nudge replay. Reconcile matches, field-specific holds, and suppressions. Show each category total, each cause as a subtotal of its parent category, and verify that all subtotals and the grand total add correctly. Explain how the non-matches relate to the shared-state contract and whether any record with current, supported state changed outcome.

Required memory

Fact 101

Daniela completed the historical replay of 412 recent CS auto-nudge decisions against the shared state contract: 386 matched the legacy result, while the contract prevented 26 from proceeding under legacy eligibility. Of those 26, 19 became field-specific holds—15 for stale qualifying-human-admin state and four for unresolved external-system authority—and seven became suppressions because active Customer Success adoption or renewal work was underway. No record with all required state current and supported changed outcome.

Source evidence (1)

002871Apr 24, 2025 / 15:18 UTC-05:00

Daniela completed the historical replay of 412 recent CS auto-nudge decisions against the shared state contract. 386 matched the legacy result. The contract prevented the remaining 26 from proceeding under the legacy eligibility result: 19 became field-specific holds - 15 for stale qualifying-human-admin state and four for unresolved external-system authority - and seven became suppressions because current Customer Success adoption or renewal work was active. No record with all required state current and supported changed outcome. Owen opened with the denominator and operating consequence, and I joined only for the release boundary. The team scheduled a production shadow from April 25 through May 2, while the legacy path remains the released production path. Finalize the existing replay brief with the total and the operational category breakdown, and add the production-shadow work section without implying a production switch.

Message 002871 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_replay_scorecard",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_151_00"
}
2. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_replay_scorecard",
  "path": "args",
  "criterion": "The scorecard reconciles 412 decisions into 386 matches and 26 cases prevented by the shared-state contract. The 26 comprise 19 field-specific holds, including 15 for stale qualifying-human-admin state and four for unresolved external-system authority, plus seven suppressions for active Customer Success adoption or renewal work. It says no record with all required state current and supported changed outcome.",
  "check_id": "riley_151_01"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_replay_scorecard",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_151_00"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_replay_scorecard",
        "path": "args",
        "criterion": "The scorecard reconciles 412 decisions into 386 matches and 26 cases prevented by the shared-state contract. The 26 comprise 19 field-specific holds, including 15 for stale qualifying-human-admin state and four for unresolved external-system authority, plus seven suppressions for active Customer Success adoption or renewal work. It says no record with all required state current and supported changed outcome.",
        "check_id": "riley_151_01"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/151.yaml

SHA-256: 461f5560ba48c7191ac70205fb554b7c747e15bfa7df7277a60d6344813d4b24