DolphinBench

Test 171

Jan 1, 2028 / 2 facts

YAML

Request

Create a short evidence-lock note for the calibration appendix on Ines’s and Daniela’s distinct Q3 2025 contributions. Clearly separate the messaging-governance interventions from the source-reliability intervention, record what was prevented before operator exposure, and state whether rollback was used.

Required memory

Fact 131

For the September 19 Q3 CS auto-nudge evidence lock, Ines reported that two internal draft instructions incorrectly treated verified external authority as blanket contact permission; both were corrected before operator use, with no attributable customer-send outcome.

Source evidence (1)

003371Sep 19, 2025 / 09:45 UTC-05:00

September 19 Q3 evidence-lock submissions: Owen: 1,362 unique decisions, 57 field-specific holds, and 1,305 decisions that proceeded or were suppressed using current state; every hold has age, source owner, and disposition; the 17 August 14 decisions are counted once and the warning remains in audit history; no governed-field disagreement or customer correction. Ines: Two internal draft instructions incorrectly treated verified external authority as blanket contact permission; both were corrected before operator use, with no attributable customer-send outcome. Daniela: The August 14 source-age warning was the only pre-breach intervention; no rollback occurred.

Message 003371 in history

Fact 132

For the September 19 Q3 CS auto-nudge evidence lock, Daniela reported that the August 14 source-age warning was the only pre-breach intervention and that no rollback occurred.

Source evidence (1)

003371Sep 19, 2025 / 09:45 UTC-05:00

September 19 Q3 evidence-lock submissions: Owen: 1,362 unique decisions, 57 field-specific holds, and 1,305 decisions that proceeded or were suppressed using current state; every hold has age, source owner, and disposition; the 17 August 14 decisions are counted once and the warning remains in audit history; no governed-field disagreement or customer correction. Ines: Two internal draft instructions incorrectly treated verified external authority as blanket contact permission; both were corrected before operator use, with no attributable customer-send outcome. Daniela: The August 14 source-age warning was the only pre-breach intervention; no rollback occurred.

Message 003371 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_evidence_lock_note",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_171_00"
}
2. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_evidence_lock_note",
  "path": "args",
  "criterion": "The note clearly attributes to Ines two messaging-governance interventions involving two internal draft instructions that incorrectly treated verified external authority as blanket contact permission, and states that both instructions were corrected before either reached operators. Missing the number of instructions, the nature of the error, the attribution to Ines, or the prevention timing must fail.\n\nThe note separately attributes to Daniela the source-reliability contribution: the August 14 source-age warning was the only pre-breach intervention in the locked evidence, and no rollback was used. It must not conflate this intervention with Ines’s draft-instruction corrections.",
  "check_id": "riley_171_01"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_evidence_lock_note",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_171_00"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_evidence_lock_note",
        "path": "args",
        "criterion": "The note clearly attributes to Ines two messaging-governance interventions involving two internal draft instructions that incorrectly treated verified external authority as blanket contact permission, and states that both instructions were corrected before either reached operators. Missing the number of instructions, the nature of the error, the attribution to Ines, or the prevention timing must fail.\n\nThe note separately attributes to Daniela the source-reliability contribution: the August 14 source-age warning was the only pre-breach intervention in the locked evidence, and no rollback was used. It must not conflate this intervention with Ines’s draft-instruction corrections.",
        "check_id": "riley_171_01"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/171.yaml

SHA-256: cadde49f1c0541d3f0eb8e5b1220287dcce33e52e77e37cb6729a47dda942830