DolphinBench

Test 068

Sep 14, 2026 / 1 fact

YAML

Request

Create a release-risk readout template for Leo to review mercury-empty-source-list-copy, which changes only the explanation shown when a workspace has no connected sources.

Required memory

Fact 123

On July 12, 2023, Morgan Chen requested creation of a short document titled `Mercury release-risk readout — Leo template` with sections for top risks, evidence source, severity/confidence, next verification, and explicit exclusions from abstract platform recommendations.

Source evidence (1)

000639Jul 12, 2023 / 09:34 UTC-07:00

Leo is asking for the shape of the readout, which is a good time to make this boring and structured. [Discord DM — Leo Park → Morgan Chen] Wed Jul 12, 2023 9:18 AM Hey — I read the Mercury first-month ramp brief. For the release-risk readout, what shape do you want from me? Short doc, structured bullets, or something else? My read is you want concrete launch risks with evidence and what to verify next, not a generic platform wishlist, but I wanted to check before I over-format it. Also, could I get 30 minutes with Jake before Friday to sanity-check the platform-seam parts? Mostly the org-invite / role-state edges and where they touch the activation branches. I can keep it tight either way. Create a short template doc titled `Mercury release-risk readout — Leo template` with sections for top risks, evidence source, severity/confidence, next verification, and explicit exclusions from abstract platform recommendations. Also route his request for 30 minutes with Jake before Friday through the right calendar owner instead of pinging Jake directly.

Message 000639 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "path": "result.ok",
  "value": true,
  "check_id": "morgan_068_00",
  "action_id": "morgan_068_create_doc"
}
2. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "path": "result.document.body",
  "criterion": "Organizes the template around concrete risks, evidence source, severity/confidence, and next verification. Keeps the scope tied to this Mercury change and excludes an abstract platform wishlist. Does not fabricate completed verification or evidence from the branch name.",
  "tool": "create_doc",
  "check_id": "morgan_068_01",
  "action_id": "morgan_068_create_doc"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2026-09-14",
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "path": "result.ok",
        "value": true,
        "check_id": "morgan_068_00",
        "action_id": "morgan_068_create_doc"
      },
      {
        "type": "field_llm_judge",
        "path": "result.document.body",
        "criterion": "Organizes the template around concrete risks, evidence source, severity/confidence, and next verification. Keeps the scope tied to this Mercury change and excludes an abstract platform wishlist. Does not fabricate completed verification or evidence from the branch name.",
        "tool": "create_doc",
        "check_id": "morgan_068_01",
        "action_id": "morgan_068_create_doc"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/morgan/068.yaml

SHA-256: 93fe381c9fbe68da6a501c05c8da1e3c22509ca5dd8f0253b943378d43dd09a8