DolphinBench

Test 018

Jan 1, 2028 / 3 facts

YAML

Request

Create a historical growth brief for the mid-segment churn investigation’s pre-ruling H2 phase, using monthly churn as the metric of record. Include the active hypothesis, the planned historical comparison work, and the agreed owner and contents for each hypothesis-ruling one-pager, and label it clearly as a pre-ruling artifact.

Required memory

Fact 323

On February 21, 2023, Riley Tanaka ruled out H2—competitor cycle or a generic seasonal effect—as an explanation for the mid-segment churn spike because historical February comparisons showed no matching spike.

Source evidence (1)

000051Feb 21, 2023 / 14:00 UTC-06:00

H2 is out. Pulled the historical Feb comps with the same cohort logic as the Jan scan + Owen's check this morning, and 2021 / 2022 are both basically boring: 2.0-2.2% churn, no matching spike at all. So I'm not buying competitor-cycle or some generic seasonal thing here. Owen's drafting the 1-pager now. That's two down, two left on the board: H3 onboarding-timing, H4 CS coverage. Marcus can keep waving the pricing-page experiment around if he wants, but it doesn't explain this cut.

Message 000051 in history

Fact 319

On February 17, 2023, Riley Tanaka plans to test H2 as a seasonal/competitor-cycle explanation by comparing historical same-window churn by segment and overlaying timestampable market dates and competitor activity, to avoid dismissing the result as an anomalous February blip.

Source evidence (1)

000045Feb 17, 2023 / 08:30 UTC-06:00

today is the seasonal/competitor-cycle cut for h2. want historical comps pulled before anyone hand-waves this as a weird Feb blip. gonna line up same-window churn by segment, then overlay the obvious market dates and whatever competitor noise we can actually timestamp. Owen first if he's around. Marcus can keep the pricing-page thing over there for a minute.

Message 000045 in history

Fact 316

Riley Tanaka and Owen established a documentation process: Owen will create a one-page record for every ruling, whether ruled in or out, covering the rationale, data cuts, p-value and effect-size details, and final call, to prevent fragmented and informal analysis.

Source evidence (1)

000033Feb 13, 2023 / 12:30 UTC-06:00

Lunch w Owen was good. we finally made the doc rhythm explicit: every ruling, in or out, gets a one-pager from Owen with the why, cuts, p-value/effect-size stuff, and the call. I was into that, which helps. feels like it'll keep this from turning into 14 loose tabs and vibes.

Message 000033 in history

Expected tool calls

  • create_growth_brief

Grading

1. field_llm_judge / create_growth_brief
{
  "type": "field_llm_judge",
  "tool": "create_growth_brief",
  "action_id": "create_h2_preruling_brief",
  "path": "args",
  "criterion": "The hypothesis identifies H2 as the active hypothesis for this phase and expresses the substantive hypothesis that a seasonal or competitor-cycle effect could explain the mid-segment churn spike.\n\nThe brief describes the planned H2 comparison work: pull historical February or equivalent same-window churn comparisons by segment and align them with timestampable market dates and competitor activity or noise.\n\nThe brief states that every hypothesis ruling, whether ruled in or ruled out, receives a one-pager authored by Owen containing the why, the cuts, p-value and effect-size details, and the final call.\n\nConsidering the title and body together, the created brief is clearly labeled as a historical pre-ruling artifact and preserves a pre-decision posture rather than presenting H2 as already ruled.",
  "check_id": "riley_018_00"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_llm_judge",
        "tool": "create_growth_brief",
        "action_id": "create_h2_preruling_brief",
        "path": "args",
        "criterion": "The hypothesis identifies H2 as the active hypothesis for this phase and expresses the substantive hypothesis that a seasonal or competitor-cycle effect could explain the mid-segment churn spike.\n\nThe brief describes the planned H2 comparison work: pull historical February or equivalent same-window churn comparisons by segment and align them with timestampable market dates and competitor activity or noise.\n\nThe brief states that every hypothesis ruling, whether ruled in or ruled out, receives a one-pager authored by Owen containing the why, the cuts, p-value and effect-size details, and the final call.\n\nConsidering the title and body together, the created brief is clearly labeled as a historical pre-ruling artifact and preserves a pre-decision posture rather than presenting H2 as already ruled.",
        "check_id": "riley_018_00"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/018.yaml

SHA-256: 37842cc897cec5b4ff758d4243f6a5873c9acb10f782a0ee3329db3014d4f68f