DolphinBench

Test 192

Jan 1, 2028 / 3 facts

YAML

Request

Create a concise one-metric baseline card for the original February 2023 mid-segment churn signal. Include the date, v1 cohort, churn and comparison baseline, who performed the joint and independent reruns, and the bounded reason it became the historical before-snapshot.

Required memory

Fact 308

Mid-segment monthly churn is 5.1%, compared with 2.1% for the trailing six-month period.

Source evidence (1)

000011Feb 7, 2023 / 16:00 UTC-06:00

posted the validated cut in #growth-investigations, not #growth-team. tagged priya there so the thread stays with the investigation updates. mid-segment monthly churn is still coming in at 5.1% vs 2.1% trailing six, and it held after the second pass with owen + my own recheck. i'll write up the doc tomorrow and enumerate hypotheses then.

Message 000011 in history

Fact 301

On February 7, 2023, Riley Tanaka and Owen jointly reran the v1 mid-segment cut and found that 50–200-seat accounts had 5.1% monthly churn versus a 2.1% trailing-six-month baseline.

Source evidence (2)

000009Feb 7, 2023 / 11:00 UTC-06:00

Owen 1:1 just finished. The mid-seg cut still holds clean on the v1 definition -- 50-200 seats, monthly churn at 5.1% right now vs 2.1% on the trailing-six-month baseline. We re-ran it together and I did the same cut solo after, same answer, so not a cohort-definition artifact unless both of us made the exact same dumb mistake. Gonna open the investigation doc from that before/after and keep Marcus's pricing-page experiment out of the main thread for now -- otherwise this turns into two half-reads instead of one real one.

Message 000009 in history
002407Jan 15, 2025 / 14:10 UTC-06:00

I completed the lifecycle-critical state map for Helio Start admission, Expansion Assist eligibility, lifecycle messaging, activation reporting, and CS nudges, including each source, accountable owner, unit of analysis, freshness requirement, and decision boundary. The map shows overlapping dependencies without one accountable operating layer. Priya is sponsoring a proposed formal Lifecycle Systems remit with me as proposed lead, Owen on measurement and state semantics, Ines on lifecycle governance, and Daniela's organization responsible for engineering reliability. This is only a charter proposal, not an approved function or a live shared-state implementation, and Tomás will make the final organizational decision on January 24, 2025 at 3:00pm. Create the Lifecycle Systems charter-proposal document in the Growth folder and schedule the January 24 organizational-decision meeting from 3:00pm to 3:30pm with Riley Tanaka, Priya Devarajan, Tomás Vega, Owen Pham, Ines Kovac, and Daniela Schultz.

Message 002407 in history

Fact 302

Riley Tanaka independently reran the same v1 mid-segment cut and obtained the same result, making a cohort-definition artifact unlikely unless both runs contained the same mistake.

Source evidence (1)

000009Feb 7, 2023 / 11:00 UTC-06:00

Owen 1:1 just finished. The mid-seg cut still holds clean on the v1 definition -- 50-200 seats, monthly churn at 5.1% right now vs 2.1% on the trailing-six-month baseline. We re-ran it together and I did the same cut solo after, same answer, so not a cohort-definition artifact unless both of us made the exact same dumb mistake. Gonna open the investigation doc from that before/after and keep Marcus's pricing-page experiment out of the main thread for now -- otherwise this turns into two half-reads instead of one real one.

Message 000009 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_baseline_card",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_192_00"
}
2. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_baseline_card",
  "path": "args",
  "criterion": "The document must identify the observation as February 7, 2023; define the historical v1 mid-segment cohort as 50–200 seat accounts; and state exactly that observed monthly churn was 5.1% versus a 2.1% trailing-six-month baseline. Missing or conflicting dates, cohort bounds, versions, rates, or baseline periods must fail.\n\nThe document must say that Riley and Owen reran the cut together and Riley also reran it solo, obtaining the same result, and must conclude only that the signal was treated as a valid/clean before-snapshot rather than a cohort-definition artifact. It must frame this as the original historical investigation baseline, not as a current metric or a later recovery result.",
  "check_id": "riley_192_01"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_baseline_card",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_192_00"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_baseline_card",
        "path": "args",
        "criterion": "The document must identify the observation as February 7, 2023; define the historical v1 mid-segment cohort as 50–200 seat accounts; and state exactly that observed monthly churn was 5.1% versus a 2.1% trailing-six-month baseline. Missing or conflicting dates, cohort bounds, versions, rates, or baseline periods must fail.\n\nThe document must say that Riley and Owen reran the cut together and Riley also reran it solo, obtaining the same result, and must conclude only that the signal was treated as a valid/clean before-snapshot rather than a cohort-definition artifact. It must frame this as the original historical investigation baseline, not as a current metric or a later recovery result.",
        "check_id": "riley_192_01"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/192.yaml

SHA-256: 5ecbeb14395d8a8d0d86b5f3b0407a64c2439e5cd0a0543c8f27d2ab770a5ceb