DolphinBench

Test 034

Jan 1, 2028 / 2 facts

YAML

Request

Create a document containing a one-metric card for our historical rollout evidence covering the third post-ship onboarding cohort; include the cohort context, metric, comparison basis, and measured change, and keep it strictly descriptive without causal or forecasting claims.

Required memory

Fact 373

The Apr. 25 third post-ship mid-segment cohort cut showed week-1 retention 3.4 percentage points above baseline, exceeding the first post-ship cohort’s 3.2-point lift.

Source evidence (1)

000202Apr 25, 2023 / 11:00 UTC-05:00

third post-ship cut came in clean. Apr 25 week-1 retention on the mid-seg cohort is +3.4pts vs baseline, so better than the first +3.2 read and not doing that annoying roll-over thing after the Apr 17 wobble. I dropped it in #growth-investigations and, small thing but not small, called it "the trend" for the first time since Feb instead of "the signal." feels like the 75-200 recut plus backing out the CS auto-nudge mess got us to the real read. still gonna watch another week obviously, but this is the first one that feels actually solid.

Message 000202 in history

Fact 363

On April 17, the second post-ship cohort measured +1.4 percentage points in week-1 retention versus baseline, below the April 10 cohort's +3.2 points and the pooled experiment result of +3.1 points; Riley Tanaka does not consider the rollout broken based on this cut.

Source evidence (1)

000180Apr 17, 2023 / 11:00 UTC-05:00

owen 1:1 was rough. second post-ship cohort came in at +1.4 pts week-1 retention vs baseline, so... not broken, but way under the apr 10 cohort at +3.2 and also under the pooled experiment read at +3.1. I flagged it internally right away and am gonna spend tue/wed slicing it every dumb way possible -- signup source, tier, geo, day-of-week. my guess is composition before i call the trigger itself, but marcus is already gonna sniff around if this sits here another cut. PostHog is clean on `lifecycle_onboarding_trigger_v2` at 100%, so this is more likely cohort shape / who entered than rollout weirdness. leaning 'don't narrate anything yet'.

Message 000180 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_historical_evidence_card",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_034_00"
}
2. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_historical_evidence_card",
  "path": "args",
  "criterion": "The document is a concise one-metric evidence card that clearly identifies the result as the third cohort after the onboarding-trigger rollout for the mid-segment cohort, uses week-1 retention as the sole performance metric, gives baseline as the comparison basis, and reports a measured increase of 3.4 percentage points. It presents this as a historical measured result and does not claim that the rollout caused the change or make a forecast.",
  "check_id": "riley_034_01"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_historical_evidence_card",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_034_00"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_historical_evidence_card",
        "path": "args",
        "criterion": "The document is a concise one-metric evidence card that clearly identifies the result as the third cohort after the onboarding-trigger rollout for the mid-segment cohort, uses week-1 retention as the sole performance metric, gives baseline as the comparison basis, and reports a measured increase of 3.4 percentage points. It presents this as a historical measured result and does not claim that the rollout caused the change or make a forecast.",
        "check_id": "riley_034_01"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/034.yaml

SHA-256: 14aac3357b84338153ade386596e1f6197e9b77799de5defee9fb1cfbd1d082f