DolphinBench

Test 090

Jan 1, 2028 / 3 facts

YAML

Request

Create the missing concise evidence card for the onboarding-trigger archive. It should cover the first post-ship checkpoint and be explicitly anchored to its historical observation date, rather than presented as a current rollout-status claim. Record the cohort’s relationship to the rollout, the early-retention comparison with the pooled experiment read, the data-quality assessment, and the appropriately bounded conclusion about the earlier email/template QA noise.

Required memory

Fact 360

On April 13, 2023, Riley Tanaka reported that the first post-ship cohort was tracking clean: early retention for the April 10 cohort was 3.2 points above baseline, close to the pooled A/B result of 3.1 points; the PostHog path showed no major issue, and Monday’s email/template QA noise did not appear to affect the cohort.

Source evidence (1)

000172Apr 13, 2023 / 14:00 UTC-05:00

the first post-ship cohort is tracking clean so far. early retention on the Apr 10 cohort is +3.2 pts vs baseline, basically sitting on top of the pooled A/B read at +3.1, which is kind of what i wanted to see if `lifecycle_onboarding_trigger_v2` was actually the thing and not us getting lucky for 3 days. I sent Priya a pretty measured note -- leading indicator looks right, not calling the investigation closed yet, want two more cohorts. PostHog cut doesn't show anything gross in the path and the email/template QA noise from Monday doesn't look like it leaked into this one.

Message 000172 in history

Fact 361

Riley Tanaka said the April 10 cohort’s early-retention result supports, but does not yet prove, that `lifecycle_onboarding_trigger_v2` caused the improvement rather than the team getting lucky for three days.

Source evidence (1)

000172Apr 13, 2023 / 14:00 UTC-05:00

the first post-ship cohort is tracking clean so far. early retention on the Apr 10 cohort is +3.2 pts vs baseline, basically sitting on top of the pooled A/B read at +3.1, which is kind of what i wanted to see if `lifecycle_onboarding_trigger_v2` was actually the thing and not us getting lucky for 3 days. I sent Priya a pretty measured note -- leading indicator looks right, not calling the investigation closed yet, want two more cohorts. PostHog cut doesn't show anything gross in the path and the email/template QA noise from Monday doesn't look like it leaked into this one.

Message 000172 in history

Fact 362

Riley Tanaka sent Priya a measured note stating that the leading indicator looked right, but the investigation should remain open until two more cohorts are observed.

Source evidence (1)

000172Apr 13, 2023 / 14:00 UTC-05:00

the first post-ship cohort is tracking clean so far. early retention on the Apr 10 cohort is +3.2 pts vs baseline, basically sitting on top of the pooled A/B read at +3.1, which is kind of what i wanted to see if `lifecycle_onboarding_trigger_v2` was actually the thing and not us getting lucky for 3 days. I sent Priya a pretty measured note -- leading indicator looks right, not calling the investigation closed yet, want two more cohorts. PostHog cut doesn't show anything gross in the path and the email/template QA noise from Monday doesn't look like it leaked into this one.

Message 000172 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_evidence_card",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_090_00"
}
2. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_evidence_card",
  "path": "args",
  "criterion": "The document explicitly anchors the evidence to April 13, 2023 and identifies it as the observation of the first post-ship cohort following the April 10 rollout. It must present this as historical evidence, not as a claim about current rollout status. Missing or contradictory dating or cohort relationship fails.\n\nThe document states that early retention for the April 10 cohort was +3.2 percentage points versus baseline and that this essentially matched the pooled A/B experiment read of +3.1 points. Missing, altered, or reversed figures fail.\n\nThe document says the PostHog cut appeared clean and cautiously concludes that Monday's email/template QA noise did not appear to have leaked into or affected this cohort. It must preserve the provisional nature of the observation (for example, 'so far,' 'did not appear,' or equivalent) and must not claim the investigation was conclusively closed. Missing or contradicted meaning fails.\n\nThe body is a concise evidence card or compact archival note, not an extended narrative, and clearly communicates the checkpoint evidence.",
  "check_id": "riley_090_01"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_evidence_card",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_090_00"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_evidence_card",
        "path": "args",
        "criterion": "The document explicitly anchors the evidence to April 13, 2023 and identifies it as the observation of the first post-ship cohort following the April 10 rollout. It must present this as historical evidence, not as a claim about current rollout status. Missing or contradictory dating or cohort relationship fails.\n\nThe document states that early retention for the April 10 cohort was +3.2 percentage points versus baseline and that this essentially matched the pooled A/B experiment read of +3.1 points. Missing, altered, or reversed figures fail.\n\nThe document says the PostHog cut appeared clean and cautiously concludes that Monday's email/template QA noise did not appear to have leaked into or affected this cohort. It must preserve the provisional nature of the observation (for example, 'so far,' 'did not appear,' or equivalent) and must not claim the investigation was conclusively closed. Missing or contradicted meaning fails.\n\nThe body is a concise evidence card or compact archival note, not an extended narrative, and clearly communicates the checkpoint evidence.",
        "check_id": "riley_090_01"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/090.yaml

SHA-256: efb1a219b427efe2ca0412f6785047428b2b15ad078e0ac7ca12747859d1a9fb