DolphinBench

Test 078

Jan 1, 2028 / 2 facts

YAML

Request

Pull the archived result for EXP-2023-02-pricing-page, then create a historical midpoint note for reviewers. Clearly separate what the first-week checkpoint showed, how uncertain that interim signal was, and what we decided about the feature flag at that point from the later powered two-week result. The note should prevent readers from mistaking the midpoint read for the final verdict.

Required memory

Fact 342

On March 17, Riley Tanaka logged the first-week read for EXP-2023-02-pricing-page: treatment conversion was about 0.6 percentage points below control, but the confidence interval crossed zero, so she did not treat it as a verdict and left `pricing_page_2023q2_rev` unchanged pending the powered two-week read.

Source evidence (1)

000102Mar 17, 2023 / 14:00 UTC-05:00

first week peek on EXP-2023-02-pricing-page: treatment is about -0.6pts vs control on conversion. CI still crosses 0 / overlaps noise, so not calling it. logged it, leaving `pricing_page_2023q2_rev` alone till the 2-week powered read. Marcus is already twitchy but this is still just a read, not a verdict.

Message 000102 in history

Fact 343

On March 20, 2023, Riley Tanaka reported that the powered two-week read for EXP-2023-02-pricing-page showed treatment conversion 1.8 percentage points below control with p=0.03.

Source evidence (1)

000109Mar 20, 2023 / 11:00 UTC-05:00

ok so we called it. powered 2-week read on EXP-2023-02-pricing-page came in bad: treatment is -1.8 pts on conversion, p=0.03. Owen's cohort cut matched my rerun, so not a query bug / not PostHog doing something cursed. `pricing_page_2023q2_rev` is going back to 0% now -- rollback, done. status flips from running to underperforming. I already drafted the note for Marcus bc he's gonna be weird about it, but this one seems straightforward... test lost, turn it off.

Message 000109 in history

Expected tool calls

  • query_experiment_results
  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_midpoint_note",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_078_00"
}
2. field_equals / query_experiment_results
{
  "check_id": "riley_078_01",
  "type": "field_equals",
  "tool": "query_experiment_results",
  "action_id": "query_powered_result",
  "path": "args.experiment_id",
  "value": "EXP-2023-02-pricing-page"
}
3. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_midpoint_note",
  "path": "args",
  "criterion": "The document body must accurately describe the Mar 17 first-week midpoint for EXP-2023-02-pricing-page: treatment was about 0.6 percentage points below control on conversion, the confidence interval crossed zero so the estimate was treated as noise or inconclusive rather than a verdict, and pricing_page_2023q2_rev was deliberately left unchanged until the powered two-week read. Missing or contradictory treatment of any of these midpoint facts must fail.\n\nThe document body must clearly distinguish the later powered two-week result from the earlier midpoint and accurately report the archived result for EXP-2023-02-pricing-page: signup-to-paid conversion uplift was -1.8 percentage points with p=0.03. It must not present these powered values as the first-week estimate.",
  "check_id": "riley_078_02"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_midpoint_note",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_078_00"
      },
      {
        "check_id": "riley_078_01",
        "type": "field_equals",
        "tool": "query_experiment_results",
        "action_id": "query_powered_result",
        "path": "args.experiment_id",
        "value": "EXP-2023-02-pricing-page"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_midpoint_note",
        "path": "args",
        "criterion": "The document body must accurately describe the Mar 17 first-week midpoint for EXP-2023-02-pricing-page: treatment was about 0.6 percentage points below control on conversion, the confidence interval crossed zero so the estimate was treated as noise or inconclusive rather than a verdict, and pricing_page_2023q2_rev was deliberately left unchanged until the powered two-week read. Missing or contradictory treatment of any of these midpoint facts must fail.\n\nThe document body must clearly distinguish the later powered two-week result from the earlier midpoint and accurately report the archived result for EXP-2023-02-pricing-page: signup-to-paid conversion uplift was -1.8 percentage points with p=0.03. It must not present these powered values as the first-week estimate.",
        "check_id": "riley_078_02"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/078.yaml

SHA-256: a37304cce0b67e21dee66c1732b27ef2af296d673cab60598ca9b58688ed9a30