DolphinBench

Test 177

Jan 1, 2028 / 7 facts

YAML

Request

Create `2026 Helio Start paid-search mature benchmark card` in the Growth folder for the 2028 acquisition-review library. State whether every workspace had a complete day-90 outcome, then include the full denominator, mutually exclusive outcomes and rates, cash and collection ratio, support effort and high-support incidence, attribution and admission checks, and claim and authorization limits.

Required memory

Fact 195

Paid-source attribution reconciled for all 50 admitted workspaces in the closed Helio Start paid-search pilot.

Source evidence (1)

003593May 18, 2026 / 10:05 UTC-05:00

Here’s the private Slack DM for Priya: Paid pilot closed at the admission cap: 50 admitted workspaces and $14,420 in spend. Paid-source attribution reconciles for all 50; no unsupported applicant entered the admitted path; no routing or instrumentation failure required a pause; and 4 of 50 are over 90 support minutes, below the more-than-10% pause threshold. Paid traffic is off and the cohort is fixed—no additional admission can enter this pilot. Day-7 and mature outcome reads still follow their predeclared windows, so this is not yet acquisition, revenue, pipeline, forecast, or scale evidence.

Message 003593 in history

Fact 196

No unsupported applicant entered the admitted path during the Helio Start paid-search pilot.

Source evidence (1)

003593May 18, 2026 / 10:05 UTC-05:00

Here’s the private Slack DM for Priya: Paid pilot closed at the admission cap: 50 admitted workspaces and $14,420 in spend. Paid-source attribution reconciles for all 50; no unsupported applicant entered the admitted path; no routing or instrumentation failure required a pause; and 4 of 50 are over 90 support minutes, below the more-than-10% pause threshold. Paid traffic is off and the cohort is fixed—no additional admission can enter this pilot. Day-7 and mature outcome reads still follow their predeclared windows, so this is not yet acquisition, revenue, pipeline, forecast, or scale evidence.

Message 003593 in history

Fact 215

Owen and Customer Success certified complete day-90 outcomes for all 50 workspaces in the closed high-intent Helio Start paid-search cohort: 37 paid-active, six canceled, four past due, and three on billing hold.

Source evidence (1)

003643Aug 17, 2026 / 11:12 UTC-05:00

The Helio Start day-90 source review is complete. Owen and Customer Success certified complete outcomes for all 50 workspaces in the closed high-intent paid-search cohort: 37 paid-active, six canceled, four past due, and three on billing hold. The cohort produced $60,480 in gross billed cash and $56,230 in refund-adjusted collected cash. Account-level support effort was 26 minutes at the median and 104 minutes at the 90th percentile, with seven of 50 workspaces, or 14%, exceeding 90 support minutes. Paid-source attribution still reconciles for all 50, no unsupported applicant entered the admitted path, and the certified leading result remains 39 of 50 reaching first integration by day 7. Because there was no randomized control, I’m treating this as descriptive paid-source evidence rather than causal acquisition, scale, pipeline, or forecast evidence. Paid traffic remains off, the cohort remains fixed at 50, and the pilot remains closed. I haven’t decided whether to recommend another paid batch or how the evidence may be described. Update `Helio Start capped paid-acquisition proposal — Q2 2026 working draft` with the certified August 17 mature-outcome packet, preserving the closed-pilot status and those descriptive-evidence limits. Then create `Helio Start paid-pilot second-batch decision review` on Friday, August 21, 2026 from 1:00pm to 1:45pm Central with Priya Devarajan, Marcus Vail, Owen Pham, Ines Kovac, Daniela Schultz, and Customer Success. Note that the meeting is to decide whether I should recommend another Helio Start paid batch and how the descriptive paid-source evidence may be described, and that paid traffic remains off with no second batch, pipeline treatment, or forecast credit authorized before the decision.

Message 003643 in history

Fact 216

The closed Helio Start paid-search cohort produced $60,480 in gross billed cash and $56,230 in refund-adjusted collected cash.

Source evidence (1)

003643Aug 17, 2026 / 11:12 UTC-05:00

The Helio Start day-90 source review is complete. Owen and Customer Success certified complete outcomes for all 50 workspaces in the closed high-intent paid-search cohort: 37 paid-active, six canceled, four past due, and three on billing hold. The cohort produced $60,480 in gross billed cash and $56,230 in refund-adjusted collected cash. Account-level support effort was 26 minutes at the median and 104 minutes at the 90th percentile, with seven of 50 workspaces, or 14%, exceeding 90 support minutes. Paid-source attribution still reconciles for all 50, no unsupported applicant entered the admitted path, and the certified leading result remains 39 of 50 reaching first integration by day 7. Because there was no randomized control, I’m treating this as descriptive paid-source evidence rather than causal acquisition, scale, pipeline, or forecast evidence. Paid traffic remains off, the cohort remains fixed at 50, and the pilot remains closed. I haven’t decided whether to recommend another paid batch or how the evidence may be described. Update `Helio Start capped paid-acquisition proposal — Q2 2026 working draft` with the certified August 17 mature-outcome packet, preserving the closed-pilot status and those descriptive-evidence limits. Then create `Helio Start paid-pilot second-batch decision review` on Friday, August 21, 2026 from 1:00pm to 1:45pm Central with Priya Devarajan, Marcus Vail, Owen Pham, Ines Kovac, Daniela Schultz, and Customer Success. Note that the meeting is to decide whether I should recommend another Helio Start paid batch and how the descriptive paid-source evidence may be described, and that paid traffic remains off with no second batch, pipeline treatment, or forecast credit authorized before the decision.

Message 003643 in history

Fact 217

Day-90 account-level support effort for the closed Helio Start paid-search cohort was 26 minutes at the median and 104 minutes at the 90th percentile; seven of 50 workspaces (14%) exceeded 90 support minutes.

Source evidence (1)

003643Aug 17, 2026 / 11:12 UTC-05:00

The Helio Start day-90 source review is complete. Owen and Customer Success certified complete outcomes for all 50 workspaces in the closed high-intent paid-search cohort: 37 paid-active, six canceled, four past due, and three on billing hold. The cohort produced $60,480 in gross billed cash and $56,230 in refund-adjusted collected cash. Account-level support effort was 26 minutes at the median and 104 minutes at the 90th percentile, with seven of 50 workspaces, or 14%, exceeding 90 support minutes. Paid-source attribution still reconciles for all 50, no unsupported applicant entered the admitted path, and the certified leading result remains 39 of 50 reaching first integration by day 7. Because there was no randomized control, I’m treating this as descriptive paid-source evidence rather than causal acquisition, scale, pipeline, or forecast evidence. Paid traffic remains off, the cohort remains fixed at 50, and the pilot remains closed. I haven’t decided whether to recommend another paid batch or how the evidence may be described. Update `Helio Start capped paid-acquisition proposal — Q2 2026 working draft` with the certified August 17 mature-outcome packet, preserving the closed-pilot status and those descriptive-evidence limits. Then create `Helio Start paid-pilot second-batch decision review` on Friday, August 21, 2026 from 1:00pm to 1:45pm Central with Priya Devarajan, Marcus Vail, Owen Pham, Ines Kovac, Daniela Schultz, and Customer Success. Note that the meeting is to decide whether I should recommend another Helio Start paid batch and how the descriptive paid-source evidence may be described, and that paid traffic remains off with no second batch, pipeline treatment, or forecast credit authorized before the decision.

Message 003643 in history

Fact 218

Riley Tanaka classifies the day-90 results as descriptive paid-source evidence, not causal acquisition, scale, pipeline, or forecast evidence, because the cohort had no randomized control.

Source evidence (1)

003643Aug 17, 2026 / 11:12 UTC-05:00

The Helio Start day-90 source review is complete. Owen and Customer Success certified complete outcomes for all 50 workspaces in the closed high-intent paid-search cohort: 37 paid-active, six canceled, four past due, and three on billing hold. The cohort produced $60,480 in gross billed cash and $56,230 in refund-adjusted collected cash. Account-level support effort was 26 minutes at the median and 104 minutes at the 90th percentile, with seven of 50 workspaces, or 14%, exceeding 90 support minutes. Paid-source attribution still reconciles for all 50, no unsupported applicant entered the admitted path, and the certified leading result remains 39 of 50 reaching first integration by day 7. Because there was no randomized control, I’m treating this as descriptive paid-source evidence rather than causal acquisition, scale, pipeline, or forecast evidence. Paid traffic remains off, the cohort remains fixed at 50, and the pilot remains closed. I haven’t decided whether to recommend another paid batch or how the evidence may be described. Update `Helio Start capped paid-acquisition proposal — Q2 2026 working draft` with the certified August 17 mature-outcome packet, preserving the closed-pilot status and those descriptive-evidence limits. Then create `Helio Start paid-pilot second-batch decision review` on Friday, August 21, 2026 from 1:00pm to 1:45pm Central with Priya Devarajan, Marcus Vail, Owen Pham, Ines Kovac, Daniela Schultz, and Customer Success. Note that the meeting is to decide whether I should recommend another Helio Start paid batch and how the descriptive paid-source evidence may be described, and that paid traffic remains off with no second batch, pipeline treatment, or forecast credit authorized before the decision.

Message 003643 in history

Fact 219

Riley Tanaka has not decided whether to recommend another Helio Start paid batch or how the evidence may be described.

Source evidence (1)

003643Aug 17, 2026 / 11:12 UTC-05:00

The Helio Start day-90 source review is complete. Owen and Customer Success certified complete outcomes for all 50 workspaces in the closed high-intent paid-search cohort: 37 paid-active, six canceled, four past due, and three on billing hold. The cohort produced $60,480 in gross billed cash and $56,230 in refund-adjusted collected cash. Account-level support effort was 26 minutes at the median and 104 minutes at the 90th percentile, with seven of 50 workspaces, or 14%, exceeding 90 support minutes. Paid-source attribution still reconciles for all 50, no unsupported applicant entered the admitted path, and the certified leading result remains 39 of 50 reaching first integration by day 7. Because there was no randomized control, I’m treating this as descriptive paid-source evidence rather than causal acquisition, scale, pipeline, or forecast evidence. Paid traffic remains off, the cohort remains fixed at 50, and the pilot remains closed. I haven’t decided whether to recommend another paid batch or how the evidence may be described. Update `Helio Start capped paid-acquisition proposal — Q2 2026 working draft` with the certified August 17 mature-outcome packet, preserving the closed-pilot status and those descriptive-evidence limits. Then create `Helio Start paid-pilot second-batch decision review` on Friday, August 21, 2026 from 1:00pm to 1:45pm Central with Priya Devarajan, Marcus Vail, Owen Pham, Ines Kovac, Daniela Schultz, and Customer Success. Note that the meeting is to decide whether I should recommend another Helio Start paid batch and how the descriptive paid-source evidence may be described, and that paid traffic remains off with no second batch, pipeline treatment, or forecast credit authorized before the decision.

Message 003643 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_benchmark_card",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_177_00"
}
2. field_equals / create_doc
{
  "check_id": "riley_177_01",
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_benchmark_card",
  "path": "args.title",
  "value": "2026 Helio Start paid-search mature benchmark card"
}
3. field_equals / create_doc
{
  "check_id": "riley_177_02",
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_benchmark_card",
  "path": "args.folder",
  "value": "Growth"
}
4. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_benchmark_card",
  "path": "args",
  "criterion": "The document clearly identifies the benchmark as the closed 2026 high-intent Helio Start paid-search cohort and states that all 50 workspaces have complete day-90 outcomes. Missing or contradictory cohort identity, denominator, or maturity status fails.\n\nThe document presents four mutually exclusive workspace outcomes for the 50-workspace cohort: 37 paid-active (74%), 6 canceled (12%), 4 past due (8%), and 3 on billing hold (6%). The categories must be represented as exhaustive and non-overlapping, with counts and corresponding rates; any incorrect or contradictory value fails.\n\nThe document reports $60,480 as gross billed cash and $56,230 as refund-adjusted collected cash, clearly distinguishes the two, and gives the collection ratio as refund-adjusted collected cash divided by gross billed cash: approximately 92.97% (93.0% when rounded to one decimal, or an equivalent decimal such as 0.9297). Incorrect labels, amounts, numerator/denominator, or ratio fail.\n\nThe document reports account-level support effort of 26 minutes at the median and 104 minutes at the 90th percentile, and states that 7 of 50 workspaces, or 14%, exceeded 90 support minutes. Any missing or contradictory support measure fails.\n\nThe document states that paid-source attribution reconciled for all 50 workspaces and that no unsupported applicant entered the admitted path. Missing or contradicted checks fail.\n\nThe document explicitly classifies the 2026 cohort packet as descriptive paid-source evidence because it had no randomized control; it does not treat the packet as causal acquisition, scale, pipeline, or forecast evidence; and it states that the historical packet/card did not itself authorize another paid batch. It must not falsely claim that no later batch was ever approved or run. Missing or contradictory limitations fail.",
  "check_id": "riley_177_03"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_benchmark_card",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_177_00"
      },
      {
        "check_id": "riley_177_01",
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_benchmark_card",
        "path": "args.title",
        "value": "2026 Helio Start paid-search mature benchmark card"
      },
      {
        "check_id": "riley_177_02",
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_benchmark_card",
        "path": "args.folder",
        "value": "Growth"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_benchmark_card",
        "path": "args",
        "criterion": "The document clearly identifies the benchmark as the closed 2026 high-intent Helio Start paid-search cohort and states that all 50 workspaces have complete day-90 outcomes. Missing or contradictory cohort identity, denominator, or maturity status fails.\n\nThe document presents four mutually exclusive workspace outcomes for the 50-workspace cohort: 37 paid-active (74%), 6 canceled (12%), 4 past due (8%), and 3 on billing hold (6%). The categories must be represented as exhaustive and non-overlapping, with counts and corresponding rates; any incorrect or contradictory value fails.\n\nThe document reports $60,480 as gross billed cash and $56,230 as refund-adjusted collected cash, clearly distinguishes the two, and gives the collection ratio as refund-adjusted collected cash divided by gross billed cash: approximately 92.97% (93.0% when rounded to one decimal, or an equivalent decimal such as 0.9297). Incorrect labels, amounts, numerator/denominator, or ratio fail.\n\nThe document reports account-level support effort of 26 minutes at the median and 104 minutes at the 90th percentile, and states that 7 of 50 workspaces, or 14%, exceeded 90 support minutes. Any missing or contradictory support measure fails.\n\nThe document states that paid-source attribution reconciled for all 50 workspaces and that no unsupported applicant entered the admitted path. Missing or contradicted checks fail.\n\nThe document explicitly classifies the 2026 cohort packet as descriptive paid-source evidence because it had no randomized control; it does not treat the packet as causal acquisition, scale, pipeline, or forecast evidence; and it states that the historical packet/card did not itself authorize another paid batch. It must not falsely claim that no later batch was ever approved or run. Missing or contradictory limitations fail.",
        "check_id": "riley_177_03"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/177.yaml

SHA-256: f7f947ff0ddf2136e13246d56c86765bb50d58af4a024750ed0caacb7fda91ac