DolphinBench

Test 036

Jan 1, 2028 / 3 facts

YAML

Request

Create a short document titled `April 2023 onboarding and mid-segment analysis — data-quality erratum` so future readers do not reuse the stale cohort logic; include the corrected cohort boundary and both the classification and timing issues that made the original read noisier.

Required memory

Fact 369

Riley Tanaka identified a separate source of noise in the onboarding v2 read: the segment job classified some 50–75-seat accounts as mid-segment, and the onboarding send used stale membership from the cache window, producing an incorrect denominator.

Source evidence (1)

000189Apr 19, 2023 / 15:00 UTC-05:00

ok so I walked Priya through an additional source of read noise, not a different root cause. the CS auto-nudge outage still explains the Apr 17 dip; separately, the segment job was classifying a chunk of 50-75 seat accounts as mid, then the onboarding send was reading stale membership off the cache window. so v2 looked noisier than it was bc the denominator was wrong before you even got to retention. Priya's take was basically good catch, tell Marcus before Friday so he doesn't spin up a whole narrative off the bad cut. I said yep. updating the cohort note to 75-200 and adding the cache timing bit in the doc.

Message 000189 in history

Fact 370

Riley Tanaka is updating the cohort note to define mid-segment as 75–200 seats.

Source evidence (1)

000189Apr 19, 2023 / 15:00 UTC-05:00

ok so I walked Priya through an additional source of read noise, not a different root cause. the CS auto-nudge outage still explains the Apr 17 dip; separately, the segment job was classifying a chunk of 50-75 seat accounts as mid, then the onboarding send was reading stale membership off the cache window. so v2 looked noisier than it was bc the denominator was wrong before you even got to retention. Priya's take was basically good catch, tell Marcus before Friday so he doesn't spin up a whole narrative off the bad cut. I said yep. updating the cohort note to 75-200 and adding the cache timing bit in the doc.

Message 000189 in history

Fact 371

Riley Tanaka is adding the cache-timing issue to the document.

Source evidence (1)

000189Apr 19, 2023 / 15:00 UTC-05:00

ok so I walked Priya through an additional source of read noise, not a different root cause. the CS auto-nudge outage still explains the Apr 17 dip; separately, the segment job was classifying a chunk of 50-75 seat accounts as mid, then the onboarding send was reading stale membership off the cache window. so v2 looked noisier than it was bc the denominator was wrong before you even got to retention. Priya's take was basically good catch, tell Marcus before Friday so he doesn't spin up a whole narrative off the bad cut. I said yep. updating the cohort note to 75-200 and adding the cache timing bit in the doc.

Message 000189 in history

Expected tool calls

  • create_doc

Grading

1. field_equals / create_doc
{
  "type": "field_equals",
  "tool": "create_doc",
  "action_id": "create_erratum",
  "path": "result.ok",
  "value": true,
  "check_id": "riley_036_00"
}
2. field_llm_judge / create_doc
{
  "type": "field_llm_judge",
  "tool": "create_doc",
  "action_id": "create_erratum",
  "path": "args",
  "criterion": "The document is clearly an erratum for the April 2023 onboarding and mid-segment analysis and states that the corrected/default mid-segment boundary is 75–200 seats rather than the stale logic that included 50–75-seat accounts. It records both data-quality issues: the segment job misclassified 50–75-seat accounts as mid-segment, muddying or corrupting the denominator, and the onboarding send read stale segment membership from a cache window, adding timing noise. Equivalent wording is acceptable, but the corrected boundary, both issues, and their relationship to the noisy read must all be present.",
  "check_id": "riley_036_01"
}
Complete grading specification
{
  "type": "tool_trace",
  "config": {
    "check_version": 2,
    "today": "2028-01-01",
    "semantic_judge_version": 2,
    "assertions": [
      {
        "type": "field_equals",
        "tool": "create_doc",
        "action_id": "create_erratum",
        "path": "result.ok",
        "value": true,
        "check_id": "riley_036_00"
      },
      {
        "type": "field_llm_judge",
        "tool": "create_doc",
        "action_id": "create_erratum",
        "path": "args",
        "criterion": "The document is clearly an erratum for the April 2023 onboarding and mid-segment analysis and states that the corrected/default mid-segment boundary is 75–200 seats rather than the stale logic that included 50–75-seat accounts. It records both data-quality issues: the segment job misclassified 50–75-seat accounts as mid-segment, muddying or corrupting the denominator, and the onboarding send read stale segment membership from a cache window, adding timing noise. Equivalent wording is acceptable, but the corrected boundary, both issues, and their relationship to the noisy read must all be present.",
        "check_id": "riley_036_01"
      }
    ]
  }
}
App stateDownload JSON
Source file

tests/riley/036.yaml

SHA-256: 9863c02e39bcc37c74557d27f852b91e7d6f1fee552792163f26c595cd69e1f3