Harness Refiner studies

Qualification tests whether each control path works. The sealed twenty-task run is a systems trace of the complete loop, not a validated quality or continual-learning benchmark. Earlier observation and transfer studies remain historical evidence.

Sealed 20-task systems trace · August 19, 2026

The full loop completed. No Harness change was accepted.

DeepSeek V4 Flash completed forty foreground attempts through OpenPond Chat. Refiner reviewed all ten adaptation tasks and returned no action each time, leaving the initial and final immutable Harness releases identical.

The result is not a learning win or a quality benchmark. Its deterministic contracts verified execution and safe abstention, but did not provide calibrated semantic quality scoring. Pass counts and token totals cannot be attributed to Refiner.

Separate mechanism qualification

6/6

Controlled scenarios passed for abstention, ownership routing, bounded mutation and rollback, fact-distinct transfer, and persistent cross-run review.

Foreground execution
40/40 attempts
20 canonical task pairs
Harness decisions
10 no action
0 accepted changes
Held-out contracts
3/10 → 4/10
Repeated run, same Harness
Observed spend
$0.89
222,498 Refiner tokens

What the run established

The released system can execute, review, recover one failed Refiner invocation, preserve immutable lineage, and safely abstain when it cannot justify a reusable update. A versioned follow-up is adding visible criteria, artifact-content evidence, calibrated semantic grading, and a repeated-run variance control before any new claim.

Complete distribution

Every task, not just the average

Each row compares foreground provider tokens for the same task. Bars are scaled within their pair; exact values and percentage changes carry the comparison. Quality is shown separately so a shorter incomplete answer is never presented as a win.

Held-out tasks

Ten distinct tasks that Refiner never saw. The Harness was unchanged, so these paired differences show repeat-run variation rather than a refinement effect.

  1. Clinic relocation brief

    Contract passed both

    Baseline135k135,371 tokens
    Repeat197k197,486 tokens

    45.9% more

  2. Payment incident review

    Contract passed both

    Baseline56.5k56,478 tokens
    Repeat60.0k59,964 tokens

    6.2% more

  3. Grant budget workbook

    Contract passed both

    Baseline16.8k16,846 tokens
    Repeat26.1k26,142 tokens

    55.2% more

  4. Shipping delay chat message

    Repeat passed

    Baseline20.4k20,446 tokens
    Repeat20.6k20,558 tokens

    0.5% more

  5. Python Requests security audit

    Contract missed both

    Baseline1.21m1,206,864 tokens
    Repeat1.01m1,014,268 tokens

    16.0% fewer

  6. Refund support reply

    Contract missed both

    Baseline21.5k21,541 tokens
    Repeat16.6k16,580 tokens

    23.0% fewer

  7. Chicago–St. Louis accessible plan

    Contract missed both

    Baseline2.11m2,111,164 tokens
    Repeat2.12m2,123,182 tokens

    0.6% more

  8. New Jersey youth grants

    Contract missed both

    Baseline2.05m2,051,958 tokens
    Repeat2.01m2,011,283 tokens

    2.0% fewer

  9. Maintenance follow-up email

    Contract missed both

    Baseline16.9k16,897 tokens
    Repeat16.6k16,553 tokens

    2.0% fewer

  10. Vendor document follow-up

    Contract missed both

    Baseline16.2k16,202 tokens
    Repeat16.3k16,306 tokens

    0.6% more

Adaptation sequence

Ten ordered tasks, each followed by a Refiner review. All reviews returned no action, so every task used the same immutable Harness release.

  1. Board launch brief

    Contract passed both

    Baseline94.9k94,934 tokens
    Repeat62.6k62,642 tokens

    34.0% fewer

  2. Latency incident review

    Contract passed both

    Baseline28.7k28,694 tokens
    Repeat57.5k57,530 tokens

    100.5% more

  3. Program budget workbook

    Contract passed both

    Baseline51.3k51,280 tokens
    Repeat29.9k29,941 tokens

    41.6% fewer

  4. Invoice correction email

    Repeat missed

    Baseline16.7k16,736 tokens
    Repeat17.0k16,965 tokens

    1.4% more

  5. Next.js security audit

    Contract missed both

    Baseline2.33m2,329,293 tokens
    Repeat2.39m2,385,840 tokens

    2.4% more

  6. Workshop reschedule email

    Contract missed both

    Baseline17.1k17,126 tokens
    Repeat16.5k16,452 tokens

    3.9% fewer

  7. Boston–DC accessible plan

    Contract missed both

    Baseline1.85m1,851,659 tokens
    Repeat2.16m2,163,689 tokens

    16.9% more

  8. ChatGPT public experiences

    Contract missed both

    Baseline987k986,830 tokens
    Repeat460k459,997 tokens

    53.4% fewer

  9. Launch delay email

    Repeat missed

    Baseline16.6k16,596 tokens
    Repeat16.9k16,865 tokens

    1.6% more

  10. Service window email

    Contract missed both

    Baseline17.0k16,956 tokens
    Repeat16.3k16,343 tokens

    3.6% fewer

How one task changes the next

Refiner does not train or alter model weights. Each adaptation task settles, Refiner reviews its bounded evidence, and any accepted immutable Harness release becomes the input to the next distinct task.

  1. 01

    Complete a task

    The task finishes normally. Its answer, tools, artifacts, quality, and usage become bounded review evidence.

  2. 02

    Review the evidence

    Refiner looks for a recurring, reusable behavior. Most reviews take no action; one justified change can advance.

  3. 03

    Use the next release

    A validated update becomes an immutable Harness release for the next distinct task. Completed work is never rewritten.

Evidence in, bounded update out

Refiner sees a bounded record of completed work: the request, visible answer, tools, recoveries, artifacts, quality, usage, and the pinned Harness. It does not see future held-out tasks.

InstructionsSkillsAgentsApproved memoryTool capabilities

A controlled comparison

  • Same OpenPond Chat model and low reasoning effort
  • Same local Work runtime and ten tools
  • Same seed, sampling configuration, and output limit
  • Deterministic output contracts for operational verification

Historical observation · August 17, 2026

Fifty varied tasks tested conservative review and routing.

50-task observation · August 17, 2026

The review loop worked. The Harness did not change.

All 50 tasks reached a terminal state and all 50 were reviewed. Forty task outcomes passed. Refiner routed five issues to another system layer, abstained on 45, and proposed no persistent Harness update.

Persistent learning

0

proposed or applied Harness changes

The initial and final immutable Harness releases were identical. This run measured conservative review and routing, not continual improvement.

Execution

50/50 terminal

50/50 received Refiner review

Task quality

40/50 passed

10 outcomes did not pass

Structural output

46/50 passed

3 requested artifacts missing

Refiner reliability

5 routed

0 review failures

Routed ownership

3 runtime

Network access or output-persistence failures

1 taskset

Fixture visibility in the online task grade

1 training

The model stopped before producing the deliverable

Earlier transfer benchmark

The August 11 run accepted one short-message rule and found a narrow held-out transfer, while its complete held-out efficiency result moved in the opposite direction. It is preserved as a separate historical trajectory.

Historical result data

Inspect the complete run

The public source includes all twenty tasks, fixtures, graders, deterministic accounting, and the content-addressed systems trace. The technical note explains what the qualification and natural-task run each establish—and what they do not.