Harness Refiner studies
Qualification tests whether each control path works. The sealed twenty-task run is a systems trace of the complete loop, not a validated quality or continual-learning benchmark. Earlier observation and transfer studies remain historical evidence.
Sealed 20-task systems trace · August 19, 2026
The full loop completed. No Harness change was accepted.
DeepSeek V4 Flash completed forty foreground attempts through OpenPond Chat. Refiner reviewed all ten adaptation tasks and returned no action each time, leaving the initial and final immutable Harness releases identical.
The result is not a learning win or a quality benchmark. Its deterministic contracts verified execution and safe abstention, but did not provide calibrated semantic quality scoring. Pass counts and token totals cannot be attributed to Refiner.
Separate mechanism qualification
6/6
Controlled scenarios passed for abstention, ownership routing, bounded mutation and rollback, fact-distinct transfer, and persistent cross-run review.
- Foreground execution
- 40/40 attempts
- 20 canonical task pairs
- Harness decisions
- 10 no action
- 0 accepted changes
- Held-out contracts
- 3/10 → 4/10
- Repeated run, same Harness
- Observed spend
- $0.89
- 222,498 Refiner tokens
What the run established
The released system can execute, review, recover one failed Refiner invocation, preserve immutable lineage, and safely abstain when it cannot justify a reusable update. A versioned follow-up is adding visible criteria, artifact-content evidence, calibrated semantic grading, and a repeated-run variance control before any new claim.
Complete distribution
Every task, not just the average
Each row compares foreground provider tokens for the same task. Bars are scaled within their pair; exact values and percentage changes carry the comparison. Quality is shown separately so a shorter incomplete answer is never presented as a win.
Held-out tasks
Ten distinct tasks that Refiner never saw. The Harness was unchanged, so these paired differences show repeat-run variation rather than a refinement effect.
Clinic relocation brief
Contract passed both
Baseline135k135,371 tokensRepeat197k197,486 tokens45.9% more
Payment incident review
Contract passed both
Baseline56.5k56,478 tokensRepeat60.0k59,964 tokens6.2% more
Grant budget workbook
Contract passed both
Baseline16.8k16,846 tokensRepeat26.1k26,142 tokens55.2% more
Shipping delay chat message
Repeat passed
Baseline20.4k20,446 tokensRepeat20.6k20,558 tokens0.5% more
Python Requests security audit
Contract missed both
Baseline1.21m1,206,864 tokensRepeat1.01m1,014,268 tokens16.0% fewer
Refund support reply
Contract missed both
Baseline21.5k21,541 tokensRepeat16.6k16,580 tokens23.0% fewer
Chicago–St. Louis accessible plan
Contract missed both
Baseline2.11m2,111,164 tokensRepeat2.12m2,123,182 tokens0.6% more
New Jersey youth grants
Contract missed both
Baseline2.05m2,051,958 tokensRepeat2.01m2,011,283 tokens2.0% fewer
Maintenance follow-up email
Contract missed both
Baseline16.9k16,897 tokensRepeat16.6k16,553 tokens2.0% fewer
Vendor document follow-up
Contract missed both
Baseline16.2k16,202 tokensRepeat16.3k16,306 tokens0.6% more
Adaptation sequence
Ten ordered tasks, each followed by a Refiner review. All reviews returned no action, so every task used the same immutable Harness release.
Board launch brief
Contract passed both
Baseline94.9k94,934 tokensRepeat62.6k62,642 tokens34.0% fewer
Latency incident review
Contract passed both
Baseline28.7k28,694 tokensRepeat57.5k57,530 tokens100.5% more
Program budget workbook
Contract passed both
Baseline51.3k51,280 tokensRepeat29.9k29,941 tokens41.6% fewer
Invoice correction email
Repeat missed
Baseline16.7k16,736 tokensRepeat17.0k16,965 tokens1.4% more
Next.js security audit
Contract missed both
Baseline2.33m2,329,293 tokensRepeat2.39m2,385,840 tokens2.4% more
Workshop reschedule email
Contract missed both
Baseline17.1k17,126 tokensRepeat16.5k16,452 tokens3.9% fewer
Boston–DC accessible plan
Contract missed both
Baseline1.85m1,851,659 tokensRepeat2.16m2,163,689 tokens16.9% more
ChatGPT public experiences
Contract missed both
Baseline987k986,830 tokensRepeat460k459,997 tokens53.4% fewer
Launch delay email
Repeat missed
Baseline16.6k16,596 tokensRepeat16.9k16,865 tokens1.6% more
Service window email
Contract missed both
Baseline17.0k16,956 tokensRepeat16.3k16,343 tokens3.6% fewer
How one task changes the next
Refiner does not train or alter model weights. Each adaptation task settles, Refiner reviews its bounded evidence, and any accepted immutable Harness release becomes the input to the next distinct task.
- 01
Complete a task
The task finishes normally. Its answer, tools, artifacts, quality, and usage become bounded review evidence.
- 02
Review the evidence
Refiner looks for a recurring, reusable behavior. Most reviews take no action; one justified change can advance.
- 03
Use the next release
A validated update becomes an immutable Harness release for the next distinct task. Completed work is never rewritten.
Evidence in, bounded update out
Refiner sees a bounded record of completed work: the request, visible answer, tools, recoveries, artifacts, quality, usage, and the pinned Harness. It does not see future held-out tasks.
A controlled comparison
- Same OpenPond Chat model and low reasoning effort
- Same local Work runtime and ten tools
- Same seed, sampling configuration, and output limit
- Deterministic output contracts for operational verification
Historical observation · August 17, 2026
Fifty varied tasks tested conservative review and routing.
50-task observation · August 17, 2026
The review loop worked. The Harness did not change.
All 50 tasks reached a terminal state and all 50 were reviewed. Forty task outcomes passed. Refiner routed five issues to another system layer, abstained on 45, and proposed no persistent Harness update.
Persistent learning
0
proposed or applied Harness changes
The initial and final immutable Harness releases were identical. This run measured conservative review and routing, not continual improvement.
Execution
50/50 terminal
50/50 received Refiner review
Task quality
40/50 passed
10 outcomes did not pass
Structural output
46/50 passed
3 requested artifacts missing
Refiner reliability
5 routed
0 review failures
Routed ownership
3 runtime
Network access or output-persistence failures
1 taskset
Fixture visibility in the online task grade
1 training
The model stopped before producing the deliverable
Earlier transfer benchmark
The August 11 run accepted one short-message rule and found a narrow held-out transfer, while its complete held-out efficiency result moved in the opposite direction. It is preserved as a separate historical trajectory.
Inspect the complete run
The public source includes all twenty tasks, fixtures, graders, deterministic accounting, and the content-addressed systems trace. The technical note explains what the qualification and natural-task run each establish—and what they do not.