Skip to introduction
Advance Copy

III · Research & methods

Foundation

Better forecasts. More legible futures.

I'm building Advance Copy to help people live in the future long enough to make better decisions today.

As a solo developer with a product background, my focus is making forecasts useful: future news, alternative timelines, and interfaces that make uncertainty legible. That work needs accurate forecasting at scale—and a system that can keep improving while I build the experience around it.

Advance Copy tells the future story. 100 Worlds makes its forecasts inspectable. Foundation explains how they were built.

01 · The system behind the stories

Helicon is that forecasting system. Amaryl preserves evidence, sources, and conflicting claims so the next forecast can build on what we've learned. Seldon directs experiments, evaluates challengers, and promotes improvements supported by results.

The names borrow from Asimov's Foundation: Seldon sets the direction; Amaryl does the research that makes it possible.

OpenRouter provides a common interface to models from different providers, making model choice another variable we can test.1 The architecture is designed to turn new research into experiments, experiments into better forecasts, and the results into the next research agenda.

Helicon · The forecasting systemArchitecture in development
How Helicon learns from each experiment The founder sets goals and budget. Amaryl supplies research and memory to Seldon. Seldon chooses experiments by value of information. Bounded experiments use model access through OpenRouter. Evaluation compares the challenger with the current baseline. Supported changes are promoted; otherwise the baseline is kept. Both outcomes return results and lessons to Amaryl, informing the next cycle. FounderGoals, budget, and priorities SeldonChoose the next experimentValue of information / cost AmarylResearch & memoryEvidence, sources, conflicting claims Bounded experimentTest a challengerHold the comparison fixed OpenRouterAccess models across providersModel choice is a testable variable EvaluateCompare against the baselineQuality, held-out checks, and cost Promote the improvementUpdate the forecasting champion Keep the baselineRetain failures and open questions Research briefModel accessSupportedNot yetResults & lessons → next cycle

FounderSets goals, budget, and priorities

  1. Amaryl · research & memoryRetains evidence, sources, and conflicting claims.
  2. Seldon · experimentsChooses what to test by value of information and cost.
  3. Run a bounded experimentTest a challenger against fixed inputs.OpenRouter: access models across providers.
  4. Evaluate against the baselineCheck quality, held-out evidence, and cost.

    SupportedPromote the improvement

    Not yetKeep the baseline

  5. Retain the resultBoth outcomes return lessons to Amaryl and shape Seldon's next experiment.↶ The next cycle starts here.
A rejected experiment still improves the next decision. This is the loop we're building; the diagram does not establish end-to-end autonomy or forecasting performance.

02 · Choosing what to learn next

Value of information guides what to try next: which experiment is most likely to change a consequential decision, given its cost? Promising ideas get bounded tests; promotion requires evidence against the current baseline and checks beyond the examples used to develop the change.

Early Seldon cycles have made the priorities clearer. More reasoning effort added cost with little improvement in one comparison. Fixing an omitted question description mattered much more on our small evaluation set. Reliable evidence and honest evaluation are part of the forecasting problem.

03 · Lessons & experiments

The latest lesson from our experiment notebook. Open the record to see the comparison, decision, and limits of the evidence, or expand the timeline to read earlier experiments.

  1. Adopted

    Classifying every candidate closed the gap

    Complete, checkable intermediate outputs make the pipeline easier to trust.

    Read the experiment
Classifying every candidate closed the gap · Adopted

SEL-EXP-2026-007 · 2026-08-23 · Adopted

Classifying every candidate closed the gap

Research-method evaluation · 2 unresolved questions

Question

Would exact-coverage classification help original sources and counterevidence survive selection?

What changed

Classified every source in bounded batches with a structured schema, rejecting malformed or incomplete output before selection. Replayed saved pools before buying fresh research.

Recorded results
MeasureAI adoptionElectricity prices
Candidates classified14 / 1416 / 16
Selected / primary-original10 / 78 / 5
Total pipeline cost$0.957363$0.996370

Both packets passed the report's evidence acceptance criteria. Subsequent operational use classified 148 of 148 new candidates.

Decision

Adopted for bounded evidence-pipeline quality and used for the remaining Issue Zero corpus. Stopped further retrieval engineering until a measured defect appeared.

What we learned

Complete, checkable intermediate outputs make the pipeline easier to trust.

Limits of this result

Adoption concerns evidence quality, not forecasting accuracy. Questions were unresolved; the sample was two, minor misclassifications remained, and no independent blind source-quality review was recorded. Later classification coverage is implementation replication, not an accuracy result.

Source: Seldon experiment ledger · SEL-EXP-2026-007-structured-evidence-classification.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.

View earlier experiments (6)Hide earlier experiments
  1. Inconclusive

    A better selector still needed better labels

    A selection rule cannot repair incomplete or incorrect classifications.

    Read the experiment
  2. Inconclusive

    Finding a source wasn't enough to retain it

    Retrieval quality is lost if the evidence packet drops the best sources.

    Read the experiment
  3. Rejected

    Three calls mostly repeated the same answer

    Aggregation needs useful diversity, not just more samples.

    Read the experiment
  4. Promoted

    The question description changed the result

    Reliable inputs mattered more than a more elaborate reasoning prompt.

    Read the experiment
  5. Rejected

    More reasoning cost more, with little gain

    More compute cannot make up for inputs the model never receives.

    Read the experiment
  6. Rejected

    A longer prompt didn't fix missing evidence

    Check what the model can see before changing how it reasons.

    Read the experiment
A better selector still needed better labels · Inconclusive

SEL-EXP-2026-006 · 2026-08-23 · Inconclusive

A better selector still needed better labels

Research-method evaluation · 2 unresolved questions

Question

Would explicit reservations for primary, YES, NO, and base-rate evidence preserve the most useful bounded packet?

What changed

Added deterministic source selection and a larger candidate pool before the existing document cap.

Recorded results
MeasureAI adoptionElectricity prices
Candidates / selected14 / 1010 / 9
Selection cost$0$0
Selection latency16 ms5 ms

The energy packet retained the EIA outlook. The AI packet still displaced an original Deloitte report with restatements. Truncated research output and memo-position labels had misclassified candidates.

Decision

Retained the selector architecture but withheld promotion of its classification inputs. Required complete structured labels before another research rerun.

What we learned

A selection rule cannot repair incomplete or incorrect classifications.

Limits of this result

Two unresolved questions and defective input labels; neither a resolved accuracy test nor a clean rejection of the selector alone. Candidate-pool formation also changed, and there was no separate post-run independent review.

Source: Seldon experiment ledger · SEL-EXP-2026-006-objective-aware-selection.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.

Finding a source wasn't enough to retain it · Inconclusive

SEL-EXP-2026-005 · 2026-08-23 · Inconclusive

Finding a source wasn't enough to retain it

Research-method evaluation · 2 unresolved questions

Question

Could bounded native web research improve source quality and provenance for the first two Issue Zero questions?

What changed

Replaced the AskNews path with one bounded Anthropic web-research request, fuller question inputs, and persistent citation provenance.

Recorded results
MeasureAI adoptionElectricity prices
Selected sources1010
Total pipeline cost$0.924395$0.887115

Source quality and traceability improved. Energy research found relevant EIA material, but capturing only the first ten citations discarded later Short-Term Energy Outlook evidence.

Decision

Retained the research capability and provenance repairs. Withheld readiness of the full pipeline and moved to evidence-selection tests.

What we learned

Retrieval quality is lost if the evidence packet drops the best sources.

Limits of this result

These questions were unresolved: changed probabilities are not evidence of improved accuracy. The two-question evaluation changed several pipeline features, had no blind quality scoring, and cannot isolate the contribution of each change.

Source: Seldon experiment ledger · SEL-EXP-2026-005-native-web-research.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.

Three calls mostly repeated the same answer · Rejected

SEL-EXP-2026-004 · 2026-08-18 · Rejected

Three calls mostly repeated the same answer

Forecast experiment · 18 synthetic cases

Question

Would three independent forecast calls reduce error enough to justify roughly triple the cost?

What changed

Compared one call with the median of three calls using the complete-input v2 prompt, the same model, and medium effort.

Recorded results · lower Brier and log loss are better
MeasureOne callThree-call median
Combined mean Brier ↓0.05280.0520
Combined log loss ↓0.21740.2151
Estimated run cost$0.1862$0.5634

Sixteen of eighteen cases had a component probability range of at most two percentage points. The tiny score improvement did not justify the additional estimated cost.

Decision

Kept one component as the default. Retained aggregation capability for a future distribution with measurable component diversity.

What we learned

Aggregation needs useful diversity, not just more samples.

Limits of this result

Eighteen reused synthetic cases, no statistical significance, and no external or prospective replication. This rejects this default change on this set, not aggregation as a general method.

Source: Seldon experiment ledger · SEL-EXP-2026-004-three-component-aggregation.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.

The question description changed the result · Promoted

SEL-EXP-2026-003 · 2026-08-18 · Promoted

The question description changed the result

Input repair · 12 development cases + 6 fresh cases

Question

Would rendering the already-stored question description repair the failure without causing blanket overconfidence?

What changed

Added the omitted description block, with no change to reasoning instructions, model, effort, or component count.

Recorded results · lower Brier and log loss are better
MeasureDescription omittedDescription included
Development Brier ↓0.25810.0601
Development log loss ↓0.69860.2295
Fair-coin control50%50%

Every development case improved or held. The new prompt subsequently scored mean Brier 0.0382 on six newly authored cases.

Decision

Promoted the repair as binary-forecast-v2. Auditing other input fields and testing on external or prospective questions became the next priorities.

What we learned

Reliable inputs mattered more than a more elaborate reasoning prompt.

Limits of this result

The large improvement is a twelve-case development result. The old prompt was not run on the six fresh cases, so that check cannot estimate a paired out-of-sample effect. No separate independent reviewer was recorded; these results do not establish real-world forecasting accuracy.

Source: Seldon experiment ledger · SEL-EXP-2026-003-render-question-description.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.

More reasoning cost more, with little gain · Rejected

SEL-EXP-2026-002 · 2026-08-18 · Rejected

More reasoning cost more, with little gain

Forecast experiment · 12 synthetic cases

Question

Would high inference effort improve quantitative reasoning relative to medium effort?

What changed

Changed only provider reasoning effort from medium to high, using the same pre-description-fix prompt and twelve cases.

Recorded results · lower Brier and log loss are better
MeasureMedium effortHigh effort
Mean Brier ↓0.25810.2579
Mean log loss ↓0.69860.6978
Estimated run cost$0.082525$0.104350

The score difference was a wash, while estimated cost rose about 26%. Three of the four flagged cases were unchanged.

Decision

Kept medium effort as the default. Revisit only when a complete-input model shows a reproducible reasoning weakness worth the extra cost.

What we learned

More compute cannot make up for inputs the model never receives.

Limits of this result

Both variants omitted the question description. The result does not rule out higher effort on other tasks or complete inputs. Sample size was twelve; no independent replication, provider-bill reconciliation, or measured latency was available.

Source: Seldon experiment ledger · SEL-EXP-2026-002-high-inference-effort.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.

A longer prompt didn't fix missing evidence · Rejected

SEL-EXP-2026-001 · 2026-08-18 · Rejected

A longer prompt didn't fix missing evidence

Forecast experiment · 12 synthetic cases

Question

Would explicit reference-class, base-rate, and adjustment stages improve probability formation?

What changed

Added a three-stage base-rate instruction while holding the model, effort, cases, and evidence policy fixed.

Recorded results · lower Brier and log loss are better
MeasureOriginal promptStaged prompt
Mean Brier ↓0.25040.2551
Mean log loss ↓0.68500.6964
Estimated run cost$0.08065$0.19632

Both scores worsened and estimated cost increased to about 2.4 times the original run. Later inspection found that both prompts omitted the question description containing decisive facts.

Decision

Rejected this exact prompt change. Input integrity and stronger evaluations moved ahead of further prompt elaboration.

What we learned

Check what the model can see before changing how it reasons.

Limits of this result

Twelve hand-authored synthetic cases; no statistical significance or independent review. This tests staging with incomplete inputs, not the value of outside-view methods in general. No separate pre-result protocol timestamp was recovered.

Source: Seldon experiment ledger · SEL-EXP-2026-001-base-rate-prompt-staging.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.

Foundation · Experiment record

What still needs to be demonstrated

AI forecasters have reported superforecaster-level results on specific benchmarks.2 The opportunity is to turn that progress into tools people can actually use. Those results aren't Helicon's results, and our small-set gains don't establish broader accuracy. Recursive improvement is the loop I'm building toward; its automation and performance still need to be demonstrated.

Learning in the open

Foundation is where I'll share methods, experiments, and failures as they become ready for publication. It's also my vehicle for learning deeply about forecasting while helping advance it—and building the accuracy needed to experiment seriously with how people understand the future.

Generalizable research will be shared after a defined partner-first window. Confidential partner inputs stay private.

References

  1. OpenRouter documentation. A common API across model providers.
  2. AIA Forecaster: Technical Report. Reports parity with human superforecasters on ForecastBench; this is a benchmark-specific result from another system.