I'm building Advance Copy to help people live in the future long enough to make better decisions today.
As a solo developer with a product background, my focus is making forecasts useful: future news, alternative timelines, and interfaces that make uncertainty legible. That work needs accurate forecasting at scale—and a system that can keep improving while I build the experience around it.
Advance Copy tells the future story. 100 Worlds makes its forecasts inspectable. Foundation explains how they were built.
01 · The system behind the stories
Helicon is that forecasting system. Amaryl preserves evidence, sources, and conflicting claims so the next forecast can build on what we've learned. Seldon directs experiments, evaluates challengers, and promotes improvements supported by results.
The names borrow from Asimov's Foundation: Seldon sets the direction; Amaryl does the research that makes it possible.
OpenRouter provides a common interface to models from different providers, making model choice another variable we can test.1 The architecture is designed to turn new research into experiments, experiments into better forecasts, and the results into the next research agenda.
FounderSets goals, budget, and priorities
- Amaryl · research & memoryRetains evidence, sources, and conflicting claims.
- Seldon · experimentsChooses what to test by value of information and cost.
- Run a bounded experimentTest a challenger against fixed inputs.OpenRouter: access models across providers.
- Evaluate against the baselineCheck quality, held-out evidence, and cost.
SupportedPromote the improvement
Not yetKeep the baseline
- Retain the resultBoth outcomes return lessons to Amaryl and shape Seldon's next experiment.↶ The next cycle starts here.
02 · Choosing what to learn next
Value of information guides what to try next: which experiment is most likely to change a consequential decision, given its cost? Promising ideas get bounded tests; promotion requires evidence against the current baseline and checks beyond the examples used to develop the change.
Early Seldon cycles have made the priorities clearer. More reasoning effort added cost with little improvement in one comparison. Fixing an omitted question description mattered much more on our small evaluation set. Reliable evidence and honest evaluation are part of the forecasting problem.
03 · Lessons & experiments
The latest lesson from our experiment notebook. Open the record to see the comparison, decision, and limits of the evidence, or expand the timeline to read earlier experiments.
Classifying every candidate closed the gap · Adopted
SEL-EXP-2026-007 · 2026-08-23 · Adopted
Classifying every candidate closed the gap
Research-method evaluation · 2 unresolved questions
Question
Would exact-coverage classification help original sources and counterevidence survive selection?
What changed
Classified every source in bounded batches with a structured schema, rejecting malformed or incomplete output before selection. Replayed saved pools before buying fresh research.
| Measure | AI adoption | Electricity prices |
|---|---|---|
| Candidates classified | 14 / 14 | 16 / 16 |
| Selected / primary-original | 10 / 7 | 8 / 5 |
| Total pipeline cost | $0.957363 | $0.996370 |
Both packets passed the report's evidence acceptance criteria. Subsequent operational use classified 148 of 148 new candidates.
Decision
Adopted for bounded evidence-pipeline quality and used for the remaining Issue Zero corpus. Stopped further retrieval engineering until a measured defect appeared.
What we learned
Complete, checkable intermediate outputs make the pipeline easier to trust.
Limits of this result
Adoption concerns evidence quality, not forecasting accuracy. Questions were unresolved; the sample was two, minor misclassifications remained, and no independent blind source-quality review was recorded. Later classification coverage is implementation replication, not an accuracy result.
Source: Seldon experiment ledger · SEL-EXP-2026-007-structured-evidence-classification.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.
View earlier experiments (6)Hide earlier experiments
-
A better selector still needed better labels
A selection rule cannot repair incomplete or incorrect classifications.
Read the experiment -
Finding a source wasn't enough to retain it
Retrieval quality is lost if the evidence packet drops the best sources.
Read the experiment -
Three calls mostly repeated the same answer
Aggregation needs useful diversity, not just more samples.
Read the experiment -
The question description changed the result
Reliable inputs mattered more than a more elaborate reasoning prompt.
Read the experiment -
More reasoning cost more, with little gain
More compute cannot make up for inputs the model never receives.
Read the experiment -
A longer prompt didn't fix missing evidence
Check what the model can see before changing how it reasons.
Read the experiment
A better selector still needed better labels · Inconclusive
SEL-EXP-2026-006 · 2026-08-23 · Inconclusive
A better selector still needed better labels
Research-method evaluation · 2 unresolved questions
Question
Would explicit reservations for primary, YES, NO, and base-rate evidence preserve the most useful bounded packet?
What changed
Added deterministic source selection and a larger candidate pool before the existing document cap.
| Measure | AI adoption | Electricity prices |
|---|---|---|
| Candidates / selected | 14 / 10 | 10 / 9 |
| Selection cost | $0 | $0 |
| Selection latency | 16 ms | 5 ms |
The energy packet retained the EIA outlook. The AI packet still displaced an original Deloitte report with restatements. Truncated research output and memo-position labels had misclassified candidates.
Decision
Retained the selector architecture but withheld promotion of its classification inputs. Required complete structured labels before another research rerun.
What we learned
A selection rule cannot repair incomplete or incorrect classifications.
Limits of this result
Two unresolved questions and defective input labels; neither a resolved accuracy test nor a clean rejection of the selector alone. Candidate-pool formation also changed, and there was no separate post-run independent review.
Source: Seldon experiment ledger · SEL-EXP-2026-006-objective-aware-selection.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.
Finding a source wasn't enough to retain it · Inconclusive
SEL-EXP-2026-005 · 2026-08-23 · Inconclusive
Finding a source wasn't enough to retain it
Research-method evaluation · 2 unresolved questions
Question
Could bounded native web research improve source quality and provenance for the first two Issue Zero questions?
What changed
Replaced the AskNews path with one bounded Anthropic web-research request, fuller question inputs, and persistent citation provenance.
| Measure | AI adoption | Electricity prices |
|---|---|---|
| Selected sources | 10 | 10 |
| Total pipeline cost | $0.924395 | $0.887115 |
Source quality and traceability improved. Energy research found relevant EIA material, but capturing only the first ten citations discarded later Short-Term Energy Outlook evidence.
Decision
Retained the research capability and provenance repairs. Withheld readiness of the full pipeline and moved to evidence-selection tests.
What we learned
Retrieval quality is lost if the evidence packet drops the best sources.
Limits of this result
These questions were unresolved: changed probabilities are not evidence of improved accuracy. The two-question evaluation changed several pipeline features, had no blind quality scoring, and cannot isolate the contribution of each change.
Source: Seldon experiment ledger · SEL-EXP-2026-005-native-web-research.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.
Three calls mostly repeated the same answer · Rejected
SEL-EXP-2026-004 · 2026-08-18 · Rejected
Three calls mostly repeated the same answer
Forecast experiment · 18 synthetic cases
Question
Would three independent forecast calls reduce error enough to justify roughly triple the cost?
What changed
Compared one call with the median of three calls using the complete-input v2 prompt, the same model, and medium effort.
| Measure | One call | Three-call median |
|---|---|---|
| Combined mean Brier ↓ | 0.0528 | 0.0520 |
| Combined log loss ↓ | 0.2174 | 0.2151 |
| Estimated run cost | $0.1862 | $0.5634 |
Sixteen of eighteen cases had a component probability range of at most two percentage points. The tiny score improvement did not justify the additional estimated cost.
Decision
Kept one component as the default. Retained aggregation capability for a future distribution with measurable component diversity.
What we learned
Aggregation needs useful diversity, not just more samples.
Limits of this result
Eighteen reused synthetic cases, no statistical significance, and no external or prospective replication. This rejects this default change on this set, not aggregation as a general method.
Source: Seldon experiment ledger · SEL-EXP-2026-004-three-component-aggregation.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.
The question description changed the result · Promoted
SEL-EXP-2026-003 · 2026-08-18 · Promoted
The question description changed the result
Input repair · 12 development cases + 6 fresh cases
Question
Would rendering the already-stored question description repair the failure without causing blanket overconfidence?
What changed
Added the omitted description block, with no change to reasoning instructions, model, effort, or component count.
| Measure | Description omitted | Description included |
|---|---|---|
| Development Brier ↓ | 0.2581 | 0.0601 |
| Development log loss ↓ | 0.6986 | 0.2295 |
| Fair-coin control | 50% | 50% |
Every development case improved or held. The new prompt subsequently scored mean Brier 0.0382 on six newly authored cases.
Decision
Promoted the repair as binary-forecast-v2. Auditing other input fields and testing on external or prospective questions became the next priorities.
What we learned
Reliable inputs mattered more than a more elaborate reasoning prompt.
Limits of this result
The large improvement is a twelve-case development result. The old prompt was not run on the six fresh cases, so that check cannot estimate a paired out-of-sample effect. No separate independent reviewer was recorded; these results do not establish real-world forecasting accuracy.
Source: Seldon experiment ledger · SEL-EXP-2026-003-render-question-description.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.
More reasoning cost more, with little gain · Rejected
SEL-EXP-2026-002 · 2026-08-18 · Rejected
More reasoning cost more, with little gain
Forecast experiment · 12 synthetic cases
Question
Would high inference effort improve quantitative reasoning relative to medium effort?
What changed
Changed only provider reasoning effort from medium to high, using the same pre-description-fix prompt and twelve cases.
| Measure | Medium effort | High effort |
|---|---|---|
| Mean Brier ↓ | 0.2581 | 0.2579 |
| Mean log loss ↓ | 0.6986 | 0.6978 |
| Estimated run cost | $0.082525 | $0.104350 |
The score difference was a wash, while estimated cost rose about 26%. Three of the four flagged cases were unchanged.
Decision
Kept medium effort as the default. Revisit only when a complete-input model shows a reproducible reasoning weakness worth the extra cost.
What we learned
More compute cannot make up for inputs the model never receives.
Limits of this result
Both variants omitted the question description. The result does not rule out higher effort on other tasks or complete inputs. Sample size was twelve; no independent replication, provider-bill reconciliation, or measured latency was available.
Source: Seldon experiment ledger · SEL-EXP-2026-002-high-inference-effort.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.
A longer prompt didn't fix missing evidence · Rejected
SEL-EXP-2026-001 · 2026-08-18 · Rejected
A longer prompt didn't fix missing evidence
Forecast experiment · 12 synthetic cases
Question
Would explicit reference-class, base-rate, and adjustment stages improve probability formation?
What changed
Added a three-stage base-rate instruction while holding the model, effort, cases, and evidence policy fixed.
| Measure | Original prompt | Staged prompt |
|---|---|---|
| Mean Brier ↓ | 0.2504 | 0.2551 |
| Mean log loss ↓ | 0.6850 | 0.6964 |
| Estimated run cost | $0.08065 | $0.19632 |
Both scores worsened and estimated cost increased to about 2.4 times the original run. Later inspection found that both prompts omitted the question description containing decisive facts.
Decision
Rejected this exact prompt change. Input integrity and stronger evaluations moved ahead of further prompt elaboration.
What we learned
Check what the model can see before changing how it reasons.
Limits of this result
Twelve hand-authored synthetic cases; no statistical significance or independent review. This tests staging with incomplete inputs, not the value of outside-view methods in general. No separate pre-result protocol timestamp was recovered.
Source: Seldon experiment ledger · SEL-EXP-2026-001-base-rate-prompt-staging.md
Summary compiled 7 September 2026 from repository revision 6d8a6a6. Costs are recorded pipeline estimates, not independently reconciled provider bills.
What still needs to be demonstrated
AI forecasters have reported superforecaster-level results on specific benchmarks.2 The opportunity is to turn that progress into tools people can actually use. Those results aren't Helicon's results, and our small-set gains don't establish broader accuracy. Recursive improvement is the loop I'm building toward; its automation and performance still need to be demonstrated.
Learning in the open
Foundation is where I'll share methods, experiments, and failures as they become ready for publication. It's also my vehicle for learning deeply about forecasting while helping advance it—and building the accuracy needed to experiment seriously with how people understand the future.
Generalizable research will be shared after a defined partner-first window. Confidential partner inputs stay private.
References
- OpenRouter documentation. A common API across model providers.
- AIA Forecaster: Technical Report. Reports parity with human superforecasters on ForecastBench; this is a benchmark-specific result from another system.