Brier-1 writes the Institute's forecasts — every prediction published with its reasoning and scored when reality arrives. Packs are how human expertise enters that loop: small, versioned, falsifiable units of judgment that have to improve the forecasts to earn their place, and are measured against reality once they do.
An expert's encoded judgment about one class of question.
Composes the applicable packs into its reasoning and predicts.
Published with its reasoning and the packs that shaped it.
Resolved and Brier-scored — credit flows back to the packs.
Most forecasting platforms aggregate predictions. The Institute aggregates judgment — and scores it. A pack is an expert asserting “this way of thinking about this kind of question makes the forecasts better,” in a form the system can test, version, attribute, and retire. It is not a contribution of text. It is a wager that has to win on a scoreboard.
If expertise transfers, it shows up as a lower Brier score. If it doesn't, the pack doesn't get in — and saying so is the whole point.
A pack manifest carries five kinds of judgment, plus the record that admitted it:
The transferable habits of a good forecaster — base-rate anchoring, mechanism decomposition, stating an interval before a point estimate, a premortem. Judgment that usually lives in one expert's head, written down where it can be reused and tested.
For questions a model can actually compute, the pack hands the agent the model. A policy-outcome pack runs PolicyEngine or the open Axiom rules engine — deterministic eligibility and benefit math, traceable to the statute — so a “what will this reform cost / who qualifies” forecast rests on computation, not a hunch.
Anyone can propose a pack. Whether it gets used is decided by a two-tier gate — cheap checks on every change, and an expensive verdict against resolved questions.
Scoring runs on a time-split holdout — train on questions resolved before a cutoff, score on questions resolved after. A pack tuned to look good on the past can't quietly overfit the test; it has to generalize forward, which is the only thing a forecast cares about.
Every forecast records the packs and versions that shaped it. Calibration can then be segmented by pack generation, and an expert's contribution becomes a measurable number: the Brier delta their pack moved. Credit isn't a byline — it's an effect on accuracy.
Illustrative figures. A pack that stops helping — because the world moved or its evidence went stale — is retired the same way it was admitted: by the score, in the open.
Brier-1 reasons through a fixed elicitation flow. Applicable packs compose into the relevant steps — they refine the method, they don't replace it.
And when a question can't be honestly answered, the agent has to say so — with a typed escape, never a confident guess. That refusal is the framework's thesis in miniature.
KPI unresolvable as stated
no defensible base rate
question ill-posed
If you think a kind of evidence, a model, a mechanism, or a piece of open software would make the forecasts better, that's a pack. Write it down, and let the scoreboard decide. The whole system is open — the agent, the packs, the gate, and the record.
Brier and its pack system are open source, of a piece with the rest of the stack — open source, open data, open weights, open predictions.