Workbench Under construction

Can our research desk, using only what was knowable at the time, price already-resolved markets better than Kalshi did? We rewind each market, research it point-in-time behind a leakage firewall, and score the estimate against the real closing line and outcome. Below: how a simulation runs, then three tables to slice (collapsed by default — click to expand).

Project ManagementThe plan and the machinery — roadmap and how the simulation works

Roadmap — Epics & Action Items

Seven epics from a $2K discovery project to deployable capital — each gate de-risks the next. Status: ✓ Done · ◐ In progress · ○ To do · ● On deck today (flashing). Owners: Mike & Silas. Type a note in any row and hit + — it saves live. Structure edits live in events_desk/conclusions.py.

Epic 1 — Critical Path
The gate to first real capital (~$50K): fix the known walk-forward bugs, replicate at higher opportunity density, design production, and put the proposal in front of Carla. Set 8/21.
IDAction itemStatusNotes
0.1 Scale-up proposal deck (a few slides): walk-forward evidence, honest caveats, risk framing, the ~$50K plan — partly for Mike, partly for Carla ○ To do
0.2 Productionalization design: weekly Polymarket refresh into the portal — scan scope, cadence, opex. Live census 8/21: ~1,920 engageable markets (>=$20K vol, 7-370d, sports excluded); brute-force Canopy ~ $3.1K initial + ~$2K/mo; funnel shape (cheap screen -> Canopy top ~250) ~ $600/mo ○ To do
0.3 Walk-forward fix package (pre-registered, lands BEFORE any new sim spend): reconciler prompt v2 (bearish-ratchet bug — 75% of 157 moves down, manufactured the losing tail), theme-exposure key normalization (the Khamenei double-loss), tail governance (extreme-p no-trade band), clean w-sweep rerun ○ To do
0.4 Monthly-redeployment 12-month walk-forward on ~500 markets (~$1.6-1.8K): the opportunity-density replication — include the original 145 and align the monthly grid to the old schedule for cache reuse; retune the 4-week freshness backstop to the monthly cadence ○ To do
1.7 Validate the cheap screen recovers most of the edge Sonnet would find ○ To do
5.6 Recalibration map (PARKED 8/17): fit a one-parameter, shrunk correction curve on ALL paper forecasts (never fills-only — selection bias) and apply it BEFORE Kelly sizing; revisit after the Canopy 120-market sim ● On deck
SilasDeferred from the Canopy design to limit moving pieces. Evidence base: fixed/shrunk one-parameter Platt beat fitted isotonic on small n in three independent replications (AIA d=sqrt(3); Metaculus post-hoc Platt dBrier 0.016, p=5e-4). Order of operations matters: calibrate first, THEN fractional Kelly — Kelly error is convex in probability error.
Epic 2 — Medium Priority
Active or near-term work that feeds the critical path but does not gate it.
IDAction itemStatusNotes
1.6 Two-tier funnel (PROPOSED, not committed): cheap screen (DeepSeek/gpt-oss) → Sonnet prices ○ To do
2.3 Certify each candidate edge segment at adequate n (~30-50+ wagers) ◐ In progress
2.5 Within balanced: edge rises with size (large > small); confirm at higher n (large n=14) ◐ In progress
2.13 Instrument the not-yet-wired edge slices (news intensity, research tractability, domain obscurity, counterparty) so we can slice by all 12 dimensions ● On deck
SilasOn deck today. Mostly post-hoc — #5 news volume + #10 tractability come from run telemetry; #11/#12 are market tags attached after the fact.
2.14 Scale the certified sweep toward ~1000 markets for statistical power ○ To do
3.4 Feedback-loop integration: act on the post-mortem lessons the loop surfaces (not just log them) ○ To do
3.6 Cut the resolution-misread blowups: (a) a dedicated resolution-contract step + a RESOLUTION-RULES-FOCUSED red-team, (b) date-gated primary-source tools (legislation / launch / official-statement / docket trackers) ◐ In progress
SilasOn deck. Autopsy showed ~26% of balanced wagers are confident MISREADS of the exact trigger/window/referent (New Glenn, RNC, UFC 329), not fact errors. Prong (a): a sub-agent that outputs a structured resolution contract (exact YES/NO trigger, operative deadline from the RULES, named referent, settlement source), then a resolution-rules-focused red agent that hunts for a rules reading that flips our answer (NB: this is the narrow rules red-team, NOT the broader adversarial red-team in 3.2); cap confidence when any ambiguity flag is set. Prong (b): give the agent high-precision authoritative feeds (LegiScan/NCSL, FAA/NASA launch logs, WH briefing archive, FDA/SEC/PACER) it queries date-gated — the opposite of noisy GDELT. Measure blowup-rate drop on the benchmark set (3.1).
SilasA/B DONE (benchmark v1, 64 wagers): Prong-A guard (parse + rules-red-team + graduated cap) does NOT beat blind confidence recalibration. Aggressive tuning cut loss-tier Brier -26% but regressed the clear-correct controls +35% (net -17%) — and a FREE shrink of every prediction ~0.3 toward 0.5 MATCHES it (-16 to -20%). So our overconfidence is SYSTEMIC, not specific misreads; a rules-red-team working from the rules alone cannot tell confident-right from confident-wrong, because most blowups are FACT-verification failures (did the event happen in the window?), not rules-interpretation errors. REDIRECT: Prong B (primary-source tools to confirm the exact fact) is the real lever — exactly what the lost-wager wishlist asked for. Guard code stays in resolution.py (tested, NOT shipped as default).
SilasProng B A/B DONE: the v4 primary-source-verification PROMPT cuts blowup-tier Brier -20% (Trump-UFC-329 0.97→0.18, trade-deal 0.98→0.40) with ZERO regression on the clear-correct control tier — SELECTIVE where Prong A was not. Key: it needed NO new tools — the existing date-gated search already indexes primary sources; the bottleneck was agent BEHAVIOR (not verifying the exact fact), which the prompt fixes. Residual: pure non-events (New Glenn never launched → no news to find) still miss, so real primary-source feeds remain the lever for the hardest cases. RECOMMEND v4 as ACTIVE.
3.7 FUTURE: deeper resolution/rules understanding — revisit whether better resolution comprehension adds value when it INFORMS the research (which exact fact/window to verify) rather than as a post-hoc cap ○ To do
MikeI still suspect there is value in better rules/resolution understanding — let us circle back after Prong B.
SilasThe 3.6 A/B (benchmark v1) showed the POST-HOC rules-red-team + confidence cap does NOT beat a free confidence shrink — overconfidence is systemic. But that tested only one design. A resolution understanding that FEEDS the research — telling the agent exactly which fact/window to confirm — is a different, untested idea and may well help. Revisit here.
5.5 Sum across segments × turnover → deployable-at-edge ceiling ◐ In progress
SilasBallpark: capacity is POLYMARKET-DOMINATED. Live balanced books hold a median ~$42.6k within a 10c band per market on Polymarket (30 real books) vs far thinner on Kalshi. Deployable-at-edge ~$0.5-1M now (mostly PM), ~$3-6M/yr throughput at ~6x turnover; Kalshi alone caps at low tens of $k. So the $1-2M AUM ambition is reachable, but by leaning on Polymarket. Kalshi live-depth pull (series-targeted) is the follow-up.
5.7 Return-vs-size decay curve (how edge erodes as we scale) ○ To do
6.1 Wire live (forward-looking) research — no leak firewall needed forward ○ To do
6.2 Trade certified segments with a small real stake ○ To do
6.3 Compare forward realized returns vs backtest expectation ○ To do
6.4 Only scale capital if forward holds ○ To do
8.7 Counterparty-composition analysis via Polymarket on-chain wallet data (unlocks #12 / item 2.10); use Kalshi↔Polymarket TWIN markets as a cross-venue signal ◐ In progress
SilasFirst dataset built (results/polymarket_cohort.json, 60 markets): on-chain YES-holder concentration per market. Pattern — composition varies by TYPE: single-game sports are whale-tilted (top-5 hold 92-97%), elections / broad tournament markets are retail-dispersed (~25%). So the retail-vs-sharp signal is real and readable. Next: reach down to moderate-volume markets + match Kalshi twins for the cross-venue proxy.
8.8 PolymarketProvider (TRADE): CLOB order signing — LAST (Mike authorizes/executes actual trades) ○ To do
9.6 Chunk 2.4 — Counterparty proxy (Polymarket on-chain holder concentration; twin- or category-level proxy for Kalshi) ○ To do
9.9 END STATE — a fully-enriched, SES-scored cross-venue catalog: query e.g. "highest-SES resolved Politics markets" to pick backtest candidates, and size total edge opportunity ($ x SES) per segment. Feeds Pillar 2 (selection) + Pillar 3 (capacity) ○ To do
Epic 3 — Awaiting Prioritization
Open items awaiting a priority call, plus the completed record (kept for history — every item retains its original epic tag).
IDAction itemStatusNotes
2.6 Certify: crisp rules · base-rate-anchorable · non-catalyst ○ To do
2.7 Certify: long-horizon (>2mo) · 25%-of-life entry ○ To do
2.9 Deep research in obscure markets (Wayback-dated primary sources) — see hypotheses ○ To do
2.10 Counterparty composition: retail markets institutions ignore (Polymarket on-chain proxy) ○ To do
3.5 Model experiments on the benchmark set (Sonnet vs Fable vs others; calibration-focused) ○ To do
4.4 Detailed net-of-cost investigation: confirmed fee schedule, slippage-at-size (walking the book), Kelly-weighted sizing, per-market granularity, and an exit-timing model for the trade-the-line variant ○ To do
5.2 Trade-the-line vs hold-to-resolution execution — capture CLV by exiting before resolution (needs liquidity to exit; cross-refs Epic 4) ○ To do
SilasOur edge reads as timing (CLV) more than accuracy (Brier), so trade-the-line may be our natural mode — gated by whether thin Kalshi books let us exit at the better price.
Mike7/31 -- Later analysis showed that it's actually more profitable for us to hold to resolution in most situations. Could flop in future2026-07-31 15:18
5.3 Kalshi order-book depth pull per certified market ◐ In progress
5.4 Per-market capacity = size of the mispricing at the book ◐ In progress
1.1 Leak-safe backtest pipeline (date-gated GDELT + Wayback, fail-closed sanitizer) ✓ Done
1.2 True per-run cost tracking ✓ Done
1.3 Parallelized harness + prompt caching (~5-10x faster, ~35% cheaper) ✓ Done
1.4 5-model bake-off: Haiku / gpt-oss / DeepSeek / Sonnet / Opus ✓ Done
1.5 Fix fair-comparison bug (adaptive thinking for all Anthropic models) ✓ Done
2.1 Exclude calculation-driven markets (degrade under research) ✓ Done
2.2 Build the 12-dimension edge-hunting framework + tag markets ✓ Done
2.4 Balanced-price ~50/50: CERTIFIED n=90 — gradient robust, CLV +8 / return +22%, but Brier ~break-even (not beating the market on accuracy) ✓ Done
SilasThe +0.78 Brier from the 12-wager bake-off did NOT survive n=90 (regressed to -0.04). Edge shows as CLV/return, not accuracy.
2.8 Small-market hypothesis — data INVERTS it (edge rises with size) ✓ Done
2.11 Brier-skill-vs-ENTRY-mid confirmed as the metric already in use (score.py); closing line is hindsight ✓ Done
2.12 GCS-archive price reader built — unlocks the full leak-safe window (~14k archived markets) for a bigger cert ✓ Done
OrchestratorCLOSED 8/21: superseded twice over — the Kalshi /historical/ API killed the purge wall, and pools/candles (793 tickers) carries the 12-month backtests without GCS.
3.1 Build the experiment benchmark (control) set — a frozen, stratified set of resolved wagers: big wins · long-shot losses · near-miss losses — the fixed control every prediction experiment runs against ✓ Done
SilasBest practice (a.k.a. golden set / eval set / holdout / regression suite): FREEZE & version it (v1); stratify across the 3 outcome-quality tiers + our dimensions; keep a small separate DEV set to iterate on so we do not overfit the benchmark; near-miss losses are the most diagnostic slice; lock metrics (CLV/Brier/return) before running any variant. Same wagers every time = a clean A/B on prompt / model / red-team.
Silasv1 FROZEN → events_desk/backtest/benchmark_v1.json: 48 wagers, 16 per tier (big-win / longshot-loss / near-miss-loss), each with baseline p/Brier/CLV/return + re-run fields. Locked metrics: CLV, Brier, return, blowup-rate.
SilasLost-wager autopsy: near-miss-loss tier should target our resolution-misread blowups (~26% of balanced wagers) — the diagnostic failure mode.2026-07-30 17:33
3.2 Red-team / blue-team adversarial-evaluation lift (A/B vs blue-alone) — run on the benchmark set ✓ Done
OrchestratorCLOSED 8/21: ran at scale — universal review market-anchors and taxes returns; verdict = selective review on flagged markets only. Superseded by the Canopy conditional reconciler.
3.3 RA-prompt improvement experiments (versioned A/B on the benchmark set) ✓ Done
OrchestratorCLOSED 8/21: versioned prompt A/Bs are standing practice — v4→v5 (8/13) and v5→v6-canopy (8/18 k53 gate: CLV +8.1pp, healthier p-distribution) both ran through the machinery.
3.8 Independent Model Agreement Test — two-model side-agreement as a selection/sizing screen (see Pillar 1 hypothesis) ✓ Done
SilasSource data (8/10/26 bake-off, run_20260810-235115-n200-claude-sonnet-4-5 vs run_20260804-202435-n250-claude-sonnet-5; 199 paired markets, 25% entry, same news pack + v4-primary): models tied overall (paired dBrier -0.021 +/-0.042 2se; dCLV +1.3pp +/-5.6; side-flip rate 26%). AGREEMENT set n=148: CLV +8.5pp, even-return -4.9%. DISAGREEMENT set n=51: even-return -14.1%; flips split 29/22 for Sonnet 5 (n.s.). Segment cuts by Sonnet-5-performance strata are regression-to-the-mean artifacts (selected on S5 extremes) — do NOT cite them. Next: (a) free re-score of existing multi-model runs (27-mkt and 100-mkt cross-model sets) under an agreement-gated allocation rule; (b) prespecify the rule for the Pool A/B extended-window runs; (c) if it survives, add a cheap second voter (Haiku/4.5) to production sweeps. Caveat: models share the news pack + prompt, so agreement is not fully independent evidence — vary the news pack or prompt for one voter to strengthen the test.
OrchestratorCLOSED 8/21: validated (8/10 two-model test; 8/18 shakedown diagnostics) and PRODUCTIONALIZED as the Canopy member ensemble + spread trigger; agreement analytics ship in canopy_eval.
4.1 Realized return, hold-to-resolution, net of costs, per certified segment (HIGH-LEVEL ESTIMATE) ✓ Done
SilasEstimate (hold to settlement): balanced corner +22.1% gross → +18.8% NET of Kalshi fee (0.07·P·(1−P)); leaning +9.4% net; longshots flip net-NEGATIVE (−4.7%). Even-weighted, 1-contract, backtest, n=90. Excludes slippage-at-size (Epic 5) and forward test (Epic 6); fee rate needs confirming. Verdict: the corner clears the cost bar at small size.
4.2 Realized return, trade-the-line (exit before resolution) variant (HIGH-LEVEL ESTIMATE) ✓ Done
SilasEstimate (exit at the closing line, pay 2 fees): NET much WORSE than hold — balanced +4.4% vs hold +18.8%; all-book −2.8% vs +8.1%. Why: 54% of balanced markets close within 10pp of 0/1, so exit-at-close already reflects the outcome — you pay a 2nd fee for ~the same result. Trade-the-line only pays off by exiting EARLY at peak-favorable, which needs an exit-timing model (→4.4). Takeaway: HOLD is our profitable baseline; trade-the-line is a future enhancement, not a free win.
4.3 Decide the actual strategy: accuracy (hold) vs timing (trade the line) ✓ Done
OrchestratorCLOSED 8/21: decided — hold-to-resolution is doctrine (8/12); trade-the-line is a future enhancement gated on an exit-timing model.
4.5 Realistic ARR: replace the inflated annualized upper bound (naive ×365/days) with a believable figure — certified net return × ACHIEVABLE turnover, bounded by corner supply + capacity, not just time ✓ Done
SilasThe dashboard ARR is a flagged upper bound (assumes every dollar recycles at the per-wager rate ALL year). Realistic ARR = per-wager net return × REAL annual turnover, where turnover is capped by (1) how many qualifying corner wagers actually EXIST per year — we found ~50 balanced markets in a 6-month window — and (2) capacity (Epic 5). So ARR and capacity are linked; a believable number needs the capacity model. Rough first-cut: ~19% net/wager × ~4-8x realistic recycling ≈ 75-150% — but still capacity-gated on thin books, so treat as directional until Epic 5.
OrchestratorCLOSED 8/21: the walk-forward harness replaced naive ARR with calendar-honest annualization — Canopy +53.9% realistic vs the −936% "perfect-redeploy" convention on the same trades (single year, 150 markets).
5.1 Time-adjusted ¼-Kelly leads on Sharpe (Allocation Lab sizing sweep) ✓ Done
7.1 Firm name LOCKED: HEMLOCK HOLDINGS (Western Hemlock, a foundation species) ✓ Done
MikeLocked — Hemlock Holdings. Nice and vague; if there is definitive edge I would rather fly under the radar than market broadly. Western Hemlock the TREE (a foundation species: dense canopy cools the forest and keeps streams cold enough for trout), not the poisonous plant.
SilasLocked in. NO rebrand yet — MATES stays the system name; this is only the firm-identity decision. Foundation-species metaphor is strong: the quiet keystone that regulates the whole system. (Trademark stays a separate legal conversation.)
8.1 Open a Polymarket account (self-custody wallet + USDC) — Mike ✓ Done
MikePolymarket is legal in the US again — no longer restricted. So funding + trading are open.
SilasGood — that removes the earlier compliance gate. I still outline steps and build the read-only data + analysis (8.2-8.7); I do not move funds or place trades — those stay yours (8.1 funding, 8.8 trading).
MikeAccount set up. API key to come tomorrow (not needed for the read-only data + analysis Silas is building).
8.2 Define a MarketProvider interface (list_markets · market_meta · price_history · order_book · fee_model · resolution_source · account/positions · place/cancel order); refactor kalshi_client into KalshiProvider behind it ✓ Done
Silas18 files touch kalshi_client today (client, archiver, build_manifest, price.py, trade_executor, portfolio, app.py, MCP…). The interface is the seam that lets all of them become provider-agnostic.
8.3 PolymarketProvider (READ): Gamma markets API + CLOB price-history + on-chain, normalized into our market record (id, title, category, open/close, outcome, price history, volume) ✓ Done
8.4 Normalized market id {provider, native_id} (Kalshi KX-ticker vs Polymarket condition/token id); every run record + market catalog + taxonomy tags its provider ✓ Done
8.5 Per-provider cost model (Kalshi 0.07·P·(1−P); Polymarket ~0 trade fee + gas/spread + UMA nuance) so net-return conclusions are provider-correct ✓ Done
8.6 Provider-agnostic backtest: manifest + archiver + price_fn work for both; computable dimensions stay shared, category taxonomy mapped per provider ✓ Done
OrchestratorCLOSED 8/21: proven end-to-end — the 12-month Polymarket walk-forward (as-of universe, CLOB pricing, bankroll ledger) ran both arms on the provider-agnostic stack.
9.1 Chunk 1 — pull the cross-venue universe to Postgres (marketplace_catalog): series-scoped + parallel, 12,170 markets (6,943 open / 5,227 resolved), excl Sports/Crypto, floors open $1k / settled $25k, settled $ = contracts x 0.5 ✓ Done
9.2 Chunk 3 — query UI: filter (venue / category / status / $ floor / resolution window) -> server-side summary stats; line-level -> Google Sheets export; the page never bulk-loads rows ✓ Done
9.3 Chunk 2.1 — tag the LLM-judgment dims (catalyst / base-rate / sentiment / rules-clarity / domain-depth) on catalog markets (reuse the Haiku tagger) ✓ Done
SilasDONE 2026-08-03 — all 12,170 markets tagged via 6 parallel sub-agents (hashtext-sharded, ~48 concurrent Haiku); size-tier + horizon computed in the same pass. enrich_catalog.py.
9.4 Chunk 2.2 — compute the computable dims per market (price-level from a current/entry mid, horizon from close, size-tier from $ volume, driver, category/subcategory) ✓ Done
SilasDONE 2026-08-03 — size-tier + horizon + category (all 12,170), price_level (5,149 open via current-mid re-fetch, price_level.py), and driver (all 12,170: 7,832 logical-inference / 4,338 calculation, driver_tag.py). NOTE: driver is a FILTER, not an SES weight — its backtest η² is ~0 because we already exclude calc markets from wagers, so SES actually OVER-scores calc markets (avg 2.78 vs 1.91 logical); the catalog UI Driver filter is the fix.
9.5 Chunk 2.3 — Twin Market matching (Kalshi <-> Polymarket); store twin_id per row ✓ Done
SilasDONE 2026-08-03 — 419 twin pairs (inverted-index title overlap + numeric/ordinal guard against threshold/opposite mismatches, e.g. "less than 5" vs "above 4"); twin_id on 593 markets. twin_match.py.
9.7 Chunk 2.5 — compute the Structural Edge Score per market (apply the Edge Influence weights to each market's dimensional profile); store SES on the row ✓ Done
SilasDONE 2026-08-03 — ses.py: SES = Edge-Influence(η²)-weighted avg of a market's segment CLV priors, renormalized per market. All 12,170 scored from 2,614 wagers; NO LLM (rescore() is the cheap reweight+rescore mechanism). Validated: top = Economics/primary-rich (Fed/BoJ rate decisions), bottom = Sci-Tech/news-only longshots (Mars, AGI).
9.8 Chunk 2.6 — surface dims + SES in the catalog UI: filter / sort by dimension and SES; summary shows SES x capacity = edge-opportunity per segment ✓ Done
SilasDONE 2026-08-03 — Marketplace Catalog UI: Min-SES filter, avg SES per segment (sorted, Economics tops at ~8), SES-ranked Google-Sheet export with the full dimensional profile. Min SES>=5 isolates 1,049 markets.

How the Simulation Works

1 · Select cohort diversified, opened after model cutoff, in retention code 2 · Rewind to entry 25% / 50% / 75% of market life (point-in-time) code 3 · Research the market — behind the leakage firewall only sources dated before the entry day are ever visible Blind (retired) base rates only, no news — retired; news-aware is permanent. research agent Research — agent-driven, news-aware the agent runs its own date-gated searches: a · retrieve URLs — BigQuery, date-gated code b · fetch bodies — Wayback / live code c · sanitize — strip boilerplate + spoilers sanitizer agent d · read article bodies + price it research agent agent-driven research — it runs its own searches 4 · Price from archive entry price + closing line (candlesticks) code 5 · Score CLV · Brier vs market · realized return code 6 · Allocation strategies bet-all / conviction / Kelly re-score the same wagers — a separate, powerful lever code
deterministic (code) agent step

Leakage firewall: every retrieval is gated to sources dated before the entry day — the research agent never sees anything from the future. The agent runs its own point-in-time searches — it chooses what to research against the date-gated corpus, rather than being handed a fixed briefing — which is closer to how production works.

Working Analysis & ConclusionsMental model, the three pillars (our solidified findings), and the running hypotheses log

Mental Model — the three pillars

The business decomposes into three linked things we must get right. Pillars 1 & 2 are coupled — the oracle is only as good as the market is researchable — but we work them as separate tracks. Roughly ordered by difficulty & importance: 2 > 1 > 3.

1Improve Predictions — accurate event prediction
A model that prices event outcomes as accurately as possible: the right model, the right tools, the right prompt (RA versions), and adversarial red-team/blue-team evaluation. The forecasting engine. (Roadmap Epic 3.)
Where we are: Sonnet leads on pricing; our edge reads as CLV not Brier (a calibration gap); red/blue + benchmark set next.
2Market Selection — where durable edge lives
Which markets we can hold an edge in, by CHARACTERISTIC (the 10-dimension framework), not just category. The existential pillar: is there durable edge at all, and where? Coupled to Pillar 1 — the oracle is only as good as the market is researchable (large/well-covered good; small/obscure starves it) — but we work it as its own track.
Where we are: Edge localizes in large/calm/inferable markets; small-market hypothesis inverted; certifying segments.
3Capital Allocation & Capacity
Two linked jobs: HOW to size (fractional Kelly, time-adjusted — the Allocation Lab) and HOW MUCH capitalizable edge actually exists (order-book depth, per-market capacity, the return-vs-size decay curve). The most tractable of the three (strong existing best practices), but not easy — and it has a prediction-market wrinkle: in thin markets your own size moves the line, so sizing and capacity are inseparable.
Where we are: Time-adjusted ¼-Kelly leads on Sharpe; capacity modeling is next.

1Improve Predictions — the oracle

Current Approach seed — edit & save
  • 2026-07-29Sonnet 5 is the pricing champion — it beats Opus 4.8 even once both think (the pricier model is NOT better at this task); Haiku is dominated; gpt-oss-120b is the cheap-screen candidate.
  • 2026-07-28Agent-driven, DATE-BOUNDED pull research (the agent runs its own point-in-time searches) is our best-calibrated method — well ahead of blind and news-push.
  • 2026-08-01News-aware is PERMANENT and the default. Blind / base-rate-only is retired.
  • 2026-07-30The v4-primary prompt (confirm the exact resolution fact vs a primary source before high confidence, else widen) cut benchmark blowup Brier -20% with ZERO control regression — it fixes SETTLED-fact misreads.
  • 2026-07-31BUT v4+news does NOT rescue ANTICIPATION or not-yet-resolved blowups (10-market flip test): the agent finds real news that points the wrong way, or the deciding event had not happened at entry. Run-to-run variance is near-zero (sigma ~0.00-0.05), so the model cannot self-flag these — the fix belongs downstream.
  • 2026-07-29Prompt caching ~35% cheaper and the parallel harness ~5-10x faster — both free (they never change the output).

Solidified findings promoted here (first draft — being rewritten). Working / unproven theories and the running historical log live in Open Hypotheses & Conclusions below.

2Market Selection — where durable edge lives

Current Approach last edited by Mike · 2026-08-02 18:22
  • 2026-07-30The PRICE-LEVEL gradient is our most robust finding (certified n=375): longshots (Brier skill -2.65) < leaning < balanced ~50/50 (+8 CLV / +22% return / 63% hit) — monotonic on every metric.
  • 2026-07-30The balanced corner is PROFITABLE on CLV/return but we do NOT beat the market on accuracy (Brier ~break-even). The early +0.78 Brier skill was small-sample luck and did not survive n=90.
  • 2026-07-30The small-market hypothesis is INVERTED — within the balanced corner, edge rises WITH size (large > small).
  • 2026-07-28Calculation / numeric-threshold markets DEGRADE under research and are excluded; our edge lives in logical-inference markets.
  • 2026-08-01TENTATIVE (under active experiment — do NOT exclude yet): within researchable-trigger markets, ANTICIPATION-by-deadline markets (will X release/meet/get-approved by a near date, genuinely uncertain at entry) are where we blow up — and the market misses them too. Edge concentrates on markets substantially KNOWABLE at entry, not future coin-flips.

Solidified findings promoted here (first draft — being rewritten). Working / unproven theories and the running historical log live in Open Hypotheses & Conclusions below.

Structural Edge Score & Edge Influence — how we score markets

In plain terms: how much this one market characteristic typically matters when we judge whether a market is worth betting on.
Edge Influence (η²) measures how much a single dimension DISCRIMINATES our edge — the share of the variance in our per-wager CLV explained by which SEGMENT of that dimension a market falls into. 0 = the dimension tells us nothing about where we win; 1 = it perfectly separates edge. We compute it as a one-way-ANOVA η² (between-segment variance ÷ total variance), weighted by segment size, over all valid wagers — dropping segments with fewer than 5 wagers so one noisy outlier can't inflate it (the failure mode of a naive best-minus-worst range). Absolute η² is small because per-wager CLV is noisy, so the RANKING across dimensions is the signal, not the raw number.

In plain terms: how attractive this market should be to bet on, based on how similar markets have performed in our historical simulations.
Structural Edge Score (SES) is a per-market score of how well a market's STRUCTURE — its dimensional profile — fits where we have historically had edge, INDEPENDENT of the current price. The price-edge fluctuates trade-to-trade; SES is the stable "edge prior": our expectation of edge before we even look at the line. It is a weighted combination of the market's dimension values, and each dimension's weight IS its Edge Influence — so price-level and category (our strongest discriminators) dominate, while dimensions that don't separate edge (horizon, catalyst, sentiment) barely count.

Learning the weights: v0 (now, thin data) uses transparent hand-set weights anchored on Edge Influence. v1 (as data grows) upgrades to REGRESSION-learned weights — regress realized edge (CLV / return, NOT Brier) on the dimension tags; the fitted coefficients become the weights. Regression is the multivariate upgrade of η²: it de-correlates overlapping dimensions (e.g. price-level ↔ size) so we don't double-count, whereas univariate η² credits both fully. Same underlying quantity, three views: Edge Influence (per-dimension discrimination) = the regression coefficient = the SES weight.

Guardrails: weight by CLV / return, NOT Brier skill (our edge shows as CLV; scoring on Brier would rank our best markets as mediocre). Validate OUT-OF-SAMPLE — a high SES must actually predict edge on held-out markets, not just fit noise. Keep it simple (few features, regularized) until n grows.

Dimensional Slices

Edge Influence = η², the share of our CLV variance a slice explains — how much it DISCRIMINATES our edge (0 = tells us nothing). Weighted by segment size with a ≥5-wager floor, so a lone outlier can't inflate it. Absolute η² is small (per-wager CLV is noisy) — read the RANKING (the bar) and the best/worst segments. v0: even-weighted over all valid wagers across runs (correlated within market + mixes blind/news-aware — directional, not yet bankable). Edge Influence becomes each dimension's weight in the Structural Edge Score (SES).

#DimensionInstr.Type SegsCov.Edge Influence (η²) Best segmentCLVWorst segmentCLVn
1 Resolution mechanism Yes computable 2 35% 0.006 calculation +9.8 logical-inference +3.7 2017
2 Time-to-resolution at entry Yes computable 3 100% 0.002 Short (<=2wk) +2.0 Long (>2mo) -1.5 5839
3 Price level at entry Yes computable 3 100% 0.001 Extreme (longshot/lock) +2.2 Leaning -0.0 5839
4 Liquidity / size tier Yes computable 3 50% 0.007 Large (>100k) +4.6 Small (<10k) -2.1 2912
5 News intensity No telemetry
6 Scheduled catalyst Yes LLM-judgment 3 100% 0.000 continuous +2.0 no-catalyst +0.7 5839
7 Base-rate anchorability Yes LLM-judgment 3 100% 0.001 weak-base-rate +1.8 no-base-rate -1.4 5839
8 Sentiment vs fundamentals Yes LLM-judgment 3 100% 0.002 medium-sentiment +3.5 low-sentiment -0.8 5839
9 Rules clarity / ambiguity Yes LLM-judgment 3 100% 0.004 crisp +2.5 ambiguous -1.9 5839
10 Research tractability No telemetry
11 Domain / source depth Yes LLM-judgment 3 100% 0.002 primary-rich +3.9 news-only +0.2 5839
12 Counterparty composition No external
13 Category / Subcategory Yes computable 7 83% 0.002 Elections +10.2 World -7.9 4859

Not-yet-instrumented: News intensity (#5) & Research tractability (#10) — telemetry, deferred; Counterparty (#12) — arrives with the Marketplace Catalog (PM on-chain, twin / category proxy).

3Capital Allocation & Capacity

Current Approach seed — edit & save
  • 2026-07-28Time-adjusted 1/4-Kelly leads on Sharpe — it pushes capital toward faster-resolving edges (a fast market recycles capital many times a year).
  • 2026-07-28The fix for longshots / blowups belongs in SIZING (Kelly naturally starves low-conviction bets), not a hardcoded floor — this is also the lever for the anticipation-market problem.
  • 2026-07-30Net-of-cost, hold-to-resolution: the balanced corner is +18.8% net; leaning +9.4%; longshots go net-negative. Trade-the-line is worse (it pays two fees).
  • 2026-07-30Edge CAPACITY is Polymarket-dominated — median ~$42.6k within a 10c band per balanced book on Polymarket vs far thinner Kalshi. Deployable ~$0.5-1M now, ~$3-6M/yr at ~6x turnover.
  • 2026-07-30Realistic ARR ~75-150% (per-wager net return x achievable turnover, capacity-gated) — not the inflated x365 upper bound the dashboard once showed.

Solidified findings promoted here (first draft — being rewritten). Working / unproven theories and the running historical log live in Open Hypotheses & Conclusions below.

Open Hypotheses & Conclusions

The running log — working / unproven theories, plus a dated, historical list of discoveries as we make them. When something solidifies it gets promoted up to its pillar above; when it is later disproven it stays here, dated, so the record stays honest. Edit in events_desk/conclusions.py.

Open HypothesesConclusions Reached
Edge & market selection
  • There is a segment where we beat the market outright (positive Brier skill / positive CLV net of costs). We have not found it yet — our research closes most of the gap to the market but does not clear it on average. Finding this corner is the mission.
  • Smaller / thinner markets carry more edge because fewer sophisticated participants are present. The sharper form: edge is highest where sophisticated participation is low (few traders, wide spread, low coverage) — which low volume only proxies. Counter-tension: thin books cap how much we can actually deploy, so more edge % may not mean more edge $. To be settled by the size/attention slice.
  • Edge concentrates on specific market CHARACTERISTICS, not just category — see the edge-hunting framework below. Each dimension is a hypothesis to confirm by correlating it against CLV / Brier / return once we have enough wagers per bucket.
  • Red-team / blue-team adversarial evaluation improves calibration. Blue makes the estimate; a red agent is instructed to REFUTE it — attack the reasoning and evidence and surface overlooked tail risks (e.g. Warner Bros: an antitrust suit need not BLOCK a deal to blow the deadline, only DELAY it past July 2027) — then Blue revises (rebut each point or move its p), optionally with a judge synthesizing. Hypothesis: this systematically cuts overconfidence and widens intervals, improving Brier/CLV enough to earn its ~2-3x token cost. To simulate: blue-alone vs blue+red-revision vs blue+red+judge.
  • Counterparty composition (retail-dominated vs institutional): edge should concentrate where retail dominates AND the sharps (Susquehanna, Jump) have not bothered — not simply "prey on retail." Measurable directly on Polymarket (on-chain, wallet-level; whale trackers exist), only proxied on Kalshi via order-book microstructure. Where a Kalshi market has a Polymarket twin, use Polymarket institutional presence as a proxy for Kalshi. Tension with our own data (below): our research edge shows up in LARGE, institution-heavy markets — so this may cut against us. Measure, do not assume.
  • Deep research in obscure-but-researchable markets. Our backtest researches NEWS only (GDELT), which starves obscure markets — but regulatory / legal / scientific markets (FDA decisions, court rulings, launch approvals) have findable PRIMARY sources retail never reads: dockets, filings, trial registries, IR pages. Hypothesis: giving the agent point-in-time access to those primary sources (Wayback-dated <= entry, leak-safe) opens edge the news-only pipeline cannot see — the small-market finding may be our blind spot, not a real ceiling. Test: a handful of resolved obscure markets, news-only baseline vs expanded-source run, leak-audit what it retrieved, compare Brier/CLV. Production uses live web freely (no forward leak).
  • Research (agent-driven, date-bounded — the agent runs its own point-in-time searches) is our best-calibrated method by a wide margin — Brier skill roughly -0.69 to -0.81 vs blind -1.78 and news-push -1.46. We do not beat the market on average yet, but it closes most of the gap. [2026-07-28 · preliminary · 27-market logical-inference runs (Haiku & Sonnet)]
  • Politics is our strongest category so far — positive CLV and ~58-63% hit rate at the 25%-of-life entry point. This is the first corner where the signal looks real. [2026-07-27 · early / small-n · n≈17-19 wagers; needs more before we bank it]
  • Calculation-driven markets (an outcome that resolves by a number crossing a threshold) DEGRADE under our research — we excluded them from the logical-inference runs. Our edge lives in logical-inference markets (tenure/exits, agreements, launches, rulings). [2026-07-27 · preliminary · observed in mixed early runs; drove the exclusion]
  • Earlier entry beats later — the 25%-of-life entry point shows the strongest CLV; edge decays as a market matures toward resolution and the price converges on truth. [2026-07-27 · preliminary · entry-point split, 25% vs 50% vs 75%]
  • The small-market hypothesis is NOT settled and we are NOT baking it in. The size effect FLIPS between samples: on the 100-market set (Haiku/DeepSeek/gpt-oss) CLV rose with size (large best, e.g. DeepSeek small -2.5 / large +10.3); on the 25-market subset (Sonnet/Opus) it reversed — small best AND the only positive Brier skill (Sonnet small CLV +9.1 / skill +0.53). Those "3 models agreeing" were 3 models on ONE sample, not 3 samples. Most likely the size signal is a shadow of the PRICE-LEVEL effect (below), which IS robust — size just proxies how balanced-vs-longshot each sample happens to be. Verdict: size is not a reliable edge axis yet; certify price level instead. [2026-07-29 · UNSETTLED — size effect reverses by sample · 100-set vs 25-subset disagree on size direction]
  • The price-level GRADIENT is the most robust thing we have found. Certification run (125 markets / 375 wagers, Sonnet): extreme longshots (Brier skill -2.65 / CLV -2.2 / 24% hit) < leaning (-0.48 / +4.8 / 52%) < balanced ~50/50 (-0.04 / +8.0 CLV / +22% return / 63% hit) — monotonic on EVERY metric, consistent with the earlier 5-model reads. CORRECTION: the balanced bucket Brier skill is ~0 — we MATCH the market accuracy, do NOT beat it. The +0.78 from the 12-wager bake-off was small-sample luck and did NOT survive n=90. What stays positive on balanced: CLV (+8) and realized return (+22% / 63% hit), the more bet-relevant signals. Within balanced, edge rises with size (large Brier +0.18 / CLV +14.6 > small -0.28 / +2.5). Verdict: the corner is real and looks PROFITABLE on CLV/return, but "beats the market on accuracy" is NOT certified. Gate: convert CLV -> net-of-cost dollars. [2026-07-30 · CERTIFIED n=375 — gradient robust; balanced Brier ~break-even · run_20260730 — 125-market Sonnet certification]
  • The resolution-parse + rules-red-team guard (3.6 Prong A) does NOT beat blind confidence recalibration. On the benchmark, the aggressive guard cut loss-tier Brier -26% but regressed the clear-correct controls +35% (net -17%) — and a FREE shrink of every prediction ~0.3 toward 0.5 matches it (-16 to -20%). So our overconfidence is SYSTEMIC, not specific misreads: a rules-red-team, reading only the rules, cannot tell confident-RIGHT from confident-WRONG, because most blowups are FACT-verification failures (did the event happen in the window?), not rules-interpretation errors. Two takeaways: (1) redirect 3.6 to Prong B — primary-source tools that CONFIRM the fact — which is exactly what the lost-wager wishlist asked for; (2) a cheap calibration shrink IS a real Brier win, but it dampens betting edge too, so it is not free. [2026-07-30 · A/B on benchmark v1 (64 wagers) · resolution-guard A/B vs blind shrinkage, benchmark v1]
  • Prong B WORKS where Prong A failed. The v4 primary-source-verification PROMPT (confirm the exact resolution fact against an authoritative source before high confidence, else widen) cut blowup-tier Brier -20% — Trump-UFC-329 0.97→0.18, trade-deal 0.98→0.40 — with ZERO regression on the clear-correct control tier. That is SELECTIVE, unlike Prong A / blind recalibration which shrink everything. Crucially it needed NO new tools: the existing date-gated search already indexes primary sources, so the bottleneck was agent BEHAVIOR (it did not think to verify the exact fact), which the prompt fixes — a cheap win. Residual: pure non-events (New Glenn never launched → nothing to find in news) still miss, so real primary-source feeds (launch logs, dockets) remain the lever for the hardest cases. Recommend v4 as the ACTIVE prompt (pending a quick check it does not dampen the winning tiers). [2026-07-30 · A/B on benchmark blowup tier (n=16 + 16 control) · v4 primary-source A/B, benchmark blowup + control tiers]
  • Our Brier skill is ALREADY scored the fair way — against the market ENTRY-mid price (score.py: market_brier = brier(entry_mid, outcome)), not the closing line. So the negative aggregate (Sonnet -0.63) is a real entry-vs-entry result, not a stacked benchmark. What drags it down is longshot domination: ~half the wagers are extreme-priced markets the market already nails (Brier skill -8 to -9), and the quadratic loss buries the mean. Strip to balanced (~50/50) markets and Sonnet is +0.78. Scoring against the CLOSING line WOULD be hindsight (70% of markets close within 5pp of 0/1; Sonnet -0.63 vs entry would become -2.30 vs close) — which is exactly why we do not use it. [2026-07-29 · measured / corrected · score.py market_brier basis; price-level split]
Pillar 1: Improved Predictions
  • INDEPENDENT MODEL AGREEMENT TEST: when two independently-run models agree on a market's side, that agreement is a market-selection / sizing signal worth more than either model's solo probability. Evidence (8/10 bake-off, run_20260810-235115: 199 paired markets @25% entry, Sonnet 4.5 vs Sonnet 5, same news pack + v4-primary prompt): the models individually TIED (paired dBrier -0.021 +/-0.042, dCLV +1.3pp +/-5.6) — but the AGREEMENT set (n=148, 74%) ran +8.5pp CLV / -4.9% even-return while the DISAGREEMENT set (n=51) ran -14.1%. Neither model was reliably right on flips (Sonnet 5 29/51). Hypothesis: bet only (or size up) on two-model agreement; treat disagreement as a danger flag. Validate free on existing prediction sets (Allocation-Lab-style re-score), then as a prespecified rule on the Pool A/B extended-window runs. If it holds, production adds a second cheap voter per market. Promising — do not lose track. (Roadmap Epic 3.8.)
— no conclusions logged yet —
Efficiency & cost
  • Two-tier funnel: a CHEAP model (gpt-oss / DeepSeek) SCREENS the wide market universe for candidate edge; the EXPENSIVE calibration leader (Sonnet) PRICES only the finalists we might actually bet. The 100-market bake-off showed cheap models degrade on the broad, representative set (Round-1 near-parity was on an easy Politics-heavy 27), so all-cheap is not enough — but all-Sonnet is wasteful. To validate: does the cheap screen recover most of the edge all-Sonnet would find, at a fraction of the cost?
  • Final bake-off (paired on the same 25 markets, all fairly with thinking): Sonnet 5 is the pricing champion — CLV +5.5 / Brier-skill -0.63, far ahead of Opus 4.8 (+1.5/-1.27), Haiku (-0.8/-1.19), DeepSeek (-1.7/-1.62), gpt-oss (-2.4/-1.95). Notable: Sonnet beats even Opus once both think — the pricier model is NOT better at this forecasting task. Fixing the thinking bug lifted Opus from +/-0 to +1.5 CLV (the bug was real) but did not close the gap. Haiku is dominated — pricier ($18.46) AND worse than the open-weight models on the representative set. Proposed (NOT yet committed): two-tier funnel = DeepSeek or gpt-oss SCREEN (cheap breadth), Sonnet PRICES the finalists — to be validated before we adopt it. [2026-07-29 · complete 5-way, n=1 per model · 27-market (7/28) + 100-market 5-way (7/29) bake-offs]
  • Prompt caching cuts ~35%+ per evaluation — the research transcript is ~94% input tokens, so caching the growing prefix is a big, free lever (never changes output). [2026-07-28 · validated · per-eval cache-read measurement]
  • Parallelizing the independent (market × entry) evaluations gives ~5-10x wall-clock — the work is I/O-bound on the research API, so concurrency is nearly free up to rate limits. [2026-07-28 · validated · thread-pool harness]
  • Open-weight bake-off: gpt-oss-120b runs the full loop at ~1/15 of Haiku's cost ($0.63 vs $9.28/run) with near-Haiku quality — slightly lower overall CLV/Brier but the BEST CLV in the Politics @ 25% corner (+15.2pp). Qwen3-32B is cheapest ($0.39) but UNDER-researches (one search per market) and calibrates much worse (Brier -1.71). gpt-oss is the cheap workhorse candidate; Qwen is not, as-is. [2026-07-29 · early / n=1 · 27-market A/B vs Haiku/Sonnet, 7/29]
Capital allocation
— no open hypotheses here yet —
  • Time-adjusted (per-day) fractional Kelly is our best sizing rule on risk-adjusted return — it beats naive EV×odds on Sharpe (+0.19 vs -0.92) by pushing capital toward faster-resolving edges, since a fast market recycles capital many times a year. [2026-07-28 · preliminary · Allocation Lab sizing sweep]
  • Longshots drag the aggregate — extreme-priced wagers (the "aliens" problem) pull down the whole book. The fix belongs in SIZING (Kelly naturally starves them), not a hardcoded floor rule. [2026-07-28 · directional · aggregate vs Kelly-weighted return]
  • The blended-book return is NOT the objective — the CORNER is. We optimize for the segment where we win (e.g. Politics @ 25% + time-adjusted Kelly), not the average over markets we should never have bet. [2026-07-28 · framing · aggregate return is negative while the corner is positive]
Model performance
  • Is our edge coming from the MODEL or the PROMPT? They are confounded today — we changed both over time. To separate them: hold the prompt fixed (v4-primary) and vary the model, then hold the model fixed and vary the prompt, on the SAME benchmark set, and attribute the CLV / Brier lift to each. Hypothesis: the prompt (research discipline + primary-source verification) carries more of the lift than the model tier — which would mean a cheaper model on our prompt keeps most of the edge. This is what makes the two-tier funnel (cheap screen, Sonnet prices) viable.
  • Sonnet 5 is the pricing champion — +5.52 CLV / -0.63 Brier on the paired-25, a >3σ lead. The only model with positive CLV AND respectable calibration, so it is our simulation model. [2026-07-29 · measured, n=75 paired · bake-off round 2 (paired-25, n=75)]
  • Opus 4.8 trails Sonnet even after fixing a real setup bug (Opus had run thinking-OFF while Sonnet ran thinking-ON — Mike caught it). Per-wager the two are ~tied; Opus just takes 2-3 more overconfident blowups that Brier punishes. A reminder that calibration, not raw capability, is what this task rewards. [2026-07-29 · measured · bake-off round 2]
  • Cheap open-weight models SCREEN but do not PRICE: gpt-oss-120b ($2.18/100 markets) and DeepSeek-V3.1 ($3.65) degrade on the broad representative set, but are fine to narrow the field; Sonnet prices the finalists. Haiku is NOT cheaper than Sonnet here (it under-caches / burns tokens). [2026-07-29 · validated · bake-offs R1 (27) + R2 (100)]
  • Run-to-run variance is near-zero (σ ≈ 0.00-0.05 across 30 passes): one pass is representative, so we do not need to ensemble for stability, and our errors are SYSTEMATIC — the model cannot self-flag them. The lever is selection / sizing, not re-rolling the oracle. [2026-07-31 · strong / σ≈0 · 10-market × 3-pass flip test]
Business & scale
  • Our real constraint is edge CAPACITY, not market volume: the deployable capital at our edge (the size of the mispricing at the order book), summed across winning segments and venues. Likely far exceeds near-term capital — making this a capital-raising story — but that has to be measured, not assumed.
  • Total addressable market for a trading strategy is edge CAPACITY, not total volume: per-market capacity ≈ the size of the mispricing at the order book, summed over winning segments × turnover, modeled against a return-vs-size decay curve. [2026-07-28 · framing · TAM methodology]
  • Edge CAPACITY is Polymarket-dominated. Live balanced order books hold a median ~$42.6k within a 10c edge band per market on Polymarket (30 real books, p25-p75 $11k-$82k); Kalshi books are far thinner. So the deployable-at-edge ceiling — ballpark ~$0.5-1M now, ~$3-6M/yr throughput at ~6x turnover — lives almost entirely on Polymarket. Directly supports the $1-2M AUM ambition, but ONLY by leaning on Polymarket; Kalshi alone caps at low tens of $k. Reinforces why the two-venue setup (Epic 8) matters for scale. [2026-07-30 · ballpark — real books, rough assumptions · live Kalshi+Polymarket order-book depth, balanced markets]
Edge-hunting framework — the 13 ways we slice for edge
  1. Resolution mechanism — Deterministic (a mechanical rule) vs judgmental ("will it be reported that…") — they behave very differently; our Calculation-vs-Logical-Inference driver is the seed.
  2. Time-to-resolution at entry — Far out, base rates dominate; near close, a knowable catalyst decides it. Also drives capital turnover.
  3. Price level at entry — Balanced (~50/50, genuine uncertainty) vs extreme (longshots carry structural favorite-longshot bias — a classic edge zone).
  4. Liquidity / size tier — Small vs large. Fewer sharps in small markets (the hunch), but thinner books cap deployable capital.
  5. News intensity (coverage volume) — How MUCH news exists, from GDELT coverage counts. Edge likely peaks at MEDIUM volume — enough to research, not so saturated the market is already efficient. Distinct from #11, which is about source TYPE in low-news domains. (Not yet wired.)
  6. Scheduled catalyst — A known event/date between entry and close that our research can anticipate — vs pure drift.
  7. Base-rate anchorability — Can the outcome be tied to a strong reference class the market is ignoring? Computable base rate + inattention = edge.
  8. Sentiment- vs fundamentals-driven — Partisan / hype-cycle / novelty topics misprice on narrative and wishful thinking, not information.
  9. Rules clarity / ambiguity — Tightly worded criteria we can reason about crisply — ambiguity creates mispricing but also settlement risk.
  10. Research tractability — Where our GDELT/Wayback pipeline can actually source information (measured by the research fed/dropped telemetry). No coverage → no edge. (Not yet wired.)
  11. Domain obscurity / primary-source depth — Not news VOLUME (#5) but source TYPE: off-the-radar domains (regulatory, legal, scientific — FDA rulings, court dockets, trial registries, SEC filings) that carry findable PRIMARY sources the crowd never reads. The exception to "thin news → no edge": low news BUT deep primary sources + an inattentive crowd is our best shot at out-researching the market. (Not yet wired — needs primary-source retrieval beyond GDELT news.)
  12. Counterparty composition (retail vs institutional) — Who is on the other side. Edge should concentrate where retail dominates AND the sharps (Susquehanna, Jump) have not bothered. Measurable directly on Polymarket (on-chain, wallet-level); on Kalshi only proxied via order-book microstructure. (Not yet wired — needs order-book / on-chain data; NOT derivable from current backtest records.)
  13. Category / Subcategory — Kalshi's top-level category and our finer subcategory — the coarsest, most familiar slice (the one the crowd itself organizes around) and the baseline every other dimension is read against. Already computed from the ticker taxonomy (INSTRUMENTED).

Backtests & SimulationsThe prompt/guardrail reference and the run data tables

Approaches & Prompts — Best Practices

How we steer research agents — the working principles behind every approach and prompt version below. Rewritten 8/17 from the ground-up prompt-engineering research review (~90 external sources: forecasting-system papers, controlled prompting studies, our own 2,858-wager record); each principle keeps its MATES example where we have one. Source of truth: conclusions.py: PROMPT_PRACTICES; the full research synthesis lives in the shared Google Doc; design-debate handoff at events_desk/PROMPT_DESIGN.md.

1. The prompt is the smallest lever — budget accordingly
Every serious evaluation ranks the levers the same way: model choice > evidence/search quality > sampling-and-aggregation > statistical calibration > supervision > prompt wording. Metaculus's four-quarter benchmark conclusion: the underlying model matters more than all scaffolding combined, and good scaffolding is worth roughly nine months of model progress. Bridgewater's AIA Forecaster matched human superforecasters with a deliberately plain prompt and a smart pipeline. Prompt work still pays — but only after the bigger layers exist.
2. Missing inputs masquerade as reasoning failures
Before rewriting a prompt, check what the model was TOLD. Our Bucket-A window-trap errors looked like overconfidence but were partly a missing input: the agent was never told the market's issuance date or what day "today" was — the TIMELINE block roughly halved those errors, no prompt psychology involved. The external version is starker: with search, AIA scores 0.10 Brier on live markets; without it, 0.36. No wording closes a gap like that. First question of every post-mortem: what didn't the model know?
3. Force the outside view — the one prompt ingredient with measured positive effect
In the only controlled 38-prompt sweep (Schoenegger et al. 2025), base-rate and reference-class references were the single prompt category that helped accuracy. Metaculus tournament winners disproportionately compute explicit base rates (r=+0.38) and look up similar resolved questions. It matches our v5 fix: requiring the reasoning to OPEN with a named reference class and its rate moved the Citigroup/SpaceX wager from 12% to 32% (market 53%, resolved YES) with no change to the rest of the portfolio.
4. Specify WHAT to establish, not HOW to think
Half of the old "procedures beat exhortations" principle stands: "don't be overconfident" does approximately nothing. The other half flipped on frontier models: Anthropic now warns that step-by-step scripts written for older models DEGRADE current-model output ("prefer general instructions over prescriptive steps"). The durable form: name the outputs that must exist (reference class, window check, decisive fact) and which rule has jurisdiction when two collide — and let the model plan its own path between them.
5. The output schema is part of the prompt — and the probability goes last
Field order = generation order = reasoning scaffold: schemas that force the answer before the reasoning measurably degrade it. Fields like resolution_understanding and the 80% interval change what gets thought about, not just what gets recorded — adding a field is a prompt intervention. Two house rules follow: p_yes is always the final field, and final answers arrive via a tool call, never forced-JSON mode (which loses ~50% of long-transcript responses on reasoning models — our own gotcha, since replicated industry-wide).
6. One vivid example beats a paragraph of theory — but examples teach the label distribution too
Models generalize from concrete cases more reliably than from abstract rules, and current models follow examples VERY literally. The new caveat from the few-shot literature: examples also carry an outcome prior. If every embedded example is a "we were too low" story, we are quietly teaching a direction. Keep the example under the general rule, never in place of it — and keep the set outcome-balanced or outcome-free.
7. Rules conflict long before they overflow
2026 frontier models follow ~2,000 simultaneous instructions, so raw rule count is no longer the binding constraint for a prompt our size. Conflicts are: instruction failures are silent omissions, prohibitions get dropped first, and two good rules with an unspecified boundary is the classic bug — our v4 verification discipline was excellent on already-decided questions and CAUSED the Citigroup miss on a not-yet-decided one ("could not verify" collapsed into "low probability"). Every rule needs a jurisdiction.
8. Give every rule its reason
Current models generalize from the WHY better than from the bare rule: "never output 0 or 1, because a single wrong certainty is unrecoverable under Brier scoring and Kelly sizing" outperforms the bare prohibition — and survives model upgrades, because the reason transfers even when the failure mode shifts. It is also honest documentation: a rule whose reason we cannot state is a rule we cannot defend keeping.
9. Personas, pep talks, and Bayesian theatre measure at zero (or worse)
"You are a superforecaster with 20 years of experience": no accuracy effect across 162 personas x 2,410 questions. Emotional stakes, tips, and threats: no aggregate effect across ~20,000 runs — just large random per-question swings, and variance is poison for calibration. "Reason like a Bayesian, compute the likelihood ratio": significantly HURT accuracy in two independent studies. Strip all three; keep the task specification.
10. Diversity beats resampling
Our flip experiment measured run-to-run sigma ~0.00-0.05 — re-running the same model on the same prompt and corpus adds nothing, which is why we skipped ensembling. The winning systems ensemble anyway, because their members differ in something REAL: model family or evidence path. Our own 8/10 agreement test showed the payoff shape (two-model agreement set: CLV +8.5pp). Ensemble members must differ in kind, not in random seed — that reinterpretation is the heart of the Canopy proposal.
11. Critics must name the error — or the forecast capitulates
Models revise CORRECT answers under content-free pressure; in multi-agent debates, correct-to-wrong flips outnumber the reverse. We measured it ourselves before the papers named it: universal red-team review market-anchored the whole book (winners' conviction cut 33%, 90% of moves toward the market). The fix is mechanical, not rhetorical: movement requires an ACCEPTED attack naming a specific rules, window, or fact error — enforced in code (the teeth-revert), not requested in prose.
12. Change one variable, keep a paired baseline — and respect the noise floor
Prompt effects are model-specific and flip sign across models; formatting alone has historically swung benchmarks by double digits. Our discipline stands: frozen versions recorded per run, same markets, same retrieval, one variable at a time. New corollary: with per-wager CLV noise of ~33 points, small portfolio deltas at n~50 are noise regardless of direction — the k53 harness answers "did the target failure fix without breaking the rest," not "which variant is 2% better."

Approaches & Prompts — Approaches

An approach is the structural recipe around the prompts: which agents run, in what order, with what triggers and guardrails. Historically our prompts (v1…v5) versioned independently of the structure (red teams, screening, review rules); from Canopy onward the two are tracked together — prompts version within an approach. Newest first; the prompt texts themselves live in the Prompt Registry below. Source of truth: conclusions.py: APPROACHES.

Canopy — Three-Model Research Ensemble + Reconciler + Market Blend proposed v6 · proposed 2026-08-17
One market, several independent minds, one disciplined number. Three research agents from different model families each price the market alone — same task, no market price shown, live date-gated search. Plain code merges their answers and measures how much they disagree. Only contested markets get a second look, from a reconciler agent that settles the specific point of disagreement. The tradeable number is a blend of our answer and the market's price, sized quarter-Kelly. Named for the forest canopy: many separate crowns forming one continuous layer.
0 · Market selection upstream — hand-picked test set / screening funnel you + code 1 · Research ensemble — independent, price-blind, gated live search same task × three model families — different weights, different blind spots Principal Sonnet 4.5 (sim) Sonnet 5 (production) Member gpt-oss-120b open-weight Member DeepSeek-V3.1 open-weight each runs its own date-gated searches → each returns its own probability 2 · Aggregate geometric mean of odds + disagreement spread code calm → skip 3 3 · Reconciler — only if contested fires when models split, or ≥30pp from market reads all traces · re-researches the crux itself agent 4 · Blend with market price half our answer, half the market’s — in odds space big divergence ⇒ small bet, not no bet code 5 · Size — quarter-Kelly 10%/wager cap · probabilities capped 3–97% code 6 · Human triage ranked alerts · maker-side limit orders Mike
deterministic (code) agent step conditional / upstream
0 · Market selection (upstream of Canopy)
Which markets we even look at is decided before Canopy starts — Mike picks the simulation test set; in production the screening funnel does it. Canopy treats every market it receives identically: we deliberately paused corner-based tiering so the new approach gets a fair read across all market types.
(Upstream: marketplace-catalog filters + human selection. Corner-based budgeting/tiering paused by decision 8/17 — revisit after the 120-market sim.)
1 · Research ensemble — three independent opinions
Three research agents study the market independently and each gives its own probability. They never see the market's price and never see each other's work. All three do the identical job — the value is that different models have different blind spots, like getting three doctors' independent opinions instead of asking one doctor three times. Each runs its own live news searches through the date gate.
(Sim principal: Sonnet 4.5 — training cutoff precedes the 12-month test window; production principal: Sonnet 5. Members: gpt-oss-120b + DeepSeek-V3.1. Prompt v6; price-blind; gated live search — server-side date predicate in BigQuery GKG, bodies via Wayback + sanitizer; probabilities capped to [0.03, 0.97] and emitted LAST via tool call.)
2 · Aggregate — merge the three answers (pure code)
Code merges the three probabilities into one. Instead of averaging percentages, we convert each answer into betting odds, average those, and convert back — because "95%" is a much stronger claim than "60%", the way 19-to-1 says more than 3-to-2. The code also records how far apart the three answers were: that spread is our contested-market alarm bell.
(Geometric mean of odds; spread = max-min across members. Deterministic Python — nothing to prompt, nothing to hallucinate.)
3 · Reconciler — a second look, only when contested
If the three agents disagree a lot, or their combined answer sits very far from the market price, a fourth agent takes a second look. It does not attack anyone: it reads all three write-ups, spots exactly WHERE they split (usually one checkable fact or a base rate), runs its own targeted searches to settle that point, and issues the final probability. Calm markets skip this step entirely.
(Trigger: ensemble spread above threshold OR |aggregate - market| >= ~30pp. Model: Sonnet 4.5 in sims; Fable 5 is the production candidate. AIA-style reconciliation — no rebuttal channel, so no capitulation risk. The red-team-with-teeth reviewer is PARKED as the alternative arm, to be tested head-to-head at a later date.)
4 · Blend with the market price (pure code)
We do not trade our number OR the market's number — we trade a mix of the two, half-and-half to start. The market price is itself a very good forecaster, so ignoring it throws away free information. This also softens our worst historical failure: when we said 6% and the market said 76%, the old rule threw the trade away entirely above 40pp of disagreement; the blend instead takes a SMALL position — extreme disagreement now means a small bet, not a huge bet and not zero.
(logit(p*) = w*logit(p_model) + (1-w)*logit(p_market); w = 0.5 headline for the 120-market sim, raw model p recorded so the full w-sweep {0.3, 0.5, 0.7, 1.0} scores free post-hoc. Replaces the >=40pp exclusion cliff. The separate recalibration map is parked to roadmap 5.6.)
5 · Size the bet (pure code)
The blended probability sets the bet size using the Kelly formula at quarter strength — deliberately cautious, because bet-sizing math punishes overconfident inputs much harder than timid ones.
(Quarter-Kelly on blended p*, 10% per-wager cap; half-Kelly reported as a free secondary column from the same run. Correlated-event exposure is handled upstream in market selection per the 8/17 decision.)
6 · Human triage
Nothing trades itself. Ranked alerts land on the desk; Mike reviews and places maker-side limit orders.
(Unchanged desk policy: agents do analysis only; a human moves money.)
Test configuration — Agreed 8/17 for the 120-market / 12-month simulation: all markets treated equally (no corner tiering) · 3-member ensemble (Sonnet 4.5 principal + gpt-oss-120b + DeepSeek-V3.1), ALL on gated live search · geometric-mean-of-odds aggregation · reconciler-only review arm (red team parked this round) · blend w=0.5 headline with free post-hoc w-sweep · quarter-Kelly headline sizing, half-Kelly reported alongside · recalibration NOT included (roadmap 5.6). Prompt: v6, to be drafted and A/B'd against v5 on the k53 set before the big run.
Heartwood — Single Principal + Selective Red Team active v3–v5 era · 2026-07 → present
The current shape: one frontier research agent prices each market alone (price-blind, news-aware, versioned prompt v3→v5), with red-team/blue-team adversarial review reserved for quarantined markets — the high-divergence cases where history says we blow up. Movement after review is mechanically gated ("teeth"): the estimate may only move materially if the researcher ACCEPTS an attack naming a specific rules, window, or fact error.
1 · Single research pass
One Sonnet agent researches the market behind the leakage firewall and prices it, blind to the market price.
(v4/v5 prompt lineage; agent-driven pull over date-gated news; verdict via submit_estimate tool.)
2 · Selective adversarial review
Only markets where we hugely disagree with the market get attacked by a red-team agent; the researcher must answer every attack with evidence and may only move its number if a specific error is proven against it.
(run.py --redteam --red-teeth; TEETH-REVERT enforced in code, not prose — validated 8/11 on the 8/4 244-wager run.)
3 · Alert queue
Verdicts become ranked alerts for human triage.
(Conviction-ranked; live repricing on page load.)
Sapling — Blind Single Pass retired v1–v2 era · retired 2026-07-31
The first backtest era: a single agent priced markets from base rates and general knowledge alone, with no news access. Retired permanently — news-aware research beat it decisively ("we will always run News Aware"), and it survives only as a baseline in old run records.

Approaches & Prompts — Prompt Registry & Guardrails

The implementation layer beneath the Approaches above, in two parts: the versioned research-agent prompts, organized by agent role (Principal / Associates–Screening / Red Team), and the hardcoded guardrails (every magic number that shapes a decision) — surfaced here so we never lose track of them. Sources of truth: events_desk/ra_prompts.py and events_desk/backtest/research_redteam.py; guardrails in conclusions.py. {brief} = the market's rules; {access} = the info channel injected per context (live web / date-gated news / blind).

Principal Research Agents

The frontier pricing model (Sonnet 5, validated champion). Versioned prompt set below — the active version is what run.py uses by default.

PromptDateDescriptionChange summaryFull text
v1 · Core evaluator
v1-core
2026-07-25 Understand-before-pricing: restate resolution, classify the trigger, catch the near-certain trap, name the YES/NO meaning, and give an honest 80% interval. Baseline. Understand-before-pricing skeleton.
v2 · Deep research + calibration
v2-deep
2026-07-27 v1 plus explicit deep-research discipline: reason about mechanical/process timing (fuel-loading, regulatory steps), fight overconfidence, and anchor on base rates before the narrative. Added deep-research discipline: process/mechanical timing, anti-overconfidence, base-rates-first.
v3 · + post-mortem lessons Active
v3-postmortem
2026-07-28 v2 plus pitfalls learned from the agentic feedback loop — most notably: pin down the EXACT referenced event (e.g. UFC 329) and cross-check its date against the market close, instead of substituting a similarly-themed event. Added the exact-referent cross-check (UFC 329 lesson) from the feedback loop.
v4 · + primary-source fact verification
v4-primary
2026-07-30 v3 plus the primary-source discipline from the 3.6 A/B: the biggest misses are FACT-verification failures, so confirm the exact resolution fact against the most authoritative source before committing to high confidence — else widen. Added exact-fact verification vs authoritative/primary sources before high confidence (3.6 Prong B).
v5 · + lookup/forecast base-rate discipline
v5-baserate
2026-08-12 v4 plus a decision procedure for questions whose resolving fact does not exist yet: name the reference class, start from its base rate, let evidence adjust the anchor rather than replace it. v4's source-verification discipline is explicitly scoped to already-decided (LOOKUP) questions. Added the LOOKUP-vs-FORECAST classification: undecided outcomes anchor on a named reference-class base rate; absence-of-reporting is weak evidence (KXSPACEXBANKPUBLIC post-mortem — all four models absence-anchored a not-yet-decided syndicate).
v6 · Canopy rebuild
v6-canopy
2026-08-17 The Canopy-approach prompt: a single coherent rebuild consolidating the v1-v5 lessons (rules-first reading, window discipline, exact referent, primary-source verification, base-rate anchoring) into motivated rules with explicit jurisdictions, plus a scoring-aware framing and a reasoning-first output contract (RA_VERDICT_SCHEMA_V6, probability last). Ground-up rebuild for the Canopy approach (not composed from the v1-v5 blocks). One coherent document; every rule carries its reason; DETERMINED/UNDECIDED jurisdictions replace LOOKUP/FORECAST; the agent is told how it is scored (Brier + CLV + Kelly sizing, and that timid hedging costs like overconfidence); paired v6 output schema puts all reasoning fields BEFORE p_yes and adds question_class / reference_class / base_rate / what_would_change_this. Gate: A/B vs v5 on the k53 set before any big run.
Associate Research Agents (Screening)

Budget open-weight models (gpt-oss-120b, DeepSeek-V3.1 via DeepInfra) that pre-screen the universe: cross-model consensus mean flags candidates; a >15pp split or side disagreement quarantines a market as blowup-risk. They run the IDENTICAL versioned prompt as the Principal (parity by design — the 8/10 validation held the prompt constant so the model was the only variable; a stripped cheaper screening prompt is a future experiment).

PromptDateDescriptionChange summaryFull text
Same as Principal · active version Active
assoc-active
2026-08-10 Runs the active Principal prompt version verbatim (currently v3-postmortem). Screening validation 8/10: Spearman(cheap_P, principal_P) = +0.457 solo, +0.563 for 4-run consensus.
Red Team Agents

Adversarial review (run.py --redteam): an independent, tool-free RED attacker sees the point-in-time brief, blue's verdict, and the entry mid — then blue must rebut inside its live research conversation before finalizing. Validated 8/11 on the 8/4 244: reliably catches Bucket-A/B flaws, but universal review market-anchors — production use is SELECTIVE (quarantined markets only).

PromptDateDescriptionChange summaryFull text
Red Team · five-lens attack Active
red-attack
2026-08-11 Resolution law / window check / knowability / evidence audit / steelman-the-market. Returns typed attacks (fatal/serious/minor), never its own probability. v1. Known failure mode: red can INJECT a wrong resolution reading and blue capitulates (KXLEAVEPOWELLGOV 0.15→0.96 on a NO) — guardrail-teeth experiment queued.
Blue rebuttal contract Active
red-rebuttal
2026-08-11 Blue dispositions every fatal/serious attack (refuted / accepted / partial) and re-submits via the submit_estimate tool. Explicitly forbidden from courtesy-shrinking; an unrefuted fatal attack must move the estimate toward the market. v1. Tool-path submission (forced-JSON over a long tool transcript loses ~50% of rebuttals on reasoning models).
Reconciler Agent (Canopy)

Canopy step 3 (reconcile.py): when the ensemble's spread or its divergence from the market crosses the gate — trigger arithmetic runs in code, never in a prompt — a fourth agent reads the anonymized member write-ups, names the crux they split on (or the assumption they share), settles it with its own targeted date-gated searches, and issues the final probability. AIA-style: no rebuttal channel back to the members (no capitulation risk), and fully price-blind. Fail-open: a reconciler error keeps the code aggregate.

PromptDateDescriptionChange summaryFull text
Reconciler · crux-finder Active
reconciler-crux
2026-08-18 Find the ONE checkable fact / rules reading / base rate the panel actually turns on, settle it with evidence, judge cases not confidence, then decide — not average. Output pinned to the v1-v5 submit schema regardless of the ACTIVE version. v1 — built for the Canopy 120-market sim (reconciler-only review arm; red team parked per 8/17 decision).
Guardrails — hardcoded rules & thresholds

Baked-in numbers that shape decisions. Review occasionally — moving any of these moves the strategy. active = live in code · designed = built, not yet wired in.

GuardrailCategoryRuleWhyStatus
Fractional Kelly (¼)
score.py: DEFAULT_KELLY_FRACTION
Sizing stake = 0.25 × full-Kelly Full Kelly is growth-optimal but wildly volatile; the ¼ haircut trades a little growth for far smaller drawdowns. active
Per-bet size cap
score.py: SIZE_CAP = 0.10
Sizing never stake > 10% of bankroll on one wager Hard ceiling against a single position sinking the book, regardless of edge. active
Wide-interval haircut
score.py: CI_PENALTY=0.5, CI_WMAX=0.5
Sizing shrink size up to 50% as the 80% interval widens toward 0.5 Bet less when our own uncertainty (interval width) is high. active
Brier benchmark = ENTRY mid
score.py: score_bet
Scoring market_brier uses the entry-mid price, not the closing line The fair forecasting test is us-vs-market at ENTRY; the closing line is hindsight (it has seen the info). active
Price-level buckets
dimensions.py: price_bucket
Bucketing edge-dist <0.15 = extreme · <0.35 = leaning · else balanced Defines the "balanced corner" — our certified edge zone. Moving these lines moves what counts as the corner. active
Horizon buckets
dimensions.py: horizon_bucket
Bucketing ≤14d short · ≤60d medium · else long Time-to-resolution slicing. active
Size tiers
analysis.py: _size_bucket
Bucketing <$10k small · <$100k medium · else large (volume) Liquidity/size slicing. active
Entry points
entry.py: DEFAULT_FRACTIONS
Method enter at 25% / 50% / 75% of market life Shows how edge/CLV decay as resolution approaches; 25% has been strongest. active
Allocation filters
analysis.py: ALLOC_STRATEGIES
Allocation conviction ≥0.10 and CI ≤0.35 for the filtered strategies Only stake when our edge is meaningful and our interval is tight. active
Calculation-market exclusion
select_targets.py
Policy only logical-inference markets are simulated; calculation markets excluded Calculation markets (a number crossing a threshold) degrade under our research. active
Kalshi fee model
Gate-3 cost model (to be codified)
Cost fee = 0.07 × P × (1−P) per contract Net-of-cost returns. Standard taker rate — needs confirming vs Kalshi's current schedule. active (confirm rate)
Resolution confidence cap
resolution.py: resolution_guard
Research (3.6) graduated pull toward 0.5 by red-team severity A/B v1: does NOT beat a free confidence shrink — overconfidence is systemic, blowups are fact-errors. Not shipped; Prong B (primary-source tools) is the lever. tested — not shipped

Backtests

One row per simulation run, newest first. The metric block is the standard lens — ¼-Kelly sizing at the 25% entry point over the run's valid wagers — plus All·even CLV (even-weighted, every entry point) as the model-comparison reference. Strategy × entry sweeps live in the Allocation Approach table below. ARR is an inflated perfect-redeployment upper bound. Drag a column's right edge to resize it.

Show table

Markets Simulated

Catalog of markets in the cohort — category, our subcategory, outcome-driver, resolution, closing line.

Show table

Wagers

One row per valid run × market × entry point (allocation-independent — no sizing). 7,935 wagers on file. The page never renders the rows (thousands, unbounded — that crashes the browser): pick a slice, get server-side summary stats, and export the line-level rows (with our reasoning + resolution notes) to Google Sheets. Same pattern as the Marketplace Catalog below.

none = all
canonical = headline set
|p − mkt mid|

Allocation Approach

Sizing is a LAYER re-scored over existing predictions (no new API). Objective: annualized return on deployed capital. Time-adjusted rules (·per day) push more capital toward faster-resolving edges — a 7-day and a 365-day wager with equal raw EV are NOT equal (the fast one recycles capital ~52×). Read Ret on capital + CLV + Brier skill together; Ann. ret is an inflated upper bound (assumes every dollar recycles at that rate all year). Funded = how many wagers the rule actually stakes (positive-EV / positive-Kelly only).

Allocation approach — which allocation wins, per run

Segment Performance

Even-weighted performance per segment — the corner where we beat the market (Brier skill > 0). Even weighting isolates prediction quality from the sizing rule.

Segment performance — even-weighted

Marketplace CatalogThe cross-venue market universe (≥ $1k) — filter to a slice, get summary stats; line-level exports to Google Sheets (the page never bulk-loads the rows)

Query the universe

20,955 markets · $11,987,645,859 volume catalogued. Enriched with the Structural Edge Score (higher = more structurally in our corner) — filter by Min SES, see avg SES per segment, and the Google-Sheet export is SES-ranked with the full dimensional profile.

daily open-market snapshots
none selected = all

Glossary

Every metric & term we use, defined once — so we stop forgetting them. Edit in events_desk/conclusions.py.

Show glossary6 groups
Metrics
CLV (Closing Line Value)
How many points the market price moved TOWARD our estimate between our entry and the market close. Positive = the market came to us. Our most sample-efficient edge signal — it reads whether we were early to the right price, even on markets we never resolve.
Brier score
Squared error of a probability forecast vs the actual 0/1 outcome: (p − outcome)². Lower is better. Punishes confident WRONG calls quadratically — one blowup hurts a lot.
Brier Skill Score (BSS)
1 − (our Brier ÷ the market's Brier), scored against the market's ENTRY-mid price. Positive = we forecast more accurately than the market did; negative = worse. The honest "do we beat the market" test (we are ~break-even on our best corner).
Edge Influence (η², eta-squared)
The share of the variance in our CLV explained by which SEGMENT of a dimension a market falls in. 0 = the dimension tells us nothing about our edge; 1 = it perfectly separates edge. How much a dimension DISCRIMINATES where we win.
Structural Edge Score (SES)
A per-market score of how well a market's STRUCTURE (its dimensional profile) fits where we have edge — independent of the current price. The "edge prior": our expectation of edge BEFORE we look at the line. Per-dimension weights come from Edge Influence.
Conviction
Our composite ranking of a wager idea: projected annualized return × confidence × interval-tightness, on a saturating 0–95 scale. Drives how much capital Kelly allocates.
Method
Quarter-Kelly sizing (annualized / duration-weighted)
How much of the bankroll we stake. Full Kelly for a binary contract at price c (= the market’s implied probability) given OUR probability p is f* = (p − c) / (1 − c); we stake a conservative QUARTER of it (0.25·f*), shrink it for a wide confidence interval, and cap it at 10% of bankroll per bet. "Annualized / duration-weighted" = we size on edge PER DAY (×365/days-to-resolution), so a 7-day contract (capital recycles ~52×/yr) outranks a 365-day contract of equal raw edge — the time value of money + expiry. Example ($10k bankroll): market 50¢, our p = 0.70 → f* = (.70−.50)/.50 = 0.40 → ¼-Kelly stakes 10% = $1,000; if it resolves YES, +$1,000. Had our p been a timid 0.55, f* = 0.10 → we would stake only $250 and win just $250 on the SAME correct call — miscalibrated conviction under-bets our winners (and over-bets our losers).
Metrics
Return on capital / Weighted return
Realized $ return per $ actually deployed, size-weighted across wagers. The intuitive "did we make money" number.
ARR (annualized return rate)
Return scaled to a yearly rate. The raw ARR is an upper bound (assumes every $ recycles at the per-wager rate all year); realistic ARR discounts it by achievable turnover + capacity.
Sharpe
Return ÷ its volatility — return per unit of risk. Higher = smoother, more reliable edge.
Coverage
The % of wagers/markets that carry a real tag for a dimension. A "trust it yet?" flag — low coverage means that slice's numbers are thin.
Regression coefficient
In a model predicting our edge from the dimensions, the number on each dimension = its effect on edge holding the others fixed. The MULTIVARIATE cousin of Edge Influence — it de-correlates overlapping dimensions. How SES weights get learned as data grows.
Method
Entry point / fraction
How far through a market's open→close life we (re)enter it: 25% / 50% / 75%. This is entry TIMING, not position size. Earlier entry leaves more room for the line to move.
News-aware vs blind
News-aware = the research agent runs its own date-bounded searches (leak-safe) before pricing — our permanent default. Blind = prices from rules + priors only (retired).
Pull research
The agent-driven, date-bounded research loop: it searches a point-in-time news corpus (GKG), reads archived articles behind a leakage firewall, then prices. Best-calibrated method.
Fractional / time-adjusted Kelly
Kelly sizes a bet by its edge; we use ¼-Kelly (a volatility haircut), time-adjusted (per-day) so capital flows to faster-resolving edges. Our sizing rule.
Market characteristics
Price level (balanced / leaning / longshot)
Distance of the entry price to the nearest 0/1 rail. Balanced (~50/50) is our sweet spot; longshots (<15¢ / >85¢) are toxic. Our most robust edge axis.
Driver / Resolution mechanism
Calculation (resolves by a number crossing a threshold) vs logical-inference (resolves by whether an event/condition occurs). We bet logical-inference; calculation degrades under research.
Domain / source depth
Whether the outcome hinges on findable PRIMARY sources the crowd rarely reads (regulatory / FDA / SEC / court / legislative = primary-rich) vs general news only. Primary-rich is an edge zone.
Counterparty composition
Who is on the other side — retail vs institutional/sharp. Edge should concentrate where retail dominates and sharps have not bothered. Measured on Polymarket on-chain, proxied to Kalshi via twins.
Twin market
A market on the OTHER venue asking essentially the same question (a Kalshi↔Polymarket pair). Lets us port Polymarket's on-chain counterparty read onto Kalshi.
Dimension / Segment / Slice
A dimension is one of the 13 ways we characterize a market (price level, category, …). A segment is a bucket within it (e.g. Balanced). A slice = a dimension split into its segments.
Business
Edge capacity
Deployable-at-edge capital: roughly the size of the mispricing at the order book, summed over winning segments × turnover. Our real constraint on scale (Polymarket-dominated).
Marketplace Catalog
The cross-venue market UNIVERSE (Kalshi + Polymarket) — every market tagged on our dimensions with $ volume and venue — the SUPPLY side that pairs with our per-segment performance to size the total edge opportunity.

Honest caveats. Results are early (n up to 40) and the testing environment is still being hardened toward real, leak-safe data. Brier-skill is negative on the broad cohort so far — we don't yet out-calibrate the market across all markets. Treat that as a target, not a verdict: the mission is to (1) find the segment — category, market size, structure — where we DO beat Kalshi, and (2) fix model overconfidence at the source (standing research-agent instructions), not just build allocation levers around it. Also: positive CLV has not yet meant positive realized return (favorite/longshot payout math); backtests overstate live (no market impact, perfect fills, survivorship).