Can our research desk, using only what was knowable at the time, price already-resolved markets better than Kalshi did? We rewind each market, research it point-in-time behind a leakage firewall, and score the estimate against the real closing line and outcome. Below: how a simulation runs, then three tables to slice (collapsed by default — click to expand).
Seven epics from a $2K discovery project to deployable capital — each gate de-risks the next.
Status: ✓ Done · ◐ In progress · ○ To do · ● On deck today (flashing). Owners: Mike & Silas.
Type a note in any row and hit + — it saves live. Structure edits live in events_desk/conclusions.py.
| ID | Action item | Status | Notes |
|---|---|---|---|
| 0.1 | Scale-up proposal deck (a few slides): walk-forward evidence, honest caveats, risk framing, the ~$50K plan — partly for Mike, partly for Carla | ○ To do |
|
| 0.2 | Productionalization design: weekly Polymarket refresh into the portal — scan scope, cadence, opex. Live census 8/21: ~1,920 engageable markets (>=$20K vol, 7-370d, sports excluded); brute-force Canopy ~ $3.1K initial + ~$2K/mo; funnel shape (cheap screen -> Canopy top ~250) ~ $600/mo | ○ To do |
|
| 0.3 | Walk-forward fix package (pre-registered, lands BEFORE any new sim spend): reconciler prompt v2 (bearish-ratchet bug — 75% of 157 moves down, manufactured the losing tail), theme-exposure key normalization (the Khamenei double-loss), tail governance (extreme-p no-trade band), clean w-sweep rerun | ○ To do |
|
| 0.4 | Monthly-redeployment 12-month walk-forward on ~500 markets (~$1.6-1.8K): the opportunity-density replication — include the original 145 and align the monthly grid to the old schedule for cache reuse; retune the 4-week freshness backstop to the monthly cadence | ○ To do |
|
| 1.7 | Validate the cheap screen recovers most of the edge Sonnet would find | ○ To do |
|
| 5.6 | Recalibration map (PARKED 8/17): fit a one-parameter, shrunk correction curve on ALL paper forecasts (never fills-only — selection bias) and apply it BEFORE Kelly sizing; revisit after the Canopy 120-market sim | ● On deck |
Deferred from the Canopy design to limit moving pieces. Evidence base: fixed/shrunk one-parameter Platt beat fitted isotonic on small n in three independent replications (AIA d=sqrt(3); Metaculus post-hoc Platt dBrier 0.016, p=5e-4). Order of operations matters: calibrate first, THEN fractional Kelly — Kelly error is convex in probability error.
|
| ID | Action item | Status | Notes |
|---|---|---|---|
| 1.6 | Two-tier funnel (PROPOSED, not committed): cheap screen (DeepSeek/gpt-oss) → Sonnet prices | ○ To do |
|
| 2.3 | Certify each candidate edge segment at adequate n (~30-50+ wagers) | ◐ In progress |
|
| 2.5 | Within balanced: edge rises with size (large > small); confirm at higher n (large n=14) | ◐ In progress |
|
| 2.13 | Instrument the not-yet-wired edge slices (news intensity, research tractability, domain obscurity, counterparty) so we can slice by all 12 dimensions | ● On deck |
On deck today. Mostly post-hoc — #5 news volume + #10 tractability come from run telemetry; #11/#12 are market tags attached after the fact.
|
| 2.14 | Scale the certified sweep toward ~1000 markets for statistical power | ○ To do |
|
| 3.4 | Feedback-loop integration: act on the post-mortem lessons the loop surfaces (not just log them) | ○ To do |
|
| 3.6 | Cut the resolution-misread blowups: (a) a dedicated resolution-contract step + a RESOLUTION-RULES-FOCUSED red-team, (b) date-gated primary-source tools (legislation / launch / official-statement / docket trackers) | ◐ In progress |
On deck. Autopsy showed ~26% of balanced wagers are confident MISREADS of the exact trigger/window/referent (New Glenn, RNC, UFC 329), not fact errors. Prong (a): a sub-agent that outputs a structured resolution contract (exact YES/NO trigger, operative deadline from the RULES, named referent, settlement source), then a resolution-rules-focused red agent that hunts for a rules reading that flips our answer (NB: this is the narrow rules red-team, NOT the broader adversarial red-team in 3.2); cap confidence when any ambiguity flag is set. Prong (b): give the agent high-precision authoritative feeds (LegiScan/NCSL, FAA/NASA launch logs, WH briefing archive, FDA/SEC/PACER) it queries date-gated — the opposite of noisy GDELT. Measure blowup-rate drop on the benchmark set (3.1). A/B DONE (benchmark v1, 64 wagers): Prong-A guard (parse + rules-red-team + graduated cap) does NOT beat blind confidence recalibration. Aggressive tuning cut loss-tier Brier -26% but regressed the clear-correct controls +35% (net -17%) — and a FREE shrink of every prediction ~0.3 toward 0.5 MATCHES it (-16 to -20%). So our overconfidence is SYSTEMIC, not specific misreads; a rules-red-team working from the rules alone cannot tell confident-right from confident-wrong, because most blowups are FACT-verification failures (did the event happen in the window?), not rules-interpretation errors. REDIRECT: Prong B (primary-source tools to confirm the exact fact) is the real lever — exactly what the lost-wager wishlist asked for. Guard code stays in resolution.py (tested, NOT shipped as default). Prong B A/B DONE: the v4 primary-source-verification PROMPT cuts blowup-tier Brier -20% (Trump-UFC-329 0.97→0.18, trade-deal 0.98→0.40) with ZERO regression on the clear-correct control tier — SELECTIVE where Prong A was not. Key: it needed NO new tools — the existing date-gated search already indexes primary sources; the bottleneck was agent BEHAVIOR (not verifying the exact fact), which the prompt fixes. Residual: pure non-events (New Glenn never launched → no news to find) still miss, so real primary-source feeds remain the lever for the hardest cases. RECOMMEND v4 as ACTIVE.
|
| 3.7 | FUTURE: deeper resolution/rules understanding — revisit whether better resolution comprehension adds value when it INFORMS the research (which exact fact/window to verify) rather than as a post-hoc cap | ○ To do |
I still suspect there is value in better rules/resolution understanding — let us circle back after Prong B. The 3.6 A/B (benchmark v1) showed the POST-HOC rules-red-team + confidence cap does NOT beat a free confidence shrink — overconfidence is systemic. But that tested only one design. A resolution understanding that FEEDS the research — telling the agent exactly which fact/window to confirm — is a different, untested idea and may well help. Revisit here.
|
| 5.5 | Sum across segments × turnover → deployable-at-edge ceiling | ◐ In progress |
Ballpark: capacity is POLYMARKET-DOMINATED. Live balanced books hold a median ~$42.6k within a 10c band per market on Polymarket (30 real books) vs far thinner on Kalshi. Deployable-at-edge ~$0.5-1M now (mostly PM), ~$3-6M/yr throughput at ~6x turnover; Kalshi alone caps at low tens of $k. So the $1-2M AUM ambition is reachable, but by leaning on Polymarket. Kalshi live-depth pull (series-targeted) is the follow-up.
|
| 5.7 | Return-vs-size decay curve (how edge erodes as we scale) | ○ To do |
|
| 6.1 | Wire live (forward-looking) research — no leak firewall needed forward | ○ To do |
|
| 6.2 | Trade certified segments with a small real stake | ○ To do |
|
| 6.3 | Compare forward realized returns vs backtest expectation | ○ To do |
|
| 6.4 | Only scale capital if forward holds | ○ To do |
|
| 8.7 | Counterparty-composition analysis via Polymarket on-chain wallet data (unlocks #12 / item 2.10); use Kalshi↔Polymarket TWIN markets as a cross-venue signal | ◐ In progress |
First dataset built (results/polymarket_cohort.json, 60 markets): on-chain YES-holder concentration per market. Pattern — composition varies by TYPE: single-game sports are whale-tilted (top-5 hold 92-97%), elections / broad tournament markets are retail-dispersed (~25%). So the retail-vs-sharp signal is real and readable. Next: reach down to moderate-volume markets + match Kalshi twins for the cross-venue proxy.
|
| 8.8 | PolymarketProvider (TRADE): CLOB order signing — LAST (Mike authorizes/executes actual trades) | ○ To do |
|
| 9.6 | Chunk 2.4 — Counterparty proxy (Polymarket on-chain holder concentration; twin- or category-level proxy for Kalshi) | ○ To do |
|
| 9.9 | END STATE — a fully-enriched, SES-scored cross-venue catalog: query e.g. "highest-SES resolved Politics markets" to pick backtest candidates, and size total edge opportunity ($ x SES) per segment. Feeds Pillar 2 (selection) + Pillar 3 (capacity) | ○ To do |
|
| ID | Action item | Status | Notes |
|---|---|---|---|
| 2.6 | Certify: crisp rules · base-rate-anchorable · non-catalyst | ○ To do |
|
| 2.7 | Certify: long-horizon (>2mo) · 25%-of-life entry | ○ To do |
|
| 2.9 | Deep research in obscure markets (Wayback-dated primary sources) — see hypotheses | ○ To do |
|
| 2.10 | Counterparty composition: retail markets institutions ignore (Polymarket on-chain proxy) | ○ To do |
|
| 3.5 | Model experiments on the benchmark set (Sonnet vs Fable vs others; calibration-focused) | ○ To do |
|
| 4.4 | Detailed net-of-cost investigation: confirmed fee schedule, slippage-at-size (walking the book), Kelly-weighted sizing, per-market granularity, and an exit-timing model for the trade-the-line variant | ○ To do |
|
| 5.2 | Trade-the-line vs hold-to-resolution execution — capture CLV by exiting before resolution (needs liquidity to exit; cross-refs Epic 4) | ○ To do |
Our edge reads as timing (CLV) more than accuracy (Brier), so trade-the-line may be our natural mode — gated by whether thin Kalshi books let us exit at the better price. 7/31 -- Later analysis showed that it's actually more profitable for us to hold to resolution in most situations. Could flop in future2026-07-31 15:18
|
| 5.3 | Kalshi order-book depth pull per certified market | ◐ In progress |
|
| 5.4 | Per-market capacity = size of the mispricing at the book | ◐ In progress |
|
| 1.1 | Leak-safe backtest pipeline (date-gated GDELT + Wayback, fail-closed sanitizer) | ✓ Done |
|
| 1.2 | True per-run cost tracking | ✓ Done |
|
| 1.3 | Parallelized harness + prompt caching (~5-10x faster, ~35% cheaper) | ✓ Done |
|
| 1.4 | 5-model bake-off: Haiku / gpt-oss / DeepSeek / Sonnet / Opus | ✓ Done |
|
| 1.5 | Fix fair-comparison bug (adaptive thinking for all Anthropic models) | ✓ Done |
|
| 2.1 | Exclude calculation-driven markets (degrade under research) | ✓ Done |
|
| 2.2 | Build the 12-dimension edge-hunting framework + tag markets | ✓ Done |
|
| 2.4 | Balanced-price ~50/50: CERTIFIED n=90 — gradient robust, CLV +8 / return +22%, but Brier ~break-even (not beating the market on accuracy) | ✓ Done |
The +0.78 Brier from the 12-wager bake-off did NOT survive n=90 (regressed to -0.04). Edge shows as CLV/return, not accuracy.
|
| 2.8 | Small-market hypothesis — data INVERTS it (edge rises with size) | ✓ Done |
|
| 2.11 | Brier-skill-vs-ENTRY-mid confirmed as the metric already in use (score.py); closing line is hindsight | ✓ Done |
|
| 2.12 | GCS-archive price reader built — unlocks the full leak-safe window (~14k archived markets) for a bigger cert | ✓ Done |
CLOSED 8/21: superseded twice over — the Kalshi /historical/ API killed the purge wall, and pools/candles (793 tickers) carries the 12-month backtests without GCS.
|
| 3.1 | Build the experiment benchmark (control) set — a frozen, stratified set of resolved wagers: big wins · long-shot losses · near-miss losses — the fixed control every prediction experiment runs against | ✓ Done |
Best practice (a.k.a. golden set / eval set / holdout / regression suite): FREEZE & version it (v1); stratify across the 3 outcome-quality tiers + our dimensions; keep a small separate DEV set to iterate on so we do not overfit the benchmark; near-miss losses are the most diagnostic slice; lock metrics (CLV/Brier/return) before running any variant. Same wagers every time = a clean A/B on prompt / model / red-team. v1 FROZEN → events_desk/backtest/benchmark_v1.json: 48 wagers, 16 per tier (big-win / longshot-loss / near-miss-loss), each with baseline p/Brier/CLV/return + re-run fields. Locked metrics: CLV, Brier, return, blowup-rate. Lost-wager autopsy: near-miss-loss tier should target our resolution-misread blowups (~26% of balanced wagers) — the diagnostic failure mode.2026-07-30 17:33
|
| 3.2 | Red-team / blue-team adversarial-evaluation lift (A/B vs blue-alone) — run on the benchmark set | ✓ Done |
CLOSED 8/21: ran at scale — universal review market-anchors and taxes returns; verdict = selective review on flagged markets only. Superseded by the Canopy conditional reconciler.
|
| 3.3 | RA-prompt improvement experiments (versioned A/B on the benchmark set) | ✓ Done |
CLOSED 8/21: versioned prompt A/Bs are standing practice — v4→v5 (8/13) and v5→v6-canopy (8/18 k53 gate: CLV +8.1pp, healthier p-distribution) both ran through the machinery.
|
| 3.8 | Independent Model Agreement Test — two-model side-agreement as a selection/sizing screen (see Pillar 1 hypothesis) | ✓ Done |
Source data (8/10/26 bake-off, run_20260810-235115-n200-claude-sonnet-4-5 vs run_20260804-202435-n250-claude-sonnet-5; 199 paired markets, 25% entry, same news pack + v4-primary): models tied overall (paired dBrier -0.021 +/-0.042 2se; dCLV +1.3pp +/-5.6; side-flip rate 26%). AGREEMENT set n=148: CLV +8.5pp, even-return -4.9%. DISAGREEMENT set n=51: even-return -14.1%; flips split 29/22 for Sonnet 5 (n.s.). Segment cuts by Sonnet-5-performance strata are regression-to-the-mean artifacts (selected on S5 extremes) — do NOT cite them. Next: (a) free re-score of existing multi-model runs (27-mkt and 100-mkt cross-model sets) under an agreement-gated allocation rule; (b) prespecify the rule for the Pool A/B extended-window runs; (c) if it survives, add a cheap second voter (Haiku/4.5) to production sweeps. Caveat: models share the news pack + prompt, so agreement is not fully independent evidence — vary the news pack or prompt for one voter to strengthen the test. CLOSED 8/21: validated (8/10 two-model test; 8/18 shakedown diagnostics) and PRODUCTIONALIZED as the Canopy member ensemble + spread trigger; agreement analytics ship in canopy_eval.
|
| 4.1 | Realized return, hold-to-resolution, net of costs, per certified segment (HIGH-LEVEL ESTIMATE) | ✓ Done |
Estimate (hold to settlement): balanced corner +22.1% gross → +18.8% NET of Kalshi fee (0.07·P·(1−P)); leaning +9.4% net; longshots flip net-NEGATIVE (−4.7%). Even-weighted, 1-contract, backtest, n=90. Excludes slippage-at-size (Epic 5) and forward test (Epic 6); fee rate needs confirming. Verdict: the corner clears the cost bar at small size.
|
| 4.2 | Realized return, trade-the-line (exit before resolution) variant (HIGH-LEVEL ESTIMATE) | ✓ Done |
Estimate (exit at the closing line, pay 2 fees): NET much WORSE than hold — balanced +4.4% vs hold +18.8%; all-book −2.8% vs +8.1%. Why: 54% of balanced markets close within 10pp of 0/1, so exit-at-close already reflects the outcome — you pay a 2nd fee for ~the same result. Trade-the-line only pays off by exiting EARLY at peak-favorable, which needs an exit-timing model (→4.4). Takeaway: HOLD is our profitable baseline; trade-the-line is a future enhancement, not a free win.
|
| 4.3 | Decide the actual strategy: accuracy (hold) vs timing (trade the line) | ✓ Done |
CLOSED 8/21: decided — hold-to-resolution is doctrine (8/12); trade-the-line is a future enhancement gated on an exit-timing model.
|
| 4.5 | Realistic ARR: replace the inflated annualized upper bound (naive ×365/days) with a believable figure — certified net return × ACHIEVABLE turnover, bounded by corner supply + capacity, not just time | ✓ Done |
The dashboard ARR is a flagged upper bound (assumes every dollar recycles at the per-wager rate ALL year). Realistic ARR = per-wager net return × REAL annual turnover, where turnover is capped by (1) how many qualifying corner wagers actually EXIST per year — we found ~50 balanced markets in a 6-month window — and (2) capacity (Epic 5). So ARR and capacity are linked; a believable number needs the capacity model. Rough first-cut: ~19% net/wager × ~4-8x realistic recycling ≈ 75-150% — but still capacity-gated on thin books, so treat as directional until Epic 5. CLOSED 8/21: the walk-forward harness replaced naive ARR with calendar-honest annualization — Canopy +53.9% realistic vs the −936% "perfect-redeploy" convention on the same trades (single year, 150 markets).
|
| 5.1 | Time-adjusted ¼-Kelly leads on Sharpe (Allocation Lab sizing sweep) | ✓ Done |
|
| 7.1 | Firm name LOCKED: HEMLOCK HOLDINGS (Western Hemlock, a foundation species) | ✓ Done |
Locked — Hemlock Holdings. Nice and vague; if there is definitive edge I would rather fly under the radar than market broadly. Western Hemlock the TREE (a foundation species: dense canopy cools the forest and keeps streams cold enough for trout), not the poisonous plant. Locked in. NO rebrand yet — MATES stays the system name; this is only the firm-identity decision. Foundation-species metaphor is strong: the quiet keystone that regulates the whole system. (Trademark stays a separate legal conversation.)
|
| 8.1 | Open a Polymarket account (self-custody wallet + USDC) — Mike | ✓ Done |
Polymarket is legal in the US again — no longer restricted. So funding + trading are open. Good — that removes the earlier compliance gate. I still outline steps and build the read-only data + analysis (8.2-8.7); I do not move funds or place trades — those stay yours (8.1 funding, 8.8 trading). Account set up. API key to come tomorrow (not needed for the read-only data + analysis Silas is building).
|
| 8.2 | Define a MarketProvider interface (list_markets · market_meta · price_history · order_book · fee_model · resolution_source · account/positions · place/cancel order); refactor kalshi_client into KalshiProvider behind it | ✓ Done |
18 files touch kalshi_client today (client, archiver, build_manifest, price.py, trade_executor, portfolio, app.py, MCP…). The interface is the seam that lets all of them become provider-agnostic.
|
| 8.3 | PolymarketProvider (READ): Gamma markets API + CLOB price-history + on-chain, normalized into our market record (id, title, category, open/close, outcome, price history, volume) | ✓ Done |
|
| 8.4 | Normalized market id {provider, native_id} (Kalshi KX-ticker vs Polymarket condition/token id); every run record + market catalog + taxonomy tags its provider | ✓ Done |
|
| 8.5 | Per-provider cost model (Kalshi 0.07·P·(1−P); Polymarket ~0 trade fee + gas/spread + UMA nuance) so net-return conclusions are provider-correct | ✓ Done |
|
| 8.6 | Provider-agnostic backtest: manifest + archiver + price_fn work for both; computable dimensions stay shared, category taxonomy mapped per provider | ✓ Done |
CLOSED 8/21: proven end-to-end — the 12-month Polymarket walk-forward (as-of universe, CLOB pricing, bankroll ledger) ran both arms on the provider-agnostic stack.
|
| 9.1 | Chunk 1 — pull the cross-venue universe to Postgres (marketplace_catalog): series-scoped + parallel, 12,170 markets (6,943 open / 5,227 resolved), excl Sports/Crypto, floors open $1k / settled $25k, settled $ = contracts x 0.5 | ✓ Done |
|
| 9.2 | Chunk 3 — query UI: filter (venue / category / status / $ floor / resolution window) -> server-side summary stats; line-level -> Google Sheets export; the page never bulk-loads rows | ✓ Done |
|
| 9.3 | Chunk 2.1 — tag the LLM-judgment dims (catalyst / base-rate / sentiment / rules-clarity / domain-depth) on catalog markets (reuse the Haiku tagger) | ✓ Done |
DONE 2026-08-03 — all 12,170 markets tagged via 6 parallel sub-agents (hashtext-sharded, ~48 concurrent Haiku); size-tier + horizon computed in the same pass. enrich_catalog.py.
|
| 9.4 | Chunk 2.2 — compute the computable dims per market (price-level from a current/entry mid, horizon from close, size-tier from $ volume, driver, category/subcategory) | ✓ Done |
DONE 2026-08-03 — size-tier + horizon + category (all 12,170), price_level (5,149 open via current-mid re-fetch, price_level.py), and driver (all 12,170: 7,832 logical-inference / 4,338 calculation, driver_tag.py). NOTE: driver is a FILTER, not an SES weight — its backtest η² is ~0 because we already exclude calc markets from wagers, so SES actually OVER-scores calc markets (avg 2.78 vs 1.91 logical); the catalog UI Driver filter is the fix.
|
| 9.5 | Chunk 2.3 — Twin Market matching (Kalshi <-> Polymarket); store twin_id per row | ✓ Done |
DONE 2026-08-03 — 419 twin pairs (inverted-index title overlap + numeric/ordinal guard against threshold/opposite mismatches, e.g. "less than 5" vs "above 4"); twin_id on 593 markets. twin_match.py.
|
| 9.7 | Chunk 2.5 — compute the Structural Edge Score per market (apply the Edge Influence weights to each market's dimensional profile); store SES on the row | ✓ Done |
DONE 2026-08-03 — ses.py: SES = Edge-Influence(η²)-weighted avg of a market's segment CLV priors, renormalized per market. All 12,170 scored from 2,614 wagers; NO LLM (rescore() is the cheap reweight+rescore mechanism). Validated: top = Economics/primary-rich (Fed/BoJ rate decisions), bottom = Sci-Tech/news-only longshots (Mars, AGI).
|
| 9.8 | Chunk 2.6 — surface dims + SES in the catalog UI: filter / sort by dimension and SES; summary shows SES x capacity = edge-opportunity per segment | ✓ Done |
DONE 2026-08-03 — Marketplace Catalog UI: Min-SES filter, avg SES per segment (sorted, Economics tops at ~8), SES-ranked Google-Sheet export with the full dimensional profile. Min SES>=5 isolates 1,049 markets.
|
Leakage firewall: every retrieval is gated to sources dated before the entry day — the research agent never sees anything from the future. The agent runs its own point-in-time searches — it chooses what to research against the date-gated corpus, rather than being handed a fixed briefing — which is closer to how production works.
The business decomposes into three linked things we must get right. Pillars 1 & 2 are coupled — the oracle is only as good as the market is researchable — but we work them as separate tracks. Roughly ordered by difficulty & importance: 2 > 1 > 3.
Solidified findings promoted here (first draft — being rewritten). Working / unproven theories and the running historical log live in Open Hypotheses & Conclusions below.
Solidified findings promoted here (first draft — being rewritten). Working / unproven theories and the running historical log live in Open Hypotheses & Conclusions below.
In plain terms: how much this one market characteristic typically matters when we judge whether a market is worth betting on.
Edge Influence (η²) measures how much a single dimension DISCRIMINATES our edge — the share of the variance in our per-wager CLV explained by which SEGMENT of that dimension a market falls into. 0 = the dimension tells us nothing about where we win; 1 = it perfectly separates edge. We compute it as a one-way-ANOVA η² (between-segment variance ÷ total variance), weighted by segment size, over all valid wagers — dropping segments with fewer than 5 wagers so one noisy outlier can't inflate it (the failure mode of a naive best-minus-worst range). Absolute η² is small because per-wager CLV is noisy, so the RANKING across dimensions is the signal, not the raw number.
In plain terms: how attractive this market should be to bet on, based on how similar markets have performed in our historical simulations.
Structural Edge Score (SES) is a per-market score of how well a market's STRUCTURE — its dimensional profile — fits where we have historically had edge, INDEPENDENT of the current price. The price-edge fluctuates trade-to-trade; SES is the stable "edge prior": our expectation of edge before we even look at the line. It is a weighted combination of the market's dimension values, and each dimension's weight IS its Edge Influence — so price-level and category (our strongest discriminators) dominate, while dimensions that don't separate edge (horizon, catalyst, sentiment) barely count.
Learning the weights: v0 (now, thin data) uses transparent hand-set weights anchored on Edge Influence. v1 (as data grows) upgrades to REGRESSION-learned weights — regress realized edge (CLV / return, NOT Brier) on the dimension tags; the fitted coefficients become the weights. Regression is the multivariate upgrade of η²: it de-correlates overlapping dimensions (e.g. price-level ↔ size) so we don't double-count, whereas univariate η² credits both fully. Same underlying quantity, three views: Edge Influence (per-dimension discrimination) = the regression coefficient = the SES weight.
Guardrails: weight by CLV / return, NOT Brier skill (our edge shows as CLV; scoring on Brier would rank our best markets as mediocre). Validate OUT-OF-SAMPLE — a high SES must actually predict edge on held-out markets, not just fit noise. Keep it simple (few features, regularized) until n grows.
Edge Influence = η², the share of our CLV variance a slice explains — how much it DISCRIMINATES our edge (0 = tells us nothing). Weighted by segment size with a ≥5-wager floor, so a lone outlier can't inflate it. Absolute η² is small (per-wager CLV is noisy) — read the RANKING (the bar) and the best/worst segments. v0: even-weighted over all valid wagers across runs (correlated within market + mixes blind/news-aware — directional, not yet bankable). Edge Influence becomes each dimension's weight in the Structural Edge Score (SES).
| # | Dimension | Instr. | Type | Segs | Cov. | Edge Influence (η²) | Best segment | CLV | Worst segment | CLV | n |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Resolution mechanism | Yes | computable | 2 | 35% | 0.006 | calculation | +9.8 | logical-inference | +3.7 | 2017 |
| 2 | Time-to-resolution at entry | Yes | computable | 3 | 100% | 0.002 | Short (<=2wk) | +2.0 | Long (>2mo) | -1.5 | 5839 |
| 3 | Price level at entry | Yes | computable | 3 | 100% | 0.001 | Extreme (longshot/lock) | +2.2 | Leaning | -0.0 | 5839 |
| 4 | Liquidity / size tier | Yes | computable | 3 | 50% | 0.007 | Large (>100k) | +4.6 | Small (<10k) | -2.1 | 2912 |
| 5 | News intensity | No | telemetry | — | — | — | — | — | — | — | — |
| 6 | Scheduled catalyst | Yes | LLM-judgment | 3 | 100% | 0.000 | continuous | +2.0 | no-catalyst | +0.7 | 5839 |
| 7 | Base-rate anchorability | Yes | LLM-judgment | 3 | 100% | 0.001 | weak-base-rate | +1.8 | no-base-rate | -1.4 | 5839 |
| 8 | Sentiment vs fundamentals | Yes | LLM-judgment | 3 | 100% | 0.002 | medium-sentiment | +3.5 | low-sentiment | -0.8 | 5839 |
| 9 | Rules clarity / ambiguity | Yes | LLM-judgment | 3 | 100% | 0.004 | crisp | +2.5 | ambiguous | -1.9 | 5839 |
| 10 | Research tractability | No | telemetry | — | — | — | — | — | — | — | — |
| 11 | Domain / source depth | Yes | LLM-judgment | 3 | 100% | 0.002 | primary-rich | +3.9 | news-only | +0.2 | 5839 |
| 12 | Counterparty composition | No | external | — | — | — | — | — | — | — | — |
| 13 | Category / Subcategory | Yes | computable | 7 | 83% | 0.002 | Elections | +10.2 | World | -7.9 | 4859 |
Not-yet-instrumented: News intensity (#5) & Research tractability (#10) — telemetry, deferred; Counterparty (#12) — arrives with the Marketplace Catalog (PM on-chain, twin / category proxy).
Solidified findings promoted here (first draft — being rewritten). Working / unproven theories and the running historical log live in Open Hypotheses & Conclusions below.
The running log — working / unproven theories, plus a dated, historical list of
discoveries as we make them. When something solidifies it gets promoted up to its pillar above;
when it is later disproven it stays here, dated, so the record stays honest. Edit in
events_desk/conclusions.py.
| Open Hypotheses | Conclusions Reached |
|---|---|
| Edge & market selection | |
|
|
| Pillar 1: Improved Predictions | |
|
— no conclusions logged yet — |
| Efficiency & cost | |
|
|
| Capital allocation | |
| — no open hypotheses here yet — |
|
| Model performance | |
|
|
| Business & scale | |
|
|
How we steer research agents — the working principles behind every approach and prompt
version below. Rewritten 8/17 from the ground-up prompt-engineering research review (~90 external sources:
forecasting-system papers, controlled prompting studies, our own 2,858-wager record); each principle keeps
its MATES example where we have one. Source of truth: conclusions.py: PROMPT_PRACTICES;
the full research synthesis lives in the shared Google Doc; design-debate handoff at
events_desk/PROMPT_DESIGN.md.
An approach is the structural recipe around the prompts: which agents run, in what
order, with what triggers and guardrails. Historically our prompts (v1…v5) versioned independently of the
structure (red teams, screening, review rules); from Canopy onward the two are tracked together — prompts
version within an approach. Newest first; the prompt texts themselves live in the Prompt Registry
below. Source of truth: conclusions.py: APPROACHES.
The implementation layer beneath the Approaches above, in two parts: the versioned
research-agent prompts, organized by agent role (Principal / Associates–Screening / Red Team), and
the hardcoded guardrails (every magic number that shapes a decision) — surfaced here so we never
lose track of them. Sources of truth:
events_desk/ra_prompts.py and events_desk/backtest/research_redteam.py; guardrails in
conclusions.py. {brief} = the
market's rules; {access} = the info channel injected per context (live web / date-gated news / blind).
The frontier pricing model (Sonnet 5, validated champion). Versioned prompt set below — the active version is what run.py uses by default.
| Prompt | Date | Description | Change summary | Full text |
|---|---|---|---|---|
| v1 · Core evaluator v1-core |
2026-07-25 | Understand-before-pricing: restate resolution, classify the trigger, catch the near-certain trap, name the YES/NO meaning, and give an honest 80% interval. | Baseline. Understand-before-pricing skeleton. | |
| v2 · Deep research + calibration v2-deep |
2026-07-27 | v1 plus explicit deep-research discipline: reason about mechanical/process timing (fuel-loading, regulatory steps), fight overconfidence, and anchor on base rates before the narrative. | Added deep-research discipline: process/mechanical timing, anti-overconfidence, base-rates-first. | |
| v3 · + post-mortem lessons Active v3-postmortem |
2026-07-28 | v2 plus pitfalls learned from the agentic feedback loop — most notably: pin down the EXACT referenced event (e.g. UFC 329) and cross-check its date against the market close, instead of substituting a similarly-themed event. | Added the exact-referent cross-check (UFC 329 lesson) from the feedback loop. | |
| v4 · + primary-source fact verification v4-primary |
2026-07-30 | v3 plus the primary-source discipline from the 3.6 A/B: the biggest misses are FACT-verification failures, so confirm the exact resolution fact against the most authoritative source before committing to high confidence — else widen. | Added exact-fact verification vs authoritative/primary sources before high confidence (3.6 Prong B). | |
| v5 · + lookup/forecast base-rate discipline v5-baserate |
2026-08-12 | v4 plus a decision procedure for questions whose resolving fact does not exist yet: name the reference class, start from its base rate, let evidence adjust the anchor rather than replace it. v4's source-verification discipline is explicitly scoped to already-decided (LOOKUP) questions. | Added the LOOKUP-vs-FORECAST classification: undecided outcomes anchor on a named reference-class base rate; absence-of-reporting is weak evidence (KXSPACEXBANKPUBLIC post-mortem — all four models absence-anchored a not-yet-decided syndicate). | |
| v6 · Canopy rebuild v6-canopy |
2026-08-17 | The Canopy-approach prompt: a single coherent rebuild consolidating the v1-v5 lessons (rules-first reading, window discipline, exact referent, primary-source verification, base-rate anchoring) into motivated rules with explicit jurisdictions, plus a scoring-aware framing and a reasoning-first output contract (RA_VERDICT_SCHEMA_V6, probability last). | Ground-up rebuild for the Canopy approach (not composed from the v1-v5 blocks). One coherent document; every rule carries its reason; DETERMINED/UNDECIDED jurisdictions replace LOOKUP/FORECAST; the agent is told how it is scored (Brier + CLV + Kelly sizing, and that timid hedging costs like overconfidence); paired v6 output schema puts all reasoning fields BEFORE p_yes and adds question_class / reference_class / base_rate / what_would_change_this. Gate: A/B vs v5 on the k53 set before any big run. |
Budget open-weight models (gpt-oss-120b, DeepSeek-V3.1 via DeepInfra) that pre-screen the universe: cross-model consensus mean flags candidates; a >15pp split or side disagreement quarantines a market as blowup-risk. They run the IDENTICAL versioned prompt as the Principal (parity by design — the 8/10 validation held the prompt constant so the model was the only variable; a stripped cheaper screening prompt is a future experiment).
| Prompt | Date | Description | Change summary | Full text |
|---|---|---|---|---|
| Same as Principal · active version Active assoc-active |
2026-08-10 | Runs the active Principal prompt version verbatim (currently v3-postmortem). | Screening validation 8/10: Spearman(cheap_P, principal_P) = +0.457 solo, +0.563 for 4-run consensus. |
Adversarial review (run.py --redteam): an independent, tool-free RED attacker sees the point-in-time brief, blue's verdict, and the entry mid — then blue must rebut inside its live research conversation before finalizing. Validated 8/11 on the 8/4 244: reliably catches Bucket-A/B flaws, but universal review market-anchors — production use is SELECTIVE (quarantined markets only).
| Prompt | Date | Description | Change summary | Full text |
|---|---|---|---|---|
| Red Team · five-lens attack Active red-attack |
2026-08-11 | Resolution law / window check / knowability / evidence audit / steelman-the-market. Returns typed attacks (fatal/serious/minor), never its own probability. | v1. Known failure mode: red can INJECT a wrong resolution reading and blue capitulates (KXLEAVEPOWELLGOV 0.15→0.96 on a NO) — guardrail-teeth experiment queued. | |
| Blue rebuttal contract Active red-rebuttal |
2026-08-11 | Blue dispositions every fatal/serious attack (refuted / accepted / partial) and re-submits via the submit_estimate tool. Explicitly forbidden from courtesy-shrinking; an unrefuted fatal attack must move the estimate toward the market. | v1. Tool-path submission (forced-JSON over a long tool transcript loses ~50% of rebuttals on reasoning models). |
Canopy step 3 (reconcile.py): when the ensemble's spread or its divergence from the market crosses the gate — trigger arithmetic runs in code, never in a prompt — a fourth agent reads the anonymized member write-ups, names the crux they split on (or the assumption they share), settles it with its own targeted date-gated searches, and issues the final probability. AIA-style: no rebuttal channel back to the members (no capitulation risk), and fully price-blind. Fail-open: a reconciler error keeps the code aggregate.
| Prompt | Date | Description | Change summary | Full text |
|---|---|---|---|---|
| Reconciler · crux-finder Active reconciler-crux |
2026-08-18 | Find the ONE checkable fact / rules reading / base rate the panel actually turns on, settle it with evidence, judge cases not confidence, then decide — not average. Output pinned to the v1-v5 submit schema regardless of the ACTIVE version. | v1 — built for the Canopy 120-market sim (reconciler-only review arm; red team parked per 8/17 decision). |
Baked-in numbers that shape decisions. Review occasionally — moving any of these moves the strategy. active = live in code · designed = built, not yet wired in.
| Guardrail | Category | Rule | Why | Status |
|---|---|---|---|---|
| Fractional Kelly (¼) score.py: DEFAULT_KELLY_FRACTION |
Sizing | stake = 0.25 × full-Kelly | Full Kelly is growth-optimal but wildly volatile; the ¼ haircut trades a little growth for far smaller drawdowns. | active |
| Per-bet size cap score.py: SIZE_CAP = 0.10 |
Sizing | never stake > 10% of bankroll on one wager | Hard ceiling against a single position sinking the book, regardless of edge. | active |
| Wide-interval haircut score.py: CI_PENALTY=0.5, CI_WMAX=0.5 |
Sizing | shrink size up to 50% as the 80% interval widens toward 0.5 | Bet less when our own uncertainty (interval width) is high. | active |
| Brier benchmark = ENTRY mid score.py: score_bet |
Scoring | market_brier uses the entry-mid price, not the closing line | The fair forecasting test is us-vs-market at ENTRY; the closing line is hindsight (it has seen the info). | active |
| Price-level buckets dimensions.py: price_bucket |
Bucketing | edge-dist <0.15 = extreme · <0.35 = leaning · else balanced | Defines the "balanced corner" — our certified edge zone. Moving these lines moves what counts as the corner. | active |
| Horizon buckets dimensions.py: horizon_bucket |
Bucketing | ≤14d short · ≤60d medium · else long | Time-to-resolution slicing. | active |
| Size tiers analysis.py: _size_bucket |
Bucketing | <$10k small · <$100k medium · else large (volume) | Liquidity/size slicing. | active |
| Entry points entry.py: DEFAULT_FRACTIONS |
Method | enter at 25% / 50% / 75% of market life | Shows how edge/CLV decay as resolution approaches; 25% has been strongest. | active |
| Allocation filters analysis.py: ALLOC_STRATEGIES |
Allocation | conviction ≥0.10 and CI ≤0.35 for the filtered strategies | Only stake when our edge is meaningful and our interval is tight. | active |
| Calculation-market exclusion select_targets.py |
Policy | only logical-inference markets are simulated; calculation markets excluded | Calculation markets (a number crossing a threshold) degrade under our research. | active |
| Kalshi fee model Gate-3 cost model (to be codified) |
Cost | fee = 0.07 × P × (1−P) per contract | Net-of-cost returns. Standard taker rate — needs confirming vs Kalshi's current schedule. | active (confirm rate) |
| Resolution confidence cap resolution.py: resolution_guard |
Research (3.6) | graduated pull toward 0.5 by red-team severity | A/B v1: does NOT beat a free confidence shrink — overconfidence is systemic, blowups are fact-errors. Not shipped; Prong B (primary-source tools) is the lever. | tested — not shipped |
One row per simulation run, newest first. The metric block is the standard lens — ¼-Kelly sizing at the 25% entry point over the run's valid wagers — plus All·even CLV (even-weighted, every entry point) as the model-comparison reference. Strategy × entry sweeps live in the Allocation Approach table below. ARR is an inflated perfect-redeployment upper bound. Drag a column's right edge to resize it.
Catalog of markets in the cohort — category, our subcategory, outcome-driver, resolution, closing line.
One row per valid run × market × entry point (allocation-independent — no sizing). 7,935 wagers on file. The page never renders the rows (thousands, unbounded — that crashes the browser): pick a slice, get server-side summary stats, and export the line-level rows (with our reasoning + resolution notes) to Google Sheets. Same pattern as the Marketplace Catalog below.
Sizing is a LAYER re-scored over existing predictions (no new API). Objective: annualized return on deployed capital. Time-adjusted rules (·per day) push more capital toward faster-resolving edges — a 7-day and a 365-day wager with equal raw EV are NOT equal (the fast one recycles capital ~52×). Read Ret on capital + CLV + Brier skill together; Ann. ret is an inflated upper bound (assumes every dollar recycles at that rate all year). Funded = how many wagers the rule actually stakes (positive-EV / positive-Kelly only).
Even-weighted performance per segment — the corner where we beat the market (Brier skill > 0). Even weighting isolates prediction quality from the sizing rule.
20,955 markets · $11,987,645,859 volume catalogued. Enriched with the Structural Edge Score (higher = more structurally in our corner) — filter by Min SES, see avg SES per segment, and the Google-Sheet export is SES-ranked with the full dimensional profile.
Every metric & term we use, defined once — so we stop forgetting them. Edit in events_desk/conclusions.py.
Honest caveats. Results are early (n up to 40) and the testing environment is still being hardened toward real, leak-safe data. Brier-skill is negative on the broad cohort so far — we don't yet out-calibrate the market across all markets. Treat that as a target, not a verdict: the mission is to (1) find the segment — category, market size, structure — where we DO beat Kalshi, and (2) fix model overconfidence at the source (standing research-agent instructions), not just build allocation levers around it. Also: positive CLV has not yet meant positive realized return (favorite/longshot payout math); backtests overstate live (no market impact, perfect fills, survivorship).