Self-hosted AI investment committee

OpenInvest

Four independent LLM roles challenge each other and a CIO synthesizes one verdict — the decision stays yours.

$ /plugin marketplace add longsizhuo/openInvest
$ /plugin install invest@openinvest

Installs the MCP server (18 tools auto-registered) + committee skills. Then say “set up invest” to onboard. No API key needed (Coordinator path).

673 Installs|4,381 Copies

Why OpenInvest

Most “AI trading” demos quote returns. We measure what can actually be tested.

At a Sharpe around 1, proving skill from returns takes decades (t = SR·√T). So instead of selling untestable P&L, OpenInvest grades calibration — and rejects the parts that don’t hold up.

80 / 76 / 74%

Calibration-first

After band calibration (γ=1.1), P10–P90 coverage moves into the pre-registered [75%, 85%] band: 76/72/69% → 80/76/74% across 30/60/90-day horizons.

0 / 3

We reject what fails

TradingAgents-style analyst agents scored 50.3 / 57.4 / 64.7% on 30-day direction — all below the 71–73% naive baseline. 0 of 3 passed the pre-registered gate.

42%

Honest about small samples

At effective n=1, probability-band coverage collapses to 42%. We suppress single-sample point estimates instead of dressing them up as precision.

0.60

Calibrated, not overconfident

Raw LLM confidence clusters at 0.60 (median, n≈2,100). We correct that bias before any verdict reaches you.

Round 2Macro StrategistVIX · rates · FXQuant Analysttechnicals · blind to holdingsRisk Officerconcentration · tail riskCIOsynthesizes · confidenceBUY · HOLD · SELL

Three analysts argue from different evidence, cross-challenge in round 2, and a CIO synthesizes one verdict.

Product Philosophy

The Investment Operating System, Not Another Chatbot

OpenInvest is not another AI assistant, and it is not trying to replace your agent.

We believe the future belongs to increasingly capable personal agents such as Claude Code, Codex, Hermes, OpenClaw, and future Agent OSs. These agents already understand their users far better than a standalone investment application ever could. They naturally accumulate long term context, remember past conversations, observe user behavior, and evolve alongside their owners.

Instead of rebuilding these capabilities, OpenInvest focuses on one thing only: becoming the world's best investment decision engine.

OpenInvest Provides

  • A verifiable investment committee
  • Evidence-based reasoning
  • Long-horizon backtesting and calibration
  • Structured decision records
  • Market intelligence and research
  • Decision APIs and protocols that any agent can consume

Your Personal Agent Provides

  • Natural conversation
  • Long-term memory
  • Understanding of your preferences and risk tolerance
  • Decision reflection
  • Personalized coaching
  • Multi-turn discussion

In other words, OpenInvest generates high quality decisions, while your agent helps you make high quality choices.

An Intentional Separation

We do not want to compete with the rapidly evolving agent ecosystem by building another chat interface or another memory system. Instead, OpenInvest is designed to become the decision infrastructure that powers those agents.

Smarter Agents, More Valuable OpenInvest

As agents become smarter, OpenInvest becomes more valuable. Every improvement in the agent ecosystem immediately benefits OpenInvest users without requiring OpenInvest to reinvent conversation, memory, or personalization.

OpenInvest doesn't build another AI agent. It builds the decision engine that AI agents deserve.

Evidence

Calibration-first, on real data — wins and honest losses

Figures are recomputed from the public research archive. Look-ahead-prone windows are flagged; we show where the committee trails, not just where it leads.

The defensive edge

+0.00%

2022 bear market — the committee protected capital

When gold fell through 2022, the committee stayed +1.95% while buy-and-hold turned negative and trend / regime rules lost ~6%.

Walk-forward replay on GC=F, 2022 (pre-training-cutoff, no look-ahead). n=251 trading days.

experiments/ta-analysts/baselines/gold_fourth_arm_result.json

The honest ledger — where we trail, and what we rejected

Risk-adjusted return across three regimes

The committee leads in the 2022 bear and edges ahead in the bull — and, honestly, trails buy-and-hold in the sharp 2020 V-shaped crash.

† 2024–26 bull carries LLM look-ahead (upper-bound only). * 2020 / 2022 are clean. Sharpe ratio, GC=F.

experiments/ta-analysts/baselines/gold_fourth_arm_result.json

TradingAgents-style analysts vs a naive baseline

Fundamental, news and sentiment analysts all missed our pre-registered hit-rate gate — a naive “always-majority” baseline beat every one.

30-day directional hit-rate, Wilson 95% CI, n≈244/analyst. The single-sided bull window makes the ~73% baseline a high bar by design.

docs/wiki/16-ta-analysts-experiment.md

Cross-model validation — one lucky cell isn’t a signal

Across windows × models, exactly one analyst cell ever cleared the gate. We require replication before believing a result — 孤证不立.

Gate = Wilson CI lower bound above both the naive baseline and the mechanical mapping.

docs/wiki/16-ta-analysts-experiment.md

Calibration & rigor

Where multi-agent value actually comes from — three mechanisms: independent information sources, de-correlated sampling + aggregation, and deterministic guardrails.
Where multi-agent value actually comes from — three mechanisms: independent information sources, de-correlated sampling + aggregation, and deterministic guardrails.
Why returns are untestable: t = SR·√T. At Sharpe ≈ 1 you need decades to prove skill — so we grade calibration, not P&L.
Why returns are untestable: t = SR·√T. At Sharpe ≈ 1 you need decades to prove skill — so we grade calibration, not P&L.
Band calibration (γ=1.1) lifts P10–P90 coverage into the pre-registered [75%, 85%] band at every horizon.
Band calibration (γ=1.1) lifts P10–P90 coverage into the pre-registered [75%, 85%] band at every horizon.
View the remaining figures
Pre-registered out-of-sample acceptance: the calibration layer lifts coverage into [75,85]% and improves Brier at every horizon (fit 2007–17, OOS 2018–26).
Pre-registered out-of-sample acceptance: the calibration layer lifts coverage into [75,85]% and improves Brier at every horizon (fit 2007–17, OOS 2018–26).
Raw LLM confidence clusters at 0.60 (median, n ≈ 2,100) — the overconfidence we correct for.
Raw LLM confidence clusters at 0.60 (median, n ≈ 2,100) — the overconfidence we correct for.
Spurious precision: at effective n=1, probability-band coverage collapses to 42%. We suppress single-sample estimates.
Spurious precision: at effective n=1, probability-band coverage collapses to 42%. We suppress single-sample estimates.
Coverage across eras × regimes — solid except documented thin-data cells (e.g. gold 2007–09 downtrend-90d, n=14).
Coverage across eras × regimes — solid except documented thin-data cells (e.g. gold 2007–09 downtrend-90d, n=14).
Conditional vs unconditional Brier by era — gold bear conditional 0.356 / unconditional 0.333.
Conditional vs unconditional Brier by era — gold bear conditional 0.356 / unconditional 0.333.
How OpenInvest compares to other multi-agent trading frameworks.
How OpenInvest compares to other multi-agent trading frameworks.
The five-stage evaluation protocol behind every claim.
The five-stage evaluation protocol behind every claim.

Methodology & Deep Dive

Under the Hood: OpenInvest’s 5 Core Systems

A rigorous walk-through of the architectural decisions, math, and validation results that power our calibrated committee.

System 01

1. Dreaming: Sleep-Cycle Memory Integration

To fix the LLM’s zero-session memory, a three-stage nightly job consolidates historical decisions. We apply volatility-aware opportunity cost thresholds to HOLD verdicts to counter over-conservatism.

Stage Mechanism

Starts at midnight. Reads all past verdicts and subsequent market price movements, tagging each with the true market regime on that decision day, generating short-term-recall files.

System 02

2. Probability Table Pathification

Instead of untestable point estimates, we query 20+ years of historical data for the active regime. Paths are sorted into 4 mutually-exclusive categories using volatility units.

90d Forward Path SimulationRegime-conditioned
0%-1 ATRDip (-1.2 ATR)End (+4.5%)
Rigorous Calibration: Conditional prior shrinkage (k=80) and band expansion (γ=1.1) ensure P10–P90 coverage lands in the pre-registered [75%, 85%] band.
System 03

3. Dual Execution Paths

The same prompts run through two distinct implementations. Disagreements between Claude and DeepSeek are used as a model divergence validation signal rather than a bug.

Coordinator Path (Claude Code)

Local sandboxed subprocesses. Zero-cost execution powered by user subscription.

Sandbox: subprocess isolation · Model: Claude 4

Direct Path (FastAPI + DeepSeek)

ThreadPool execution for cron-based automation and real-time live SSE stream.

Concurrency: ThreadPoolExecutor · Model: DeepSeek-Chat
System 04

4. Bayesian Optimization & Negative Results

We optimize prompts and allocation rules programmatically. Rather than hiding failed assumptions, we document them to ensure scientific integrity.

Calibration Formulas
λ = eff_n / (eff_n + k)k = 80

Small-sample shrinkage — pull thin conditional buckets toward the asset’s unconditional distribution.

q′ = γ · qγ = 1.1 · P10/P90

Bandwidth expansion on the P10/P90 & downside quantiles — fixes the structural under-coverage.

t ≈ SR · √TSharpe 1 → ~4 yr

Why we grade calibration, not P&L: proving Sharpe-1 skill takes years; calibration is testable today.

Reward Signal
Reward Function: Annualized Return - 0.5 × MaxDrawdown + 0.2 × (Sharpe - 1)
Optimization Impact (Reward)
v0 Baseline (100% HOLD)0.00
v1 Prompt (Cash Opportunity)+0.398
Optuna Best / DSPy Optimized+0.422 (Dev accuracy +10pp)
Intellectual Honesty & Negative Findings

Intellectual Honesty: Optuna revealed that multi-round debate is a placebo (1 round = 3 rounds in performance). TradingAgents subagents also performed below naive baselines. Focus remains on calibrated CIO veto rights.

System 05

5. Markdown DB & Concurrency Locks

We use YAML frontmatter for schema validation (Pydantic) and the Markdown body for direct LLM ingestion. Atomic transactions are secured via fcntl file locks.

portfolio.mdYAML + GFM Markdown
---
schema_version: 2
cash:
  CNY: 50000
holdings:
  - symbol: NDQ.AX
    units: 50
    avg_cost: 38.50
---

# Current Holdings
- CNY Cash: ¥50,000
- NDQ.AX: 50 units @ A$38.50
Pydantic validation is run prior to atomic file writes
Concurrency Guarantees

Fcntl read-modify-write (RMW) locking prevents race conditions between chat bots, schedulers, and APIs.

Stress Test (0 lost updates)
50 threads × cash +1 CNY = precise 50.00 CNY delta

Reference Documentation

Browse and read the raw, unedited Markdown documentation and Architecture Decision Records (ADRs) directly from our codebase.