v0.2.0Apache 2.0Python 3.12+

ResearchForge Documentation

Local-first AI research and benchmarking workflow for teams that need evidence, not guesses.

💡
Start here: the open-source workflow is local and reproducible, while the enterprise layer adds shared control and team-scale coordination.
Open source
Included in the Apache-licensed CLI
  • • IDE-first workflow with Claude Code and Cursor
  • • literature search and ranking
  • • baseline, run, validate, and ship
  • • local worktrees, protected paths, and audit log
  • • Docker and local Python execution
Enterprise add-on
Optional team controls and shared infrastructure
  • • self-hosted Hub and approval workflow
  • • multi-user coordination across machines
  • • air-gapped deployment
  • • workload tagging and shared lineage
  • • policy and governance layers for regulated teams
Start here

The shortest path to a real ResearchForge run

2 min
01 · Install
Python 3.12+ and Git are enough to begin.
02 · Install for your IDE
Register the Claude/Cursor workflow once for your machine.
03 · Start in the IDE
Type the slash command or @mention and approve the steps.
bashbash
pip install "researchforge[serve]"
researchforge all install --user

# Open Claude Code or Cursor and run:
/researchforge-start
# or
@researchforge-start

This is the recommended entry point. The detailed command reference below is for advanced workflows, CI/CD, and automation — not the default path most users should start with.

How ResearchForge works

ResearchForge implements a six-stage loop that converts a research question into a validated, shippable result with full lineage:

01
Search
arXiv → ranked papers → local knowledge base
AI generates domain-specific queries from your objective
02
Hypotheses
Papers + domain → testable hypotheses
Claude / Cursor reads papers and writes hypotheses.yaml
03
Baseline
Freeze the current metric — immovable reference
Every experiment is measured against this exact number
04
Run
Isolated git worktrees — one per experiment
Your checkout is never touched; each variant runs independently
05
Validate
Re-run the winner N times to confirm stability
"Validated" is only earned here — one run is never enough
06
Ship
Clean branch + engineering report + full lineage
One commit on the frozen baseline, never auto-pushed
↺autorun closes the loop — re-synthesizes hypotheses from measured results and expands the experiment graph until it stalls, hits the target, or runs out of time.

The key design principle: nothing moves until it has evidence. The baseline is immovable. Experiments run in isolation. The winner is only shipped after validation confirms it isn't a lucky seed.

💡
You write the eval script once — or Claude Code writes it for you. ResearchForge runs it to measure the baseline, then runs it again inside each experiment worktree. The AI (Claude Code / Cursor) only patches your implementation code — it never touches your benchmark or evaluation logic once the contract is approved.

Where does the eval script come from?

This depends on your project. ResearchForge handles all three cases:

✅
You already have a benchmark script
ResearchForge scans your repo and auto-detects scripts in benchmarks/, evaluate.py, scripts/eval.py, etc. It pre-fills the contract with full_command: python benchmarks/evaluate.py. You just review and approve.
🤖
You have tests but no benchmark script
Claude Code / Cursor writes the eval script for you during the project setup phase, before any experiments run. It writes artifacts/results.json in the standard format. You review it, then approve the contract.
✍️
You write it yourself
The contract wizard outputs full_command: "# TODO: command that writes the result file" as a placeholder. You write the eval script (or have Claude Code or Cursor write it), point the contract at it, and approve.
⚠
Once you approve the contract, the eval script's path is locked into protected_paths. No experiment can modify it. This is the guarantee that your benchmark stays stable across the entire experiment run.

Requirements

ResearchForge requires Python 3.12+ and Git — nothing else. An IDE (Claude Code or Cursor) is recommended but optional: set any AI API key and use the standalone CLI path instead. No Node.js, no Docker (unless you want container isolation), no cloud account needed.

💡
ResearchForge works in three modes: Claude Code (/researchforge-start), Cursor (@researchforge-start), or standalone with any AI API key (ANTHROPIC_API_KEY, GEMINI_API_KEY, or OPENAI_API_KEY). No IDE required for the standalone path — researchforge research synthesize calls the AI directly.
RequirementVersionWhy
Python3.12+The ResearchForge CLI and execution engine
Gitany recentWorktree isolation — one worktree per experiment
Claude Code (optional)latestRecommended AI layer: reads papers, writes patches via /researchforge-* skills.
Cursor (optional)latestAlternative AI layer: same capabilities via @researchforge-start MDC rules.
API key (optional)anyStandalone mode: set ANTHROPIC_API_KEY, GEMINI_API_KEY, or OPENAI_API_KEY — no IDE needed.

Install

Install ResearchForge with a single pip command. The base package bundles all three AI provider SDKs (Anthropic, Google, OpenAI) so no extra install is needed for standalone mode. Add [serve] only if you want the live web monitor dashboard.

Requires Python 3.12+ and Git. No Node.js required.

Standard install

pipbash
pip install researchforge          # includes built-in AI providers (Anthropic, Gemini, OpenAI)
pip install "researchforge[serve]" # add only if you want the live web monitor

Install with IDE integrations

bashbash
# After pip install, register skills/rules:
researchforge all install --user

# Claude Code only:
researchforge claude install

# Cursor only:
researchforge cursor install

Install from source

bashbash
git clone https://github.com/forger-labs-hq/researchforge
cd researchforge
pip install -e ".[serve,dev]"

Docker (no Python on host)

bashbash
docker run --rm -v "$PWD":/workspace -w /workspace \
  ghcr.io/forger-labs-hq/researchforge:latest \
  researchforge research search "your query"
ℹ
The Docker image includes Python 3.12, all dependencies, and Git. Mount your project at /workspace.

The IDE-first workflow

The recommended way to use ResearchForge is through Claude Code (/researchforge-start) or Cursor (@researchforge-start). The IDE handles all the creative steps — reading papers, forming hypotheses, writing patches — while the CLI handles the deterministic steps: freezing baselines, running experiments, validating winners, shipping. You type three things total: approve (contract), approve (plan), ship.

bashbash
# Recommended — start in the IDE
/researchforge-start
# or
@researchforge-start
💡
This is the path most users should follow first. The CLI commands below are the underlying engine for CI/CD, automation, and advanced users.

Quickstart (2 minutes)

Install ResearchForge, open your IDE, and start the guided workflow. This is the shortest route for real usage.

bashbash
# 1. Install
pip install "researchforge[serve]"

# 2. Open Claude Code or Cursor
#    Type one of these:
/researchforge-start
# or
@researchforge-start

# 3. Approve the contract, run the baseline, and let the agent do the rest
ℹ
The commands shown in this marketing site are examples of the ResearchForge CLI surface. This repo is a Next.js marketing site, not the Python CLI implementation itself, so they are documentation examples rather than commands you can execute in this project folder.

The IDE-first workflow

The intended way to use ResearchForge is through your IDE. Type one slash command or @mention and Claude Code / Cursor takes over: scans your repo, writes the eval script if needed, searches literature, generates hypotheses, runs experiments, and presents results — asking your approval at every consequential step. You approve; they execute.

💡
The CLI commands further in this page are what the IDE runs under the hood. You can run them manually for CI/CD, but you never have to write them yourself.
Claude Code / Cursor creates all of this for you
✓The eval script — benchmarks/evaluate.py that writes artifacts/results.json — Claude writes it from scratch if you don't have one. You never need to.
✓The contract — .researchforge/contract.yaml — objective, metric, protected paths, eval commands
✓Hypotheses — Testable ideas grounded in arXiv papers — for your approval before anything runs
✓Experiment patches — Git diffs applied to src/ only — eval scripts and tests are cryptographically locked
✓Engineering report — Full lineage, results, and reasoning shipped alongside the winner branch

The full loop — search to shipped branch

The complete research pipeline as it runs inside your IDE. Claude Code or Cursor drives every step — you only type your objective and approvals. Each dashboard panel below is presented inline in the chat, exactly as you’d see it in a real session.

01 Query
arXiv full-text search + relevance ranking
02 Landscape
Research directions · landmark papers · hypotheses
03 Experiments
Isolated worktrees · eval script · results dashboard
04 Ship
Validation × N · clean branch · engineering report
Claude Code · researchforge · ROGII wellbore geology
You
/researchforge-start
Claude Code / Cursor
Scanning repository...
📁 Python project · src/predictor/ · tests/ · pyproject.toml
⚠️ No benchmark script detected.
What do you want to improve?
You
Improve RMSE on the wellbore geology prediction task. Input is LWD log sequences, target is TVT (rock layer position).
Claude Code / Cursor
IDE sets up: writing benchmarks/evaluate.py · locking evaluation contract · defining metric & constraints
Searching arXiv for relevant literature...
01 · Literature Queryfull arXiv search · relevance scoring · stored to knowledge base847 candidates · 20 stored
[0.94]Sequence-based lithology prediction using deep learning
[0.91]Transfer learning for geophysical core analysis
[0.88]Attention mechanisms in subsurface sequential data
[0.85]Self-supervised pre-training on well log representations
[0.82]Multi-task learning for formation evaluation from LWD
02 · Research Landscapedirections · landmark papers · evidence claims · generated by IDE3 directions · 7 hypotheses
Sequential Modelling8 papers
🏆 "Sequence-based lithology prediction" [0.94]
LSTM/GRU encoders outperform feature-based — 3 studies agree
Multi-task Learning4 papers
🏆 "Multi-task formation evaluation" [0.85]
Auxiliary gamma+resistivity prediction → +10–15% on primary
Transfer & Self-supervised5 papers
🏆 "Transfer learning for geophysical" [0.91]
Pre-training on adjacent basins helps; limited evidence for LWD self-supervised
02 · Hypothesesgrounded in landmark papers · scored by evidence strength · your approval required
✓hyp-001Sequential encoding with depth positional normalisation
✓hyp-002Multi-task auxiliary prediction (gamma + resistivity)
✓hyp-003Per-well feature normalisation before training
○hyp-004Self-supervised pre-training on unlabelled logs
○hyp-005-0073 more lower-priority hypotheses
Approve hyp-001, hyp-002, hyp-003?
You
Approved. Run them all.
Claude Code / Cursor
Baseline frozen: rmse=15.2441·9 experiments · isolated worktrees, one at a time · main branch untouched
03 · Experiment Dashboardparallel git worktrees · eval script runs in each · live results vs baselinebaseline=15.24 · rmse lower is better
IDStatusHypothesis patchRMSEΔ
exp-001PASSPer-well normalisation (hyp-003)12.89+2.35
exp-002PASSGRU encoder + depth encoding (hyp-001)11.44+3.80
exp-003PASSMulti-task: gamma + resistivity (hyp-002)9.22+5.98
exp-004FAILSelf-supervised pre-training16.11−0.87
exp-005PASSMulti-scale rolling window (5/20/50)7.88+7.36
exp-006PASSGRU + per-well norm (builds on exp-001)7.11+8.13
exp-007REJETransformer encoder — constraint ✗ p95>budget9.03+6.21
exp-008REJEEnsemble weighting — Δ below threshold15.33−0.09
exp-009WINNGRU + multi-task + multi-scale (exp-003 × exp-005)6.10+9.14
04 · Validation — exp-009 × 5 runs
Re-running winner with different seeds to confirm stability before shipping
6.0891 · 6.1204 · 6.0977 · 6.1108 · 6.0985
✓ mean=6.1033 stdev=0.0114 — Stable
Ship as feat/multi-task-gru-winner?
You
Ship it.
Claude Code / Cursor
04 · Engineering Reportauto-generated · full lineage · shipped with branchexp-009 · feat/multi-task-gru-winner
Primary metric:rmse 15.24 → 6.10
Improvement:60.0% reduction
Papers:20 top results from arXiv
Hypotheses:7 generated · 3 tested
Experiments:9 run · 3 rejected · 1 winner
Validation:mean=6.10 stdev=0.011 n=5
✓ Branch: feat/multi-task-gru-winner
✓ Only winning changes applied — checkout clean
✓ Lineage: 9 experiments · 3 rejected with reasons · full audit trail
✓ Report: .researchforge/reports/exp-009-final.json

Claude Code — full walkthrough

From first command to shipped branch. You type 5 things; Claude does the rest.

Claude Code — chat session
You
/researchforge-start
Claude Code
Scanning repo... 📁 src/classifier/, tests/
⚠️ No benchmark script detected.
What do you want to improve?
You
Text classification. Improve F1, keep p95 latency under 200ms.
Claude Code
Writing benchmarks/evaluate.py:
import json, pathlib, time
from src.classifier import Classifier
clf = Classifier()
latencies, correct = [], []
for item in load_test_data():
    t0 = time.perf_counter()
    correct.append(clf.predict(item["text"]) == item["label"])
    latencies.append((time.perf_counter()-t0)*1000)
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
    "schema_version": 1,
    "primary_metric": {"name": "f1", "value": compute_f1(correct)},
    "secondary_metrics": {
        "p95_latency_ms": sorted(latencies)[int(len(latencies)*0.95)]
    },
    "sample_count": len(correct), "seed": 42
}))
Contract: f1 maximize · constraint: p95_ms ≤ 200
editable: src/, config/ · protected: benchmarks/, tests/
Approve?
You
Approved.
Claude Code
✓ Locked. Searching arXiv... 3 hypotheses generated:
hyp-001: Early-exit head on layer 4 (paper [0.88])
hyp-002: Knowledge distillation 12→4 layers (paper [0.91])
hyp-003: INT8 quantization of attention (paper [0.94])
Approve all three?
You
Yes.
Claude Code
3 isolated worktrees, run one at a time · main branch untouched
exp-001 f1=0.889 p95=94ms Δ=+0.021 ✓ PASS
exp-002 f1=0.901 p95=88ms Δ=+0.033 ✓ PASS ← winner
exp-003 f1=0.864 p95=71ms Δ=-0.004 ✗ REJECTED
Validated × 3 · mean=0.9012 stdev=0.0003 ✓
Ship as feat/distillation-winner?
You
Ship it.
Claude Code
✓ Branch: feat/distillation-winner
✓ Report: .researchforge/reports/exp-002-final.json
✓ Checkout clean · 3 experiments in lineage

Cursor — full walkthrough

Same workflow via @researchforge-start. This repo already has a benchmark script — RF detects it automatically.

Cursor — chat session
You
@researchforge-start Reduce RAG cost without ROUGE-L dropping below 0.85.
Cursor
Found benchmarks/eval_rag.py ✓
Contract: rouge maximize · constraint: cost ≤ $0.005/q
Approve? (y/n)
You
y
Cursor
2 hypotheses from literature:
hyp-001: gpt-4o-mini for non-critical steps
hyp-002: Result caching for repeated queries
Run both?
You
y
Cursor
exp-001 rouge=0.871 cost=$0.0028 ✓
exp-002 rouge=0.889 cost=$0.0041 ✓ winner
Validated × 3. Ship?
You
Ship.
Cursor
✓ Branch + report + lineage ready.

✨ Cross-IDE state sharing

This is one of ResearchForge's strongest features and almost always overlooked. Both Claude Code and Cursor read and write the exact same .researchforge/ directory. They share 100% of state — papers, hypotheses, baselines, experiment results, lineage — in real time.

🔄Start in Claude Code. Continue in Cursor. Or run both simultaneously.
🤖
One searches, one ships
Have Claude Code run the literature search and generate hypotheses while Cursor reviews results and prepares the ship command — both working on the same project state.
🔁
Resume anywhere
Start a research loop in Claude Code, close it, open Cursor, type @researchforge-results — it picks up exactly where Claude left off. No re-running, no lost state.
👥
Two AI brains, one experiment set
Claude Code generates hypotheses from papers. Cursor writes the experiment patches. Both write to the same lineage. The results are indistinguishable from a single-IDE run.

What exactly is shared

File / directoryWhat it containsBoth IDEs can
.researchforge/researchforge.dbAll project state: papers, hypotheses, contract, baseline, plans, experimentsRead via CLI commands
.researchforge/synthesis/landscape.yaml + hypotheses.yaml the AI writesRead, write (AI writes, CLI validates)
.researchforge/experiments/patches/One unified diff per experiment variantRead, write (AI writes patches)
.researchforge/worktrees/Isolated git checkouts at the baseline commitRead (managed by RF)
.researchforge/artifacts/Per-execution stdout, stderr, diff, results.jsonRead results, interpret
.researchforge/reports/Engineering + research reportsRead, build with report build
.researchforge/research-log.mdLiving context autorun feeds back to the AIRead (autorun writes)

Example — resume mid-loop in a different IDE

Session A: Claude Code (earlier today)
✓ researchforge research search "wellbore geology"
✓ researchforge hypotheses generate → 7 hypotheses
✓ researchforge baseline run → rmse=15.24
⚡ Claude Code session closed — laptop restarted
Session B: Cursor (right now — different IDE, same project directory)
You
@researchforge-results
Cursor
Reading .researchforge/ ...
📋 Project: wellbore-geology · baseline: rmse=15.24
📚 20 papers stored · 7 hypotheses (3 approved, awaiting run)
Ready to run experiments. Approve and I'll start?
You
Yes, run them.
Cursor
Running 3 experiments from where Claude Code left off... (same lineage, same baseline)
💡
Commit .researchforge/ to git and your whole team shares the research state. Every team member's Claude Code or Cursor session will see the same papers, hypotheses, and results — regardless of machine.

Install IDE skills/rules

bashbash
pip install "researchforge[serve]"
researchforge all install --user   # → ~/.claude/skills/ and ~/.cursor/rules/
researchforge all status
Claude Code
/researchforge-start
/researchforge-baseline
/researchforge-run
/researchforge-results
/researchforge-ship
Cursor
@researchforge-start
@researchforge-baseline
@researchforge-run
@researchforge-results
@researchforge-ship
ℹ
Both IDEs share .researchforge/ state — start in one, continue in the other.

Core concepts

Baseline
A frozen measurement of your metric committed before any experiments run. Immovable unless explicitly reset. The reference all improvements are measured against.
Hypothesis
A testable idea grounded in retrieved literature. e.g. “Layer norm before attention (paper-003) will improve accuracy by ≥2%”. Each hypothesis maps to one or more AI-generated patches.
Experiment
One isolated run of your code with a specific config applied. Gets its own git worktree, its own venv, its own env vars. Completely independent from every other run.
Worktree
A git worktree is a secondary checkout of your repo at the baseline commit. Your main branch is never touched. When the experiment finishes, the worktree is cleaned up.
Subagent
The process running inside a worktree. Runs the eval script after the AI patch is applied. Reports results via artifacts/results.json.
Lineage
The full directed acyclic graph of baseline → experiments → rejections → promotions. Every result, every rejection reason, every config is immutably stored.
Stall
N consecutive experiments with no improvement over the current best. When stall is reached, the run loop stops automatically.
Protected path
A path declared in permissions.protected_paths. A patch touching it is recorded as rejected at import — the experiment never runs at all. The path guard reads the diff before execution, so there is no race window.

How metrics are captured — the results.json contract

ResearchForge does not scan stdout. Your benchmark script writes a structured artifacts/results.json file after every run. ResearchForge reads that file to compare experiments against the baseline.

💡
You write this script once when setting up the project. It lives in a protected path (e.g. benchmarks/) and is never modified by the AI during experiments. The AI only patches your implementation code insrc/ or config/.
benchmarks/evaluate.py — your eval script (write this once)python
"""
Your benchmark script. Lives in a protected path.
ResearchForge runs this to measure the baseline, then runs it again
inside each experiment worktree (with the AI's patch applied to src/).
"""
import json
import pathlib
from my_model import load_and_eval  # ← AI can patch this

# Run your evaluation
accuracy, p95_ms, cost = load_and_eval(dataset="benchmark-v2")

# Write results in the standard ResearchForge format
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
    "schema_version": 1,
    "primary_metric": {"name": "accuracy", "value": accuracy},
    "secondary_metrics": {
        "p95_latency_ms": p95_ms,
        "average_cost_usd": cost,
    },
    "sample_count": 1200,
    "seed": 42,
    "metadata": {"dataset_version": "benchmark-v2"},
}), encoding="utf-8")

print("evaluation complete")

What happens during an experiment

1. ResearchForge creates a git worktree at the baseline commit
.researchforge/worktrees/exp-003/ ← isolated copy at baseline commit
2. Claude Code / Cursor generates a git patch from the hypothesis
diff --git a/src/model.py b/src/model.py
+NORMALIZATION = "layer"
+USE_POSITIONAL_ENCODING = True
3. Patch applied to worktree (editable paths only)
git apply change.patch ✓
4. Path guard checks patch didn't touch protected paths
benchmarks/ → protected ✓ untouched
5. Your eval script runs in the worktree
python benchmarks/evaluate.py --subset full
6. ResearchForge reads artifacts/results.json
accuracy=0.891 → Δ+0.031 vs baseline 0.860 → PASS ✓

Multiple metrics & constraints

Your eval script can write as many secondary metrics as needed. Specify hard constraints to automatically reject experiments that trade too much quality for speed (or cost).

artifacts/results.json schemajson
{
  "schema_version": 1,
  "primary_metric": {"name": "accuracy", "value": 0.891},
  "secondary_metrics": {
    "p95_latency_ms": 143.2,
    "average_cost_usd": 0.0031,
    "f1_macro": 0.877
  },
  "sample_count": 1200,
  "seed": 42,
  "metadata": {"dataset_version": "benchmark-v2", "model_params": 7340032}
}
contract — define constraints in the objectiveyaml
objective:
  description: >
    Improve accuracy on the classification benchmark while keeping
    p95 latency under 200ms and cost under $0.005 per query.
  primary_metric:
    name: accuracy
    direction: maximize
  hard_constraints:
    - name: p95_latency_ms
      operator: <=
      value: 200
    - name: average_cost_usd
      operator: <=
      value: 0.005
💡
The contract wizard guesses metric name and direction from plain-English objectives. "Improve accuracy" → accuracy / maximize. "Reduce p95 latency below 200ms" → latency_ms / minimize. You can always edit the contract YAML manually afterward.

Screening funnel

For slow full benchmarks, define a fast screening subset. Experiments must beat the baseline on the cheap screen before the expensive full eval runs.

contract YAMLyaml
execution:
  screening_command: python benchmarks/evaluate.py --subset screening
  full_command:      python benchmarks/evaluate.py --subset full
  result_file: artifacts/results.json
benchmarks/evaluate.py — handle --subsetpython
import sys
import json, pathlib

subset = "screening" if "--subset" in sys.argv and "screening" in sys.argv else "full"
# screening = fast 10% sample; full = complete eval
dataset_size = 120 if subset == "screening" else 1200

accuracy = run_eval(n_samples=dataset_size)

pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
    "schema_version": 1,
    "primary_metric": {"name": "accuracy", "value": accuracy},
    "sample_count": dataset_size,
    "seed": 42,
}), encoding="utf-8")
ℹ
Experiments that fail screening are marked REJECTED (screen) in the lineage — they never run the expensive full eval, saving significant compute on non-promising hypotheses.

Searches arXiv end-to-end, ranks results by relevance, stores top papers in the local knowledge base. The query is inferred from your objective by default; use --query to override.

bashbash
researchforge research search [--query/-q TEXT]... [flags]
FlagDefaultDescription
--query / -qfrom objectiveRepeatable: add extra search terms (e.g. -q "YOLOv5 pruning" -q "knowledge distillation")
--select / --n20Max papers to store in the knowledge base
--min-score0.6Minimum relevance score (0–1.0)
--categories / -callarXiv category filter, repeatable (e.g. --categories cs.CV --categories cs.LG)
--max-candidates200Fetch this many candidates before deduplication and ranking
--providernoneAI provider (anthropic|google|openai) — generates domain-specific queries instead of keyword fallback
--forcefalseRe-run search even if papers already exist
--jsonfalseOutput results as JSON
💡
For best results, make your objective domain-specific (e.g. "Improve YOLOv5s mAP@0.5 on COCO128 object detection"). Generic objectives produce irrelevant papers. Use --provider anthropic to let the AI generate targeted queries.

papers (manage knowledge base)

bashbash
# List stored papers
researchforge papers list

# Show details for a specific paper
researchforge papers show paper-003

baseline run

Runs your benchmark script once to freeze the current metric as the immovable reference. Must be run after researchforge contract approve. The objective, timeout, and metric come from the approved contract — not from flags.

bashbash
researchforge baseline run [--check] [--json]
FlagDefaultDescription
--checkfalseDry-run: validate the contract and benchmark command without executing
--jsonfalseOutput result as JSON
⚠
Once frozen, the baseline is immovable for the lifetime of the project. This is intentional — it prevents baseline creep.
bashbash
# Check current baseline
researchforge baseline show
researchforge baseline status    # alias for show

hypotheses

Hypotheses are written by your AI (Claude Code, Cursor, or direct API) and imported. hypotheses generate calls an AI provider directly — no IDE needed.

bashbash
# Generate hypotheses using a direct AI API key (no IDE needed)
researchforge hypotheses generate [--provider anthropic|google|openai] [--model TEXT] [--no-import] [--json]
# Equivalent alias:
researchforge research synthesize [--provider ...] [--model ...] [--no-import] [--json]

# Import a hypotheses.yaml written by Claude/Cursor (validates schema)
researchforge hypotheses import .researchforge/synthesis/hypotheses.yaml

# List all hypotheses and their status
researchforge hypotheses list

# Show details for one hypothesis
researchforge hypotheses show hyp-002
💡
hypotheses generate requires ANTHROPIC_API_KEY, GEMINI_API_KEY, or OPENAI_API_KEY to be set. In the IDE workflow, Claude Code or Cursor writes hypotheses.yaml directly and you run hypotheses import to validate and store it.

experiment (plan & manage)

There is no plan top-level command. Experiment plans are managed via researchforge experiment.

bashbash
# Generate a plan template for a hypothesis (Claude/Cursor fills in the patches)
researchforge experiment plan <hyp-id>

# Import a plan.yaml (written by Claude/Cursor or by hand)
researchforge experiment import plan.yaml

# Start a run (import + approval + execute)
researchforge experiment start plan.yaml
researchforge run plan.yaml          # alias

# List experiments and status
researchforge experiment list

# Resume or discard an interrupted run
researchforge experiment resume run-001
researchforge experiment abandon run-001

# Validate the contract (not the plan)
researchforge contract validate

autorun — the autonomous loop

The headline feature: a fully autonomous research loop that runs overnight. Each round searches the experiment graph for the best node to expand, tries available hypotheses there, measures results, and re-synthesizes new ideas from what it found. Stops when it stalls, hits --target, or runs out of --max-hours.

⚠
The contract still needs your typed approval before the first batch. --yes skips the first-batch prompt only — never the contract gate. Ctrl-C is safe and expected; autorun --resume continues with the same stall counter.
bashbash
# Start an overnight run
researchforge serve --background        # start live monitor first
researchforge autorun --target 0.85 --max-hours 8 --yes

# Short run to see how the loop behaves before committing overnight
researchforge autorun --max-rounds 2 --observe

# Continue an interrupted loop
researchforge autorun --resume
FlagDefaultDescription
--stall INT2Stop a plan after N consecutive non-improvements
--global-stall INT3Stop the whole loop after N rounds with no improvement anywhere
--max-rounds INTnoneHard cap on synthesis rounds
--max-hours FLOATnoneWall-clock limit (overnight safety cap)
--target FLOATcontract's target_valueStop as soon as the primary metric reaches this
--compound / --no-compoundonBuild each round on a node of the graph instead of on the baseline
--explore FLOAT0.0UCB1 constant — 0 always expands the current best; higher revisits under-explored branches
--merge / --no-mergeoffTry combining two independent winners each round
--observe / --no-observeoffAI reads each run's logs and records a paragraph on what it showed
--resynthesize / --no-resynthesizeonGenerate new hypotheses from measured results each round
-p / --provider TEXTautoanthropic | google | openai
-y / --yesoffUnattended — skip the first-batch approval prompt
--resume—Continue the interrupted loop from .researchforge/autorun.json
💡
Start with --max-rounds 2 --observe to watch two rounds end to end, read the results, and confirm the loop is reasoning about your benchmark before running overnight.

run

Alias for researchforge experiment start. Imports the plan, asks for one typed approval, then executes all experiments in isolated git worktrees. Stall, parallel, timeout, and threshold come from the approved contract — not from flags.

bashbash
researchforge run <plan.yaml> [--yes] [--monitor/--no-monitor] [--json]
FlagDefaultDescription
--yesfalseSkip the typed approval prompt (for CI/CD)
--monitor / --no-monitorautoStart/skip the live web monitor
--jsonfalseOutput progress as JSON lines
ℹ
--parallel, --stall, --worker, and --tags are not flags on this command — stall and parallel are set in researchforge.yaml; worker/tags are enterprise roadmap features.

validate

Re-runs the winner N times (from the contract's validation.repeat_finalists) to confirm the result is stable, not a lucky seed.

bashbash
researchforge validate <run-id> [--experiment/-e TEXT] [--yes] [--json]
FlagDefaultDescription
run-id—Required: the run to validate (e.g. run-001)
--experiment / -ebest in runValidate a specific experiment ID instead of the best
--yesfalseSkip confirmation prompt
--jsonfalseOutput as JSON
ℹ
The number of repeat runs and pass threshold come from validation.repeat_finalists in the approved contract, not from flags.

ship

ship is a subcommand group, not a single command. Use ship branch for a local clean branch, then optionally ship pr to open a draft PR. Run researchforge report build separately for the engineering report.

bashbash
# Create a clean local branch from the frozen baseline
researchforge ship branch [experiment_id] [--branch TEXT] [--yes] [--json]

# Build the engineering report
researchforge report build

# OPT-IN: push branch + open a DRAFT PR (requires gh CLI + 3-gate approval)
researchforge ship pr [experiment_id] [--yes] [--json]
⚠
ship pr only runs when: (1) shipping.allow_draft_pr: true in the approved contract, (2) gh CLI is authenticated, and (3) you type push at the confirmation prompt. Nothing is pushed without all three gates.

hub

The local hub shows all projects on your machine with their folder locations, status, and live activity. It starts automatically once the serve extra is installed.

bashbash
# Start the machine-wide hub dashboard (http://127.0.0.1:9000)
researchforge hub --background

# Start the per-project live monitor
researchforge serve --background
ℹ
Multi-user Hub controls (hub experiments, hub approve) are enterprise features — not available in the OSS CLI.

all install

bashbash
# Install both Claude Code skills and Cursor rules
researchforge all install [--user] [--global]

# --user: installs to ~/.claude/skills/ and ~/.cursor/rules/
# --global: installs to system-wide config (requires admin)

# Verify installation
researchforge all status

researchforge.yaml — complete reference

🚨
Copy this schema exactly — the contract uses extra="forbid" on every section. An invented key causes researchforge contract validate to reject the file outright.
researchforge.yamlyaml
# researchforge.yaml — complete reference (every key the contract accepts)
version: 1                      # required, literal 1

project:
  name: my-project
  mode: improve_repository      # improve_repository | explore_research_idea

objective:
  description: "Improve detection mAP@0.5 without exceeding the inference budget"
  primary_metric:
    name: map50
    direction: maximize         # maximize | minimize
    target_value: 0.85          # optional — autorun stops when reached
  hard_constraints:             # optional, repeatable
    - name: inference_ms
      operator: "<="            # <= | >= | < | > | ==
      value: 200
  secondary_metrics:            # optional — recorded, never enforced
    - inference_ms

repository:
  baseline_ref: main            # the ref the baseline commit is resolved from

execution:
  mode: auto                    # auto (default) | docker | venv
                                # auto prefers Docker when available
  trusted_repository: false
  setup_command: "pip install -r requirements.txt"
  screening_command: "python benchmarks/evaluate.py --quick"   # optional
  test_command: null            # optional — must pass before benchmark runs
  full_command: "python benchmarks/evaluate.py"                # required
  result_file: artifacts/results.json
  timeout_minutes: 20
  cpu_limit: 2
  memory_mb: 4096
  max_experiments: 8
  stall: 3                      # stop after N consecutive non-improvements

permissions:
  editable_paths:               # the only paths a patch may touch
    - src/
  protected_paths:              # patch touching these → rejected at import, never runs
    - benchmarks/
    - tests/

network:
  mode: none                    # none (default) | enabled
                                # use enabled for repos that download weights

secrets:
  forward_environment_variables: []   # nothing forwarded unless named here

validation:
  repeat_finalists: 3           # repeats to earn "validated"
  require_existing_tests: true

shipping:
  allow_branch_creation: true
  allow_draft_pr: false         # gate 1 of 3 for ship pr
ℹ
execution.mode: auto prefers Docker when a Dockerfile and daemon are present, falls back to venv. Set mode: venv explicitly if venv is what you want. network.mode only accepts none and enabled — it is a different field from execution.mode.

plan.yaml — experiment plan format

Generated by researchforge experiment plan <hyp-id> --synthesize. The importer forbids unknown keys — copy this schema exactly.

plan.yamlyaml
hypothesis_id: hyp-001          # required — must be a stored hypothesis
approach_summary: "Tune confidence and NMS thresholds"

experiments:
  # A) patch variant — a unified diff applied to the worktree
  - key: conf-low               # ^[a-z0-9][a-z0-9-]{0,40}$
    title: "Lower the confidence threshold"
    change_summary: "CONF 0.001 → 0.0001 in src/config.py"
    patch_file: patches/conf-low.patch
    expected_effect: improvement   # optional
    notes: "Grounded in paper-004"  # optional

  # B) env-only variant — no patch file needed
  - key: iou-065
    title: "Raise the NMS IoU threshold"
    change_summary: "IOU 0.6 → 0.65"
    env_overrides:              # injected into the experiment subprocess
      IOU: "0.65"

  # C) build on a measured ancestor (builds on exp-NNN already in the DB)
  - key: conf-plus-size
    title: "Confidence tuning at larger input size"
    change_summary: "Adds IMGSZ 800 on top of the conf-low winner"
    parent: conf-low            # key in this plan, or exp-NNN already measured
    patch_file: patches/conf-plus-size.patch

  # D) merge two independent winners
  - key: merge-001
    title: "Combine both winners"
    change_summary: "Composes conf-low and iou-065"
    parents: [conf-low, iou-065]
    # patch_file / env_overrides may be omitted for a pure merge
⚠
Set either patch_file or env_overrides, never both. Patches must live inside .researchforge/experiments/patches/. A patch touching a protected path is recorded as rejected at import and never runs. Use parent (single) or parents (list), not depends_on or requires_pass.

Environment variables

shellbash
# AI providers (built in — no extra install, auto-detected)
ANTHROPIC_API_KEY=sk-ant-...          # → claude-opus-4-5 (default)
GEMINI_API_KEY=...                    # → gemini-2.0-flash (default)
GOOGLE_API_KEY=...                    # alias for GEMINI_API_KEY
OPENAI_API_KEY=sk-...                 # → gpt-4o (default)
RESEARCHFORGE_LLM=claude-opus-4-5    # override the model for any provider
RESEARCHFORGE_LLM=http://localhost:11434/api   # or point at Ollama (air-gapped)

# Local behaviour
RESEARCHFORGE_HOME=~/.researchforge   # where machine-wide state (hub registry) lives
RESEARCHFORGE_NO_HUB=1               # opt out of the hub auto-starting
ℹ
Nothing reaches the network except research search (arXiv) and AI provider calls. Air-gapped machines simply skip those two steps and use papers import to supply literature.

Claude Code

After researchforge claude install, the following slash commands are available in any Claude Code session:

CommandWhat it does
/researchforge-startThe full loop from the top — search, baseline, hypotheses, run, ship
/researchforge-doctorCheck the install and the project's next step
/researchforge-papersSearch and manage the literature
/researchforge-landscapeWrite the research landscape from stored papers
/researchforge-hypothesesWrite and import hypotheses
/researchforge-planWrite plan.yaml + patches for a hypothesis
/researchforge-baselineFreeze the baseline
/researchforge-runRun an approved plan
/researchforge-resultsRead the lineage and results
/researchforge-validateRepeat-run the finalist to confirm stability
/researchforge-shipShip the validated winner as a clean branch
/researchforge-paperBuild the research package (BibTeX, outline, evidence matrix)

Cursor

After researchforge cursor install, use @researchforge-start in Cursor chat. The MDC rule instructs Cursor to follow the RF workflow automatically.

💡
You can use both Claude Code and Cursor simultaneously — they share the same .researchforge/ state directory, so experiments started in one IDE are visible in the other.

scikit-learn

train.pypython
import os
from sklearn.ensemble import GradientBoostingRegressor, RandomForestRegressor
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import mean_squared_error
import numpy as np

# ResearchForge injects these via worktree env vars
model_type = os.environ.get("MODEL_TYPE", "gbm")
n_estimators = int(os.environ.get("N_ESTIMATORS", "100"))
normalize = os.environ.get("NORMALIZE", "false").lower() == "true"

X_train, X_val, y_train, y_val = load_data()

if normalize:
    scaler = StandardScaler()
    X_train = scaler.fit_transform(X_train)
    X_val = scaler.transform(X_val)

if model_type == "rf":
    model = RandomForestRegressor(n_estimators=n_estimators, random_state=42)
else:
    model = GradientBoostingRegressor(n_estimators=n_estimators, random_state=42)

model.fit(X_train, y_train)
preds = model.predict(X_val)
rmse = np.sqrt(mean_squared_error(y_val, preds))

# Write results.json — NOT print(RF_METRIC)
import json, pathlib
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
    "schema_version": 1,
    "primary_metric": {"name": "rmse", "value": float(rmse)},
    "sample_count": len(y_val),
    "seed": 42,
}), encoding="utf-8")
plan.yaml (sklearn)yaml
hypothesis_id: hyp-001
approach_summary: "Test model type and normalisation"
experiments:
  - key: gbm-200
    title: "GBM 200 estimators"
    change_summary: "MODEL_TYPE=gbm N_ESTIMATORS=200"
    env_overrides: { MODEL_TYPE: gbm, N_ESTIMATORS: "200" }
  - key: rf-200
    title: "Random forest 200 estimators"
    change_summary: "MODEL_TYPE=rf N_ESTIMATORS=200"
    env_overrides: { MODEL_TYPE: rf, N_ESTIMATORS: "200" }
  - key: gbm-norm
    title: "GBM with normalisation"
    change_summary: "Adds NORMALIZE=true on top of gbm-200 winner"
    parent: gbm-200
    env_overrides: { NORMALIZE: "true" }

PyTorch / Lightning

train.pypython
import os
import torch
import pytorch_lightning as pl

lr = float(os.environ.get("LR", "1e-3"))
hidden = int(os.environ.get("HIDDEN_SIZE", "256"))
dropout = float(os.environ.get("DROPOUT", "0.1"))
use_batchnorm = os.environ.get("BATCHNORM", "false") == "true"

class MyModel(pl.LightningModule):
    def __init__(self):
        super().__init__()
        self.net = build_net(hidden, dropout, use_batchnorm)
        self.lr = lr

    def training_step(self, batch, idx):
        loss = self.net(batch)
        return loss

    def validation_step(self, batch, idx):
        val_loss = self.net(batch)
        # Emit to ResearchForge
        self.log("rf_val_loss", val_loss)
        return val_loss

trainer.fit(model, train_dl, val_dl)

# Write results.json from best checkpoint metrics
import json, pathlib
best_val = trainer.callback_metrics.get("val_loss", float("inf"))
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
    "schema_version": 1,
    "primary_metric": {"name": "val_loss", "value": float(best_val)},
    "sample_count": len(val_dl.dataset),
    "seed": 42,
}), encoding="utf-8")

HuggingFace Transformers

train.pypython
import os
from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer
from datasets import load_dataset
import numpy as np

model_name = os.environ.get("MODEL_NAME", "distilbert-base-uncased")
lr = float(os.environ.get("LR", "2e-5"))
epochs = int(os.environ.get("EPOCHS", "3"))
warmup = float(os.environ.get("WARMUP_RATIO", "0.1"))

model = AutoModelForSequenceClassification.from_pretrained(model_name)

args = TrainingArguments(
    output_dir="./out",
    learning_rate=lr,
    num_train_epochs=epochs,
    warmup_ratio=warmup,
    evaluation_strategy="epoch",
    save_strategy="no",
    load_best_model_at_end=False,
    report_to="none",  # disable wandb/mlflow — RF handles tracking
)

trainer = Trainer(model=model, args=args, ...)
trainer.train()
results = trainer.evaluate()

# Write results.json
import json, pathlib
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
    "schema_version": 1,
    "primary_metric": {"name": "f1", "value": results["eval_f1"]},
    "secondary_metrics": {"eval_loss": results["eval_loss"]},
    "sample_count": len(eval_dataset),
    "seed": 42,
}), encoding="utf-8")

Python environments

Each worktree gets an isolated venv cloned from the baseline environment. Every experiment starts from exactly the same dependency state.

bashbash
# Activate your env first, then run baseline:
source .venv/bin/activate
researchforge baseline run

# Worktrees are created at:
# .researchforge/worktrees/exp-001/   ← git worktree at baseline commit
# venv is cloned inside the worktree

# See all file locations:
researchforge paths

Docker execution

Set execution.mode: docker (or auto, which prefers Docker when available). ResearchForge uses your existing Dockerfile automatically. If you don't have one, generate it:

researchforge.yamlyaml
execution:
  mode: docker     # or auto (default — Docker when Dockerfile + daemon exist)
bashbash
# Generate a Dockerfile from the repo scan (no API key needed)
researchforge generate dockerfile

# AI-tailored version
researchforge generate dockerfile --provider anthropic

# GPU base image
researchforge generate dockerfile --cuda
ℹ
With Docker mode, each experiment gets its own container built from the same image. First run takes ~2 min to build (cached after). Best for any public repo you cloned — full process isolation, no dependency conflicts.

GitHub Actions / CI

GitHub Actions is an example of how to run the ResearchForge loop in CI, not a built-in product connector. The CLI still reads your benchmark output file and runs worktrees locally.

.github/workflows/research.ymlyaml
name: ResearchForge experiments

on:
  workflow_dispatch:
  push:
    branches: [research/**]

jobs:
  run-experiments:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }

      - name: Install ResearchForge
        run: pip install "researchforge[serve]"

      - name: Install project deps
        run: pip install -r requirements.txt

      - name: Freeze baseline
        run: researchforge baseline run

      - name: Run experiments
        run: researchforge experiment start .researchforge/experiments/plan.yaml --yes

Hub & local monitor

The hub and local monitor are first-class features in the ResearchForge CLI. They are local-only dashboards that help you inspect runs, project state, and experiment lineage.

bashbash
# Start the local monitoring server for this project
researchforge serve --background

# Start the machine-wide hub dashboard
researchforge hub --background

# Inspect project state and monitor status
researchforge status
researchforge paths

Stall & convergence

The stall parameter stops the run loop after N consecutive experiments that all fail to improve on the current best result. Set it in the contract (applies to every run) or override it per-run:

researchforge.yamlyaml
execution:
  stall: 3          # stop after 3 consecutive non-improvements
bashbash
# Override for one run
researchforge experiment run <plan-id> --stall 3

# In autorun
researchforge autorun --stall 2 --global-stall 3
exp-001 accuracy=0.891 best=0.860 Δ=+0.031 stall=0 → PASS
exp-002 accuracy=0.858 best=0.891 Δ=-0.033 stall=1 → REJECTED
exp-003 accuracy=0.862 best=0.891 Δ=-0.029 stall=2 → REJECTED
exp-004 accuracy=0.890 best=0.891 Δ=-0.001 stall=3 → STOP
⚠
Stall counts consecutive non-improvements from the current best, not the original baseline. It resets to 0 whenever a new best is found.

Isolation & worktrees

Experiments run one at a time, each in a fully isolated git worktree at the baseline commit with its own venv or Docker container. Parallel execution is not available in the current release.

process diagrambash
Orchestrator
│
├─ experiment exp-001  (.researchforge/worktrees/exp-001/)
│    env: MODEL_TYPE=gru, LR=1e-3
│    runs: python benchmarks/evaluate.py --subset full
│    writes: .researchforge/artifacts/exp-001/results.json
│    cleaned up after: ✓
│
├─ experiment exp-002  (.researchforge/worktrees/exp-002/)  ← runs next
│    ...
│
└─ experiment exp-003  (.researchforge/worktrees/exp-003/)  ← runs after
     ...
ℹ
If an experiment crashes, times out, or the setup fails, the orchestrator records it as FAIL or FAILED_SETUP, cleans up the worktree, and continues with the remaining experiments. Your main checkout is never touched. Use researchforge experiment resume run-001 to continue an interrupted run.

Protected paths

Protected paths are declared in permissions.protected_paths of the contract. A patch touching a protected path is recorded as rejected at import time — the experiment never executes at all. Use researchforge contract show to see what is currently protected.

researchforge.yamlyaml
permissions:
  editable_paths:        # AI may only touch these
    - src/
    - config/
  protected_paths:       # patch touching any of these → rejected at import
    - benchmarks/        # eval script
    - tests/             # test suite
ℹ
The approved contract itself is pinned by a SHA-256 of the YAML bytes at approval time. Editing the file after approval is detected and requires re-approval.

Security model

ResearchForge is a local-only CLI — no data leaves your machine. The security model has three properties: the approved contract is SHA-256 pinned so post-approval edits are detected; each experiment runs in a detached git worktree at the exact baseline commit so your checkout is never touched; and patches are inspected at import time, so a patch touching a protected path is rejected before it ever executes.

🔒
Contract pinning
The approved contract is pinned by a SHA-256 of its YAML bytes. Editing it after approval is detected and forces re-approval before anything runs.
🔒
Worktree isolation
Each experiment runs at the exact baseline commit in a detached worktree. Uncommitted changes in your checkout cannot leak into a run. This is isolation for trusted code — not a hostile-code sandbox; use Docker mode for repos you did not write.
🔒
Protected path enforcement
The patch is inspected at import time. A patch touching a protected path is recorded as rejected and never executes — there is no race window between checking and running.

Lineage & audit trail

The audit trail is derived from the project database — not a separate log file. There is no audit.log to tamper with, and it works on projects created before the feature existed.

bashbash
# View the trail (oldest first)
researchforge audit log
researchforge audit log --last 20
researchforge audit log --kind contract_approved   # filter by event kind

# Export as JSON (includes gate findings)
researchforge audit export audit.json

Event kinds: project_created, papers_searched, landscape_imported, hypothesis_reviewed, contract_approved, baseline_measured, plan_created, plan_approved, run_started, run_completed, benchmark_ran, experiment_decided, deliverable_created.

ℹ
audit export includes gate findings — plans that reached execution with no approval on record. That is the compliance-relevant part.

Enterprise add-ons: Hub setup

These capabilities are additive to the open-source CLI. The base ResearchForge product remains local-first and framework-agnostic; the Enterprise layer adds shared infrastructure, governance, and team coordination.

ℹ
The Enterprise Hub is not required for the OSS workflow. It is an optional shared control plane for teams that want a single dashboard, approvals, and multi-user coordination.

The Hub is a self-hosted server that aggregates team experiments, provides a shared dashboard, and exposes the approval queue. Runs as a Docker container inside your VPC.

bashbash
# Pull and start the hub
docker pull ghcr.io/forger-labs-hq/researchforge-hub:latest

docker run -d \
  --name rf-hub \
  -p 8080:8080 \
  -v /data/rf-hub:/data \
  -e RF_SECRET_KEY=$(openssl rand -hex 32) \
  -e RF_ADMIN_EMAIL=admin@yourcompany.com \
  ghcr.io/forger-labs-hq/researchforge-hub:latest

# Dashboard available at http://your-server:8080

The Hub dashboard shows: all team experiments with full lineage, live run status, metric history across days/weeks, approval queue for team lead review, and the full audit log export.

API key & workloads

~/.researchforgercbash
RESEARCHFORGE_HUB_URL=https://hub.yourcompany.com
RESEARCHFORGE_API_KEY=rf_live_xxxxxxxxxxxx
bash (per-run override)bash
RF_HUB_URL=https://hub.yourcompany.com \
RF_API_KEY=rf_live_xxxx \
researchforge run --workload search-ranking-v3

Workloads are stable project identifiers (for example a x-rf-workload header). All experiments tagged with the same workload are grouped together in the Hub dashboard for cross-run comparison.

CI in the OSS workflow

ResearchForge is a local CLI — there is no built-in CI connector. You can run it in GitHub Actions by treating it as a CLI step. The plan must already exist (written by Claude/Cursor locally and committed).

.github/workflows/experiments.ymlyaml
name: ResearchForge experiments

on:
  workflow_dispatch:
  push:
    branches: [research/**]

jobs:
  run-experiments:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install researchforge
      - run: pip install -r requirements.txt
      - run: researchforge contract approve --yes
      - run: researchforge baseline run
      - run: researchforge run .researchforge/experiments/plan.yaml --yes
      - run: researchforge results show run-001 --json

Air-gapped use

ResearchForge is local-only by design. Nothing reaches the network except research search (arXiv) and AI provider calls. On an air-gapped machine, skip those two steps and supply the literature manually:

bashbash
# On a connected machine: export pip wheels
pip download researchforge -d ./rf-wheels/
tar -czf rf-wheels.tar.gz rf-wheels/

# On a connected machine: export your paper set
researchforge papers export papers.json

# Transfer both to the air-gapped machine, then:
tar -xzf rf-wheels.tar.gz
pip install --no-index --find-links=./rf-wheels researchforge
researchforge papers import papers.json

# Use Ollama for AI calls locally:
export RESEARCHFORGE_LLM=http://localhost:11434/api

Multi-user coordination

Multi-machine experiment distribution and shared approval queues are enterprise roadmap features. In the current OSS release, share the project state by committing .researchforge/ to git — every team member's Claude Code or Cursor session sees the same papers, hypotheses, and results.

bashbash
# Share project state via git
git add .researchforge/
git commit -m "share research state"
git push

# Teammate pulls and continues:
git pull
researchforge status          # exact next step
researchforge results show run-001

Output artifacts

All project state lives in .researchforge/. The primary store is a SQLite database; loose files are the AI hand-off artifacts. Run researchforge paths to see every location.

directory structurebash
.researchforge/
├── researchforge.db          ← all state: papers, hypotheses, contract,
│                               baseline, plans, experiments, runs, decisions
├── config.json               ← project id + settings
├── synthesis/                ← landscape.yaml + hypotheses.yaml the AI writes
├── experiments/
│   ├── context.json          ← planning context exported for the AI
│   ├── plan.yaml             ← the plan being imported / last imported
│   └── patches/              ← one unified diff per variant
├── worktrees/                ← isolated checkouts at the baseline commit
├── artifacts/                ← per-execution stdout, stderr, diff, results.json
├── reports/                  ← engineering + research reports
├── research-log.md           ← living context autorun feeds back to the AI
└── autorun.json              ← loop state for autorun --resume
ℹ
There is no lineage.json or audit.log file — the DAG and audit trail are derived from the database and rendered by researchforge dashboard and researchforge audit log.

Troubleshooting

✗ artifacts/results.json not found after experiment run
→ Your eval script did not write artifacts/results.json. Check that it runs to completion and that pathlib.Path("artifacts/results.json").write_text(...) is called even on early stopping or error paths.
✗ Baseline already frozen — use baseline reset
→ You're trying to run baseline run when one already exists. Run researchforge baseline status to see what's frozen. Use researchforge baseline reset --confirm to clear it.
✗ Protected path violation: config/prod.yaml
→ An experiment modified a protected file. Check your experiment code for writes to that path. The experiment is already marked FAIL in the lineage.
✗ Worktree checkout failed: unstaged changes
→ Git can't create a worktree when there are unstaged changes to tracked files. Run git stash or git add + git commit before running experiments.
✗ venv clone failed: pip not found
→ ResearchForge couldn't find pip in the active venv. Make sure a venv is active (source .venv/bin/activate) before running baseline run.
✗ Hub connection refused
→ The Hub server isn't reachable at the URL in RESEARCHFORGE_HUB_URL. Check that the Docker container is running (docker ps | grep rf-hub) and the port is open.

Upgrading

bashbash
# Upgrade to latest
pip install --upgrade researchforge

# Re-register IDE skills/rules after upgrading
researchforge all install --user

# Confirm the installed version
pip show researchforge

# Verify git/Python/Docker are still usable
researchforge doctor
ResearchForge — Apache 2.0 · Made by Forger Labs HQ