ResearchForge Documentation
Local-first AI research and benchmarking workflow for teams that need evidence, not guesses.
- • IDE-first workflow with Claude Code and Cursor
- • literature search and ranking
- • baseline, run, validate, and ship
- • local worktrees, protected paths, and audit log
- • Docker and local Python execution
- • self-hosted Hub and approval workflow
- • multi-user coordination across machines
- • air-gapped deployment
- • workload tagging and shared lineage
- • policy and governance layers for regulated teams
The shortest path to a real ResearchForge run
pip install "researchforge[serve]" researchforge all install --user # Open Claude Code or Cursor and run: /researchforge-start # or @researchforge-start
This is the recommended entry point. The detailed command reference below is for advanced workflows, CI/CD, and automation — not the default path most users should start with.
How ResearchForge works
ResearchForge implements a six-stage loop that converts a research question into a validated, shippable result with full lineage:
The key design principle: nothing moves until it has evidence. The baseline is immovable. Experiments run in isolation. The winner is only shipped after validation confirms it isn't a lucky seed.
Where does the eval script come from?
This depends on your project. ResearchForge handles all three cases:
protected_paths. No experiment can modify it. This is the guarantee that your benchmark stays stable across the entire experiment run.Requirements
ResearchForge requires Python 3.12+ and Git — nothing else. An IDE (Claude Code or Cursor) is recommended but optional: set any AI API key and use the standalone CLI path instead. No Node.js, no Docker (unless you want container isolation), no cloud account needed.
/researchforge-start), Cursor (@researchforge-start), or standalone with any AI API key (ANTHROPIC_API_KEY, GEMINI_API_KEY, or OPENAI_API_KEY). No IDE required for the standalone path — researchforge research synthesize calls the AI directly.| Requirement | Version | Why |
|---|---|---|
| Python | 3.12+ | The ResearchForge CLI and execution engine |
| Git | any recent | Worktree isolation — one worktree per experiment |
| Claude Code (optional) | latest | Recommended AI layer: reads papers, writes patches via /researchforge-* skills. |
| Cursor (optional) | latest | Alternative AI layer: same capabilities via @researchforge-start MDC rules. |
| API key (optional) | any | Standalone mode: set ANTHROPIC_API_KEY, GEMINI_API_KEY, or OPENAI_API_KEY — no IDE needed. |
Install
Install ResearchForge with a single pip command. The base package bundles all three AI provider SDKs (Anthropic, Google, OpenAI) so no extra install is needed for standalone mode. Add [serve] only if you want the live web monitor dashboard.
Requires Python 3.12+ and Git. No Node.js required.
Standard install
pip install researchforge # includes built-in AI providers (Anthropic, Gemini, OpenAI) pip install "researchforge[serve]" # add only if you want the live web monitor
Install with IDE integrations
# After pip install, register skills/rules: researchforge all install --user # Claude Code only: researchforge claude install # Cursor only: researchforge cursor install
Install from source
git clone https://github.com/forger-labs-hq/researchforge cd researchforge pip install -e ".[serve,dev]"
Docker (no Python on host)
docker run --rm -v "$PWD":/workspace -w /workspace \ ghcr.io/forger-labs-hq/researchforge:latest \ researchforge research search "your query"
/workspace.The IDE-first workflow
The recommended way to use ResearchForge is through Claude Code (/researchforge-start) or Cursor (@researchforge-start). The IDE handles all the creative steps — reading papers, forming hypotheses, writing patches — while the CLI handles the deterministic steps: freezing baselines, running experiments, validating winners, shipping. You type three things total: approve (contract), approve (plan), ship.
# Recommended — start in the IDE /researchforge-start # or @researchforge-start
Quickstart (2 minutes)
Install ResearchForge, open your IDE, and start the guided workflow. This is the shortest route for real usage.
# 1. Install pip install "researchforge[serve]" # 2. Open Claude Code or Cursor # Type one of these: /researchforge-start # or @researchforge-start # 3. Approve the contract, run the baseline, and let the agent do the rest
The IDE-first workflow
The intended way to use ResearchForge is through your IDE. Type one slash command or @mention and Claude Code / Cursor takes over: scans your repo, writes the eval script if needed, searches literature, generates hypotheses, runs experiments, and presents results — asking your approval at every consequential step. You approve; they execute.
The full loop — search to shipped branch
The complete research pipeline as it runs inside your IDE. Claude Code or Cursor drives every step — you only type your objective and approvals. Each dashboard panel below is presented inline in the chat, exactly as you’d see it in a real session.
Claude Code — full walkthrough
From first command to shipped branch. You type 5 things; Claude does the rest.
import json, pathlib, time
from src.classifier import Classifier
clf = Classifier()
latencies, correct = [], []
for item in load_test_data():
t0 = time.perf_counter()
correct.append(clf.predict(item["text"]) == item["label"])
latencies.append((time.perf_counter()-t0)*1000)
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
"schema_version": 1,
"primary_metric": {"name": "f1", "value": compute_f1(correct)},
"secondary_metrics": {
"p95_latency_ms": sorted(latencies)[int(len(latencies)*0.95)]
},
"sample_count": len(correct), "seed": 42
}))Cursor — full walkthrough
Same workflow via @researchforge-start. This repo already has a benchmark script — RF detects it automatically.
✨ Cross-IDE state sharing
This is one of ResearchForge's strongest features and almost always overlooked. Both Claude Code and Cursor read and write the exact same .researchforge/ directory. They share 100% of state — papers, hypotheses, baselines, experiment results, lineage — in real time.
What exactly is shared
| File / directory | What it contains | Both IDEs can |
|---|---|---|
| .researchforge/researchforge.db | All project state: papers, hypotheses, contract, baseline, plans, experiments | Read via CLI commands |
| .researchforge/synthesis/ | landscape.yaml + hypotheses.yaml the AI writes | Read, write (AI writes, CLI validates) |
| .researchforge/experiments/patches/ | One unified diff per experiment variant | Read, write (AI writes patches) |
| .researchforge/worktrees/ | Isolated git checkouts at the baseline commit | Read (managed by RF) |
| .researchforge/artifacts/ | Per-execution stdout, stderr, diff, results.json | Read results, interpret |
| .researchforge/reports/ | Engineering + research reports | Read, build with report build |
| .researchforge/research-log.md | Living context autorun feeds back to the AI | Read (autorun writes) |
Example — resume mid-loop in a different IDE
.researchforge/ to git and your whole team shares the research state. Every team member's Claude Code or Cursor session will see the same papers, hypotheses, and results — regardless of machine.Install IDE skills/rules
pip install "researchforge[serve]" researchforge all install --user # → ~/.claude/skills/ and ~/.cursor/rules/ researchforge all status
.researchforge/ state — start in one, continue in the other.Core concepts
How metrics are captured — the results.json contract
ResearchForge does not scan stdout. Your benchmark script writes a structured artifacts/results.json file after every run. ResearchForge reads that file to compare experiments against the baseline.
benchmarks/) and is never modified by the AI during experiments. The AI only patches your implementation code insrc/ or config/."""
Your benchmark script. Lives in a protected path.
ResearchForge runs this to measure the baseline, then runs it again
inside each experiment worktree (with the AI's patch applied to src/).
"""
import json
import pathlib
from my_model import load_and_eval # ← AI can patch this
# Run your evaluation
accuracy, p95_ms, cost = load_and_eval(dataset="benchmark-v2")
# Write results in the standard ResearchForge format
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
"schema_version": 1,
"primary_metric": {"name": "accuracy", "value": accuracy},
"secondary_metrics": {
"p95_latency_ms": p95_ms,
"average_cost_usd": cost,
},
"sample_count": 1200,
"seed": 42,
"metadata": {"dataset_version": "benchmark-v2"},
}), encoding="utf-8")
print("evaluation complete")What happens during an experiment
Multiple metrics & constraints
Your eval script can write as many secondary metrics as needed. Specify hard constraints to automatically reject experiments that trade too much quality for speed (or cost).
{
"schema_version": 1,
"primary_metric": {"name": "accuracy", "value": 0.891},
"secondary_metrics": {
"p95_latency_ms": 143.2,
"average_cost_usd": 0.0031,
"f1_macro": 0.877
},
"sample_count": 1200,
"seed": 42,
"metadata": {"dataset_version": "benchmark-v2", "model_params": 7340032}
}objective:
description: >
Improve accuracy on the classification benchmark while keeping
p95 latency under 200ms and cost under $0.005 per query.
primary_metric:
name: accuracy
direction: maximize
hard_constraints:
- name: p95_latency_ms
operator: <=
value: 200
- name: average_cost_usd
operator: <=
value: 0.005"Improve accuracy" → accuracy / maximize. "Reduce p95 latency below 200ms" → latency_ms / minimize. You can always edit the contract YAML manually afterward.Screening funnel
For slow full benchmarks, define a fast screening subset. Experiments must beat the baseline on the cheap screen before the expensive full eval runs.
execution: screening_command: python benchmarks/evaluate.py --subset screening full_command: python benchmarks/evaluate.py --subset full result_file: artifacts/results.json
import sys
import json, pathlib
subset = "screening" if "--subset" in sys.argv and "screening" in sys.argv else "full"
# screening = fast 10% sample; full = complete eval
dataset_size = 120 if subset == "screening" else 1200
accuracy = run_eval(n_samples=dataset_size)
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
"schema_version": 1,
"primary_metric": {"name": "accuracy", "value": accuracy},
"sample_count": dataset_size,
"seed": 42,
}), encoding="utf-8")REJECTED (screen) in the lineage — they never run the expensive full eval, saving significant compute on non-promising hypotheses.research search
Searches arXiv end-to-end, ranks results by relevance, stores top papers in the local knowledge base. The query is inferred from your objective by default; use --query to override.
researchforge research search [--query/-q TEXT]... [flags]
| Flag | Default | Description |
|---|---|---|
| --query / -q | from objective | Repeatable: add extra search terms (e.g. -q "YOLOv5 pruning" -q "knowledge distillation") |
| --select / --n | 20 | Max papers to store in the knowledge base |
| --min-score | 0.6 | Minimum relevance score (0–1.0) |
| --categories / -c | all | arXiv category filter, repeatable (e.g. --categories cs.CV --categories cs.LG) |
| --max-candidates | 200 | Fetch this many candidates before deduplication and ranking |
| --provider | none | AI provider (anthropic|google|openai) — generates domain-specific queries instead of keyword fallback |
| --force | false | Re-run search even if papers already exist |
| --json | false | Output results as JSON |
"Improve YOLOv5s mAP@0.5 on COCO128 object detection"). Generic objectives produce irrelevant papers. Use --provider anthropic to let the AI generate targeted queries.papers (manage knowledge base)
# List stored papers researchforge papers list # Show details for a specific paper researchforge papers show paper-003
baseline run
Runs your benchmark script once to freeze the current metric as the immovable reference. Must be run after researchforge contract approve. The objective, timeout, and metric come from the approved contract — not from flags.
researchforge baseline run [--check] [--json]
| Flag | Default | Description |
|---|---|---|
| --check | false | Dry-run: validate the contract and benchmark command without executing |
| --json | false | Output result as JSON |
# Check current baseline researchforge baseline show researchforge baseline status # alias for show
hypotheses
Hypotheses are written by your AI (Claude Code, Cursor, or direct API) and imported. hypotheses generate calls an AI provider directly — no IDE needed.
# Generate hypotheses using a direct AI API key (no IDE needed) researchforge hypotheses generate [--provider anthropic|google|openai] [--model TEXT] [--no-import] [--json] # Equivalent alias: researchforge research synthesize [--provider ...] [--model ...] [--no-import] [--json] # Import a hypotheses.yaml written by Claude/Cursor (validates schema) researchforge hypotheses import .researchforge/synthesis/hypotheses.yaml # List all hypotheses and their status researchforge hypotheses list # Show details for one hypothesis researchforge hypotheses show hyp-002
hypotheses generate requires ANTHROPIC_API_KEY, GEMINI_API_KEY, or OPENAI_API_KEY to be set. In the IDE workflow, Claude Code or Cursor writes hypotheses.yaml directly and you run hypotheses import to validate and store it.experiment (plan & manage)
There is no plan top-level command. Experiment plans are managed via researchforge experiment.
# Generate a plan template for a hypothesis (Claude/Cursor fills in the patches) researchforge experiment plan <hyp-id> # Import a plan.yaml (written by Claude/Cursor or by hand) researchforge experiment import plan.yaml # Start a run (import + approval + execute) researchforge experiment start plan.yaml researchforge run plan.yaml # alias # List experiments and status researchforge experiment list # Resume or discard an interrupted run researchforge experiment resume run-001 researchforge experiment abandon run-001 # Validate the contract (not the plan) researchforge contract validate
autorun — the autonomous loop
The headline feature: a fully autonomous research loop that runs overnight. Each round searches the experiment graph for the best node to expand, tries available hypotheses there, measures results, and re-synthesizes new ideas from what it found. Stops when it stalls, hits --target, or runs out of --max-hours.
--yes skips the first-batch prompt only — never the contract gate. Ctrl-C is safe and expected; autorun --resume continues with the same stall counter.# Start an overnight run researchforge serve --background # start live monitor first researchforge autorun --target 0.85 --max-hours 8 --yes # Short run to see how the loop behaves before committing overnight researchforge autorun --max-rounds 2 --observe # Continue an interrupted loop researchforge autorun --resume
| Flag | Default | Description |
|---|---|---|
| --stall INT | 2 | Stop a plan after N consecutive non-improvements |
| --global-stall INT | 3 | Stop the whole loop after N rounds with no improvement anywhere |
| --max-rounds INT | none | Hard cap on synthesis rounds |
| --max-hours FLOAT | none | Wall-clock limit (overnight safety cap) |
| --target FLOAT | contract's target_value | Stop as soon as the primary metric reaches this |
| --compound / --no-compound | on | Build each round on a node of the graph instead of on the baseline |
| --explore FLOAT | 0.0 | UCB1 constant — 0 always expands the current best; higher revisits under-explored branches |
| --merge / --no-merge | off | Try combining two independent winners each round |
| --observe / --no-observe | off | AI reads each run's logs and records a paragraph on what it showed |
| --resynthesize / --no-resynthesize | on | Generate new hypotheses from measured results each round |
| -p / --provider TEXT | auto | anthropic | google | openai |
| -y / --yes | off | Unattended — skip the first-batch approval prompt |
| --resume | — | Continue the interrupted loop from .researchforge/autorun.json |
--max-rounds 2 --observe to watch two rounds end to end, read the results, and confirm the loop is reasoning about your benchmark before running overnight.run
Alias for researchforge experiment start. Imports the plan, asks for one typed approval, then executes all experiments in isolated git worktrees. Stall, parallel, timeout, and threshold come from the approved contract — not from flags.
researchforge run <plan.yaml> [--yes] [--monitor/--no-monitor] [--json]
| Flag | Default | Description |
|---|---|---|
| --yes | false | Skip the typed approval prompt (for CI/CD) |
| --monitor / --no-monitor | auto | Start/skip the live web monitor |
| --json | false | Output progress as JSON lines |
--parallel, --stall, --worker, and --tags are not flags on this command — stall and parallel are set in researchforge.yaml; worker/tags are enterprise roadmap features.validate
Re-runs the winner N times (from the contract's validation.repeat_finalists) to confirm the result is stable, not a lucky seed.
researchforge validate <run-id> [--experiment/-e TEXT] [--yes] [--json]
| Flag | Default | Description |
|---|---|---|
| run-id | — | Required: the run to validate (e.g. run-001) |
| --experiment / -e | best in run | Validate a specific experiment ID instead of the best |
| --yes | false | Skip confirmation prompt |
| --json | false | Output as JSON |
validation.repeat_finalists in the approved contract, not from flags.ship
ship is a subcommand group, not a single command. Use ship branch for a local clean branch, then optionally ship pr to open a draft PR. Run researchforge report build separately for the engineering report.
# Create a clean local branch from the frozen baseline researchforge ship branch [experiment_id] [--branch TEXT] [--yes] [--json] # Build the engineering report researchforge report build # OPT-IN: push branch + open a DRAFT PR (requires gh CLI + 3-gate approval) researchforge ship pr [experiment_id] [--yes] [--json]
ship pr only runs when: (1) shipping.allow_draft_pr: true in the approved contract, (2) gh CLI is authenticated, and (3) you type push at the confirmation prompt. Nothing is pushed without all three gates.hub
The local hub shows all projects on your machine with their folder locations, status, and live activity. It starts automatically once the serve extra is installed.
# Start the machine-wide hub dashboard (http://127.0.0.1:9000) researchforge hub --background # Start the per-project live monitor researchforge serve --background
hub experiments, hub approve) are enterprise features — not available in the OSS CLI.all install
# Install both Claude Code skills and Cursor rules researchforge all install [--user] [--global] # --user: installs to ~/.claude/skills/ and ~/.cursor/rules/ # --global: installs to system-wide config (requires admin) # Verify installation researchforge all status
researchforge.yaml — complete reference
extra="forbid" on every section. An invented key causes researchforge contract validate to reject the file outright.# researchforge.yaml — complete reference (every key the contract accepts)
version: 1 # required, literal 1
project:
name: my-project
mode: improve_repository # improve_repository | explore_research_idea
objective:
description: "Improve detection mAP@0.5 without exceeding the inference budget"
primary_metric:
name: map50
direction: maximize # maximize | minimize
target_value: 0.85 # optional — autorun stops when reached
hard_constraints: # optional, repeatable
- name: inference_ms
operator: "<=" # <= | >= | < | > | ==
value: 200
secondary_metrics: # optional — recorded, never enforced
- inference_ms
repository:
baseline_ref: main # the ref the baseline commit is resolved from
execution:
mode: auto # auto (default) | docker | venv
# auto prefers Docker when available
trusted_repository: false
setup_command: "pip install -r requirements.txt"
screening_command: "python benchmarks/evaluate.py --quick" # optional
test_command: null # optional — must pass before benchmark runs
full_command: "python benchmarks/evaluate.py" # required
result_file: artifacts/results.json
timeout_minutes: 20
cpu_limit: 2
memory_mb: 4096
max_experiments: 8
stall: 3 # stop after N consecutive non-improvements
permissions:
editable_paths: # the only paths a patch may touch
- src/
protected_paths: # patch touching these → rejected at import, never runs
- benchmarks/
- tests/
network:
mode: none # none (default) | enabled
# use enabled for repos that download weights
secrets:
forward_environment_variables: [] # nothing forwarded unless named here
validation:
repeat_finalists: 3 # repeats to earn "validated"
require_existing_tests: true
shipping:
allow_branch_creation: true
allow_draft_pr: false # gate 1 of 3 for ship prexecution.mode: auto prefers Docker when a Dockerfile and daemon are present, falls back to venv. Set mode: venv explicitly if venv is what you want. network.mode only accepts none and enabled — it is a different field from execution.mode.plan.yaml — experiment plan format
Generated by researchforge experiment plan <hyp-id> --synthesize. The importer forbids unknown keys — copy this schema exactly.
hypothesis_id: hyp-001 # required — must be a stored hypothesis
approach_summary: "Tune confidence and NMS thresholds"
experiments:
# A) patch variant — a unified diff applied to the worktree
- key: conf-low # ^[a-z0-9][a-z0-9-]{0,40}$
title: "Lower the confidence threshold"
change_summary: "CONF 0.001 → 0.0001 in src/config.py"
patch_file: patches/conf-low.patch
expected_effect: improvement # optional
notes: "Grounded in paper-004" # optional
# B) env-only variant — no patch file needed
- key: iou-065
title: "Raise the NMS IoU threshold"
change_summary: "IOU 0.6 → 0.65"
env_overrides: # injected into the experiment subprocess
IOU: "0.65"
# C) build on a measured ancestor (builds on exp-NNN already in the DB)
- key: conf-plus-size
title: "Confidence tuning at larger input size"
change_summary: "Adds IMGSZ 800 on top of the conf-low winner"
parent: conf-low # key in this plan, or exp-NNN already measured
patch_file: patches/conf-plus-size.patch
# D) merge two independent winners
- key: merge-001
title: "Combine both winners"
change_summary: "Composes conf-low and iou-065"
parents: [conf-low, iou-065]
# patch_file / env_overrides may be omitted for a pure mergepatch_file or env_overrides, never both. Patches must live inside .researchforge/experiments/patches/. A patch touching a protected path is recorded as rejected at import and never runs. Use parent (single) or parents (list), not depends_on or requires_pass.Environment variables
# AI providers (built in — no extra install, auto-detected) ANTHROPIC_API_KEY=sk-ant-... # → claude-opus-4-5 (default) GEMINI_API_KEY=... # → gemini-2.0-flash (default) GOOGLE_API_KEY=... # alias for GEMINI_API_KEY OPENAI_API_KEY=sk-... # → gpt-4o (default) RESEARCHFORGE_LLM=claude-opus-4-5 # override the model for any provider RESEARCHFORGE_LLM=http://localhost:11434/api # or point at Ollama (air-gapped) # Local behaviour RESEARCHFORGE_HOME=~/.researchforge # where machine-wide state (hub registry) lives RESEARCHFORGE_NO_HUB=1 # opt out of the hub auto-starting
research search (arXiv) and AI provider calls. Air-gapped machines simply skip those two steps and use papers import to supply literature.Claude Code
After researchforge claude install, the following slash commands are available in any Claude Code session:
| Command | What it does |
|---|---|
| /researchforge-start | The full loop from the top — search, baseline, hypotheses, run, ship |
| /researchforge-doctor | Check the install and the project's next step |
| /researchforge-papers | Search and manage the literature |
| /researchforge-landscape | Write the research landscape from stored papers |
| /researchforge-hypotheses | Write and import hypotheses |
| /researchforge-plan | Write plan.yaml + patches for a hypothesis |
| /researchforge-baseline | Freeze the baseline |
| /researchforge-run | Run an approved plan |
| /researchforge-results | Read the lineage and results |
| /researchforge-validate | Repeat-run the finalist to confirm stability |
| /researchforge-ship | Ship the validated winner as a clean branch |
| /researchforge-paper | Build the research package (BibTeX, outline, evidence matrix) |
Cursor
After researchforge cursor install, use @researchforge-start in Cursor chat. The MDC rule instructs Cursor to follow the RF workflow automatically.
.researchforge/ state directory, so experiments started in one IDE are visible in the other.scikit-learn
import os
from sklearn.ensemble import GradientBoostingRegressor, RandomForestRegressor
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import mean_squared_error
import numpy as np
# ResearchForge injects these via worktree env vars
model_type = os.environ.get("MODEL_TYPE", "gbm")
n_estimators = int(os.environ.get("N_ESTIMATORS", "100"))
normalize = os.environ.get("NORMALIZE", "false").lower() == "true"
X_train, X_val, y_train, y_val = load_data()
if normalize:
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_val = scaler.transform(X_val)
if model_type == "rf":
model = RandomForestRegressor(n_estimators=n_estimators, random_state=42)
else:
model = GradientBoostingRegressor(n_estimators=n_estimators, random_state=42)
model.fit(X_train, y_train)
preds = model.predict(X_val)
rmse = np.sqrt(mean_squared_error(y_val, preds))
# Write results.json — NOT print(RF_METRIC)
import json, pathlib
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
"schema_version": 1,
"primary_metric": {"name": "rmse", "value": float(rmse)},
"sample_count": len(y_val),
"seed": 42,
}), encoding="utf-8")hypothesis_id: hyp-001
approach_summary: "Test model type and normalisation"
experiments:
- key: gbm-200
title: "GBM 200 estimators"
change_summary: "MODEL_TYPE=gbm N_ESTIMATORS=200"
env_overrides: { MODEL_TYPE: gbm, N_ESTIMATORS: "200" }
- key: rf-200
title: "Random forest 200 estimators"
change_summary: "MODEL_TYPE=rf N_ESTIMATORS=200"
env_overrides: { MODEL_TYPE: rf, N_ESTIMATORS: "200" }
- key: gbm-norm
title: "GBM with normalisation"
change_summary: "Adds NORMALIZE=true on top of gbm-200 winner"
parent: gbm-200
env_overrides: { NORMALIZE: "true" }PyTorch / Lightning
import os
import torch
import pytorch_lightning as pl
lr = float(os.environ.get("LR", "1e-3"))
hidden = int(os.environ.get("HIDDEN_SIZE", "256"))
dropout = float(os.environ.get("DROPOUT", "0.1"))
use_batchnorm = os.environ.get("BATCHNORM", "false") == "true"
class MyModel(pl.LightningModule):
def __init__(self):
super().__init__()
self.net = build_net(hidden, dropout, use_batchnorm)
self.lr = lr
def training_step(self, batch, idx):
loss = self.net(batch)
return loss
def validation_step(self, batch, idx):
val_loss = self.net(batch)
# Emit to ResearchForge
self.log("rf_val_loss", val_loss)
return val_loss
trainer.fit(model, train_dl, val_dl)
# Write results.json from best checkpoint metrics
import json, pathlib
best_val = trainer.callback_metrics.get("val_loss", float("inf"))
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
"schema_version": 1,
"primary_metric": {"name": "val_loss", "value": float(best_val)},
"sample_count": len(val_dl.dataset),
"seed": 42,
}), encoding="utf-8")HuggingFace Transformers
import os
from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer
from datasets import load_dataset
import numpy as np
model_name = os.environ.get("MODEL_NAME", "distilbert-base-uncased")
lr = float(os.environ.get("LR", "2e-5"))
epochs = int(os.environ.get("EPOCHS", "3"))
warmup = float(os.environ.get("WARMUP_RATIO", "0.1"))
model = AutoModelForSequenceClassification.from_pretrained(model_name)
args = TrainingArguments(
output_dir="./out",
learning_rate=lr,
num_train_epochs=epochs,
warmup_ratio=warmup,
evaluation_strategy="epoch",
save_strategy="no",
load_best_model_at_end=False,
report_to="none", # disable wandb/mlflow — RF handles tracking
)
trainer = Trainer(model=model, args=args, ...)
trainer.train()
results = trainer.evaluate()
# Write results.json
import json, pathlib
pathlib.Path("artifacts").mkdir(exist_ok=True)
pathlib.Path("artifacts/results.json").write_text(json.dumps({
"schema_version": 1,
"primary_metric": {"name": "f1", "value": results["eval_f1"]},
"secondary_metrics": {"eval_loss": results["eval_loss"]},
"sample_count": len(eval_dataset),
"seed": 42,
}), encoding="utf-8")Python environments
Each worktree gets an isolated venv cloned from the baseline environment. Every experiment starts from exactly the same dependency state.
# Activate your env first, then run baseline: source .venv/bin/activate researchforge baseline run # Worktrees are created at: # .researchforge/worktrees/exp-001/ ← git worktree at baseline commit # venv is cloned inside the worktree # See all file locations: researchforge paths
Docker execution
Set execution.mode: docker (or auto, which prefers Docker when available). ResearchForge uses your existing Dockerfile automatically. If you don't have one, generate it:
execution: mode: docker # or auto (default — Docker when Dockerfile + daemon exist)
# Generate a Dockerfile from the repo scan (no API key needed) researchforge generate dockerfile # AI-tailored version researchforge generate dockerfile --provider anthropic # GPU base image researchforge generate dockerfile --cuda
GitHub Actions / CI
GitHub Actions is an example of how to run the ResearchForge loop in CI, not a built-in product connector. The CLI still reads your benchmark output file and runs worktrees locally.
name: ResearchForge experiments
on:
workflow_dispatch:
push:
branches: [research/**]
jobs:
run-experiments:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- name: Install ResearchForge
run: pip install "researchforge[serve]"
- name: Install project deps
run: pip install -r requirements.txt
- name: Freeze baseline
run: researchforge baseline run
- name: Run experiments
run: researchforge experiment start .researchforge/experiments/plan.yaml --yesHub & local monitor
The hub and local monitor are first-class features in the ResearchForge CLI. They are local-only dashboards that help you inspect runs, project state, and experiment lineage.
# Start the local monitoring server for this project researchforge serve --background # Start the machine-wide hub dashboard researchforge hub --background # Inspect project state and monitor status researchforge status researchforge paths
Stall & convergence
The stall parameter stops the run loop after N consecutive experiments that all fail to improve on the current best result. Set it in the contract (applies to every run) or override it per-run:
execution: stall: 3 # stop after 3 consecutive non-improvements
# Override for one run researchforge experiment run <plan-id> --stall 3 # In autorun researchforge autorun --stall 2 --global-stall 3
Isolation & worktrees
Experiments run one at a time, each in a fully isolated git worktree at the baseline commit with its own venv or Docker container. Parallel execution is not available in the current release.
Orchestrator
│
├─ experiment exp-001 (.researchforge/worktrees/exp-001/)
│ env: MODEL_TYPE=gru, LR=1e-3
│ runs: python benchmarks/evaluate.py --subset full
│ writes: .researchforge/artifacts/exp-001/results.json
│ cleaned up after: ✓
│
├─ experiment exp-002 (.researchforge/worktrees/exp-002/) ← runs next
│ ...
│
└─ experiment exp-003 (.researchforge/worktrees/exp-003/) ← runs after
...researchforge experiment resume run-001 to continue an interrupted run.Protected paths
Protected paths are declared in permissions.protected_paths of the contract. A patch touching a protected path is recorded as rejected at import time — the experiment never executes at all. Use researchforge contract show to see what is currently protected.
permissions:
editable_paths: # AI may only touch these
- src/
- config/
protected_paths: # patch touching any of these → rejected at import
- benchmarks/ # eval script
- tests/ # test suiteSecurity model
ResearchForge is a local-only CLI — no data leaves your machine. The security model has three properties: the approved contract is SHA-256 pinned so post-approval edits are detected; each experiment runs in a detached git worktree at the exact baseline commit so your checkout is never touched; and patches are inspected at import time, so a patch touching a protected path is rejected before it ever executes.
Lineage & audit trail
The audit trail is derived from the project database — not a separate log file. There is no audit.log to tamper with, and it works on projects created before the feature existed.
# View the trail (oldest first) researchforge audit log researchforge audit log --last 20 researchforge audit log --kind contract_approved # filter by event kind # Export as JSON (includes gate findings) researchforge audit export audit.json
Event kinds: project_created, papers_searched, landscape_imported, hypothesis_reviewed, contract_approved, baseline_measured, plan_created, plan_approved, run_started, run_completed, benchmark_ran, experiment_decided, deliverable_created.
audit export includes gate findings — plans that reached execution with no approval on record. That is the compliance-relevant part.Enterprise add-ons: Hub setup
These capabilities are additive to the open-source CLI. The base ResearchForge product remains local-first and framework-agnostic; the Enterprise layer adds shared infrastructure, governance, and team coordination.
The Hub is a self-hosted server that aggregates team experiments, provides a shared dashboard, and exposes the approval queue. Runs as a Docker container inside your VPC.
# Pull and start the hub docker pull ghcr.io/forger-labs-hq/researchforge-hub:latest docker run -d \ --name rf-hub \ -p 8080:8080 \ -v /data/rf-hub:/data \ -e RF_SECRET_KEY=$(openssl rand -hex 32) \ -e RF_ADMIN_EMAIL=admin@yourcompany.com \ ghcr.io/forger-labs-hq/researchforge-hub:latest # Dashboard available at http://your-server:8080
The Hub dashboard shows: all team experiments with full lineage, live run status, metric history across days/weeks, approval queue for team lead review, and the full audit log export.
API key & workloads
RESEARCHFORGE_HUB_URL=https://hub.yourcompany.com RESEARCHFORGE_API_KEY=rf_live_xxxxxxxxxxxx
RF_HUB_URL=https://hub.yourcompany.com \ RF_API_KEY=rf_live_xxxx \ researchforge run --workload search-ranking-v3
Workloads are stable project identifiers (for example a x-rf-workload header). All experiments tagged with the same workload are grouped together in the Hub dashboard for cross-run comparison.
CI in the OSS workflow
ResearchForge is a local CLI — there is no built-in CI connector. You can run it in GitHub Actions by treating it as a CLI step. The plan must already exist (written by Claude/Cursor locally and committed).
name: ResearchForge experiments
on:
workflow_dispatch:
push:
branches: [research/**]
jobs:
run-experiments:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install researchforge
- run: pip install -r requirements.txt
- run: researchforge contract approve --yes
- run: researchforge baseline run
- run: researchforge run .researchforge/experiments/plan.yaml --yes
- run: researchforge results show run-001 --jsonAir-gapped use
ResearchForge is local-only by design. Nothing reaches the network except research search (arXiv) and AI provider calls. On an air-gapped machine, skip those two steps and supply the literature manually:
# On a connected machine: export pip wheels pip download researchforge -d ./rf-wheels/ tar -czf rf-wheels.tar.gz rf-wheels/ # On a connected machine: export your paper set researchforge papers export papers.json # Transfer both to the air-gapped machine, then: tar -xzf rf-wheels.tar.gz pip install --no-index --find-links=./rf-wheels researchforge researchforge papers import papers.json # Use Ollama for AI calls locally: export RESEARCHFORGE_LLM=http://localhost:11434/api
Multi-user coordination
Multi-machine experiment distribution and shared approval queues are enterprise roadmap features. In the current OSS release, share the project state by committing .researchforge/ to git — every team member's Claude Code or Cursor session sees the same papers, hypotheses, and results.
# Share project state via git git add .researchforge/ git commit -m "share research state" git push # Teammate pulls and continues: git pull researchforge status # exact next step researchforge results show run-001
Output artifacts
All project state lives in .researchforge/. The primary store is a SQLite database; loose files are the AI hand-off artifacts. Run researchforge paths to see every location.
.researchforge/ ├── researchforge.db ← all state: papers, hypotheses, contract, │ baseline, plans, experiments, runs, decisions ├── config.json ← project id + settings ├── synthesis/ ← landscape.yaml + hypotheses.yaml the AI writes ├── experiments/ │ ├── context.json ← planning context exported for the AI │ ├── plan.yaml ← the plan being imported / last imported │ └── patches/ ← one unified diff per variant ├── worktrees/ ← isolated checkouts at the baseline commit ├── artifacts/ ← per-execution stdout, stderr, diff, results.json ├── reports/ ← engineering + research reports ├── research-log.md ← living context autorun feeds back to the AI └── autorun.json ← loop state for autorun --resume
lineage.json or audit.log file — the DAG and audit trail are derived from the database and rendered by researchforge dashboard and researchforge audit log.Troubleshooting
Upgrading
# Upgrade to latest pip install --upgrade researchforge # Re-register IDE skills/rules after upgrading researchforge all install --user # Confirm the installed version pip show researchforge # Verify git/Python/Docker are still usable researchforge doctor