
The research engine forAI teams that need proof
ResearchForge turns papers, experiments, and model iterations into a repeatable, auditable loop — from Claude Code or Cursor to validated, ship-ready results.
ResearchForge
From papers to proof in 90 seconds.
A concise walkthrough of the ResearchForge loop: search, test, validate, and ship with traceable evidence.
What it covers
- Research and hypothesis generation from arXiv and local code
- Baseline locking and isolated experiment worktrees
- Traceable lineage, rejection reasons, and validation proof
- Shipping the winner into a clean, reproducible branch
ResearchForge
Without it vs. With it.
Most ML teams run experiments the same messy way for years. ResearchForge changes that.
- No systematic way to survey the relevant literature
- No record of what failed or why
- Main branch polluted by broken experiments
- Baseline drifts — “improvement” is guesswork
- Good ideas stay unvalidated because testing is slow
- No lineage: can't reproduce last month's winner
- arXiv searched end-to-end — top papers ranked and stored automatically
- Full experiment lineage — every result, every rejection, every reason
- Git worktrees keep your checkout untouched, always
- Frozen baseline before anything runs — reproducible from day one
- Validate, ship, or reject hypotheses in minutes
- One-line proof: ship the winner as a clean branch + report
ResearchForge
How It Works
Three commands. Papers → hypotheses → validated results.
Search Literature
$ researchforge research searcharXiv searched → best papers ranked → stored
Form & Test Hypotheses
$ researchforge baseline runFrozen baseline. No guessing.
Validate & Ship
$ researchforge validatemean=0.90 stdev=0.0 n=2
Every experiment. Every rejection. All in the record.
The engine tracks the full lineage — so you always know why a result is what it is.
Run experiments in parallel isolation
Every experiment gets its own isolated git worktree at the baseline commit. Your checkout is never touched — no matter how many run at once.
ResearchForge
Built for the research stack you already trust.
Install once. Claude Code and Cursor. One consistent experiment loop.
Claude Code
Skills in ~/.claude/skills/
/researchforge-startCursor
Rules in ~/.cursor/rules/
@researchforge-start↕ or both at once
$ researchforge all install --userResearchForge
Grounded in the way research teams actually work.
ResearchForge sits on top of your repo, your paper sources, and your CI — without replacing the stack you already use.
* Enterprise tier — MLflow, W&B, CI/CD, GPU runners, Slack, custom adapters
Explore enterprise features →ResearchForge
What the research loop looks like in practice.
Two live competitions. Real results. Traceable experiments. No cherry-picking.
RMSE 15.2 → 6.1
ROGII Wellbore Geology Prediction
“The winning idea came from a paper I'd never have found manually. ResearchForge retrieved it, linked it to a hypothesis, ran it in isolation, and told me exactly why it beat the baseline.”
Rank #203 / ~2,500 teams
ARC-AGI-3 — Fluid Intelligence Benchmark
“Graph-frontier exploration (hyp-002) outperformed every hand-tuned probe approach. ResearchForge tracked the chain from base score 0.08 to 1.21 across 4 experiment variants, all with full lineage.”
ResearchForge
Why teams upgrade to Enterprise
When experiments affect product decisions, the audit trail, isolation, and lineage become operational requirements — not optional extras.
Air-gapped deployment
Run entirely within your VPC. No data leaves your network.
Full audit trail
Cryptographic contract — every change, every result, every decision logged.
Multi-user hub
Shared experiment dashboard across your whole ML team.
CI/CD integration
GitHub Actions + GitLab CI — experiments triggered on every PR.
MLflow / W&B bridge
Layer validated experiments on top of your existing tracking stack.
Custom LLM support
Use your private fine-tuned model, not just Claude or Cursor.
Cloud execution
AWS Batch / GCP Cloud Run — scale beyond local machines.
SOC2-ready
Protected path enforcement, compliance audit log, reproducibility on demand.
ResearchForge is free forever for individuals. Enterprise teams get air-gapped deployment, SSO, multi-user hub, and a dedicated support SLA.
ResearchForge
ResearchForge is free forever.
Individual researchers and open-source projects always get the full CLI — no nags, no limits.
- Unlimited local experiments
- Full experiment lineage
- Claude Code + Cursor skills
- Git worktree isolation
- Community support
- Multi-user hub dashboard
- CI/CD plugin (Actions / GitLab)
- Slack / Teams notifications
- Email support SLA
- SSO / SAML ready
- Air-gapped / VPC deployment
- Okta, Azure AD (SSO)
- MLflow / W&B integration
- SOC2 audit trail
- Dedicated support + SLA
MIT is for libraries.
ResearchForge is infrastructure.
Apache 2.0 means you can use it, modify it, and build on it — including commercially.