Skip to content

Deploy GPT Researcher v3.7.0: Open Source Research AI

Install GPT Researcher v3.7.0 with Docker or pip. Specs, copy-paste commands, Jev defaults, verification steps, and fixes for WeasyPrint, empty /my-docs mounts, and :8000 404s.

8 min readIntermediate

Two ways to stand up GPT Researcher as an open source research AI: drop the pip package into your own code, or deploy the full FastAPI + frontend stack. The package wins when you only need the agent class inside an existing service. For a research workstation that ships cited 2,000+ word reports, Docker (or a clean git clone) wins – UI, Deep Research tree mode, and Jev filtering come wired.

v3.7.0 (GitHub release 26 Sep 2026) is the version this guide pins. Planner breaks the brief into sub-questions, scrapers hit 20+ sources in parallel, context is scored for usefulness instead of embeddings-only, then Markdown/PDF/DOCX drops with inline citations. Commands below target that tag only – not whatever a two-year-old blog still ranks for.

Research tooling ages badly when install posts freeze on old runtimes. If you have ever followed a “works on 3.10” README after a hard Python floor bump, you already know the evening you are trying to skip here.

System requirements for GPT Researcher v3.7.0

Linux deployment notes list a floor that actually completes light web runs: 1 vCPU, 2 GB RAM, 50 GB SSD, stable outbound access for APIs and scrapes. As of v3.7.0, Python 3.12+ is mandatory – 3.10/3.11 envs fail at install or import after a fresh pull.

Recommended when you run Deep Research often: 2-4 vCPU, 4-8 GB RAM, 50-100 GB SSD. Concurrent scrapers and long synthesis push RAM first; local Ollama-class models jump you into 16-32 GB+ and prefer GPU. Docker images also stack Chromium layers – disk climbs quietly in outputs/.

Tier CPU RAM Disk Notes
Minimum 1 vCPU 2 GB 50 GB Single research_report only
Recommended 2-4 vCPU 4-8 GB 50-100 GB UI + Deep Research
Local LLM 4+ / GPU 16-32 GB+ 100 GB+ Ollama etc.

Undersized boxes do not fail politely. Parallel scrapers stall; long context synthesis OOMs mid-report. API spend is a separate budget line – default quality path expects Tavily + OpenAI keys (or pluggable equivalents).

Official download and sources

Primary source is the GitHub repo. Skip random mirrors.

Clone when you want full control; otherwise follow the published Docker flow. Pin the tag if you need bit-for-bit reruns later.

Install open source research AI: recommended Docker path

Docker keeps Chromium, WeasyPrint system libs, and the Node UI off your host. That is why it is the default path here. Pip-only and bare git+venv sit after for the cases they actually fit.

  1. Install Docker Desktop (or engine + compose) from Docker’s site.
  2. Clone and prepare env:
    git clone https://github.com/assafelovic/gpt-researcher.git
    cd gpt-researcher
    cp .env.example .env
    # Edit .env: OPENAI_API_KEY=sk-... and TAVILY_API_KEY=tvly-...
    # Optional: OPENAI_BASE_URL, RETRIEVER, FAST_LLM/SMART_LLM, DOC_PATH
  3. Build and run:
    docker compose up --build
    # or: docker-compose up --build
  4. Backend: http://localhost:8000 – React UI: http://localhost:3000 (compose defaults). Open the UI and fire a short query.

CLI-only container (no compose UI):

docker run -it --name gpt-researcher -p 8000:8000 --env-file .env 
 -v /absolute/path/to/your/docs:/my-docs gpt-researcher

The volume is not optional for local/hybrid docs. Skip -v ...:/my-docs and the container only sees an empty baked-in directory – research “succeeds” while your files never appear.

Pro tip: v3.7.0 defaults the context filter to Jev, with keyword BM25 as the no-key fallback. Standard runs no longer need an embeddings provider. Add TYPESAFE_API_KEY for the full Jev scorer; CONTEXT_FILTER=embeddings restores the older path. Docs homepage cites usefulness scoring around 73% for Jev vs ~51% keyword vs ~46% embeddings on their benchmark – your corpus will differ.

Pip library only (agent class, no UI):

python3.12 -m venv .venv && source .venv/bin/activate # Windows: .venvScriptsactivate
pip install -U gpt-researcher
export OPENAI_API_KEY=... TAVILY_API_KEY=...

Then from gpt_researcher import GPTResearcher and the async conduct_research / write_report path.

Bare metal full app: same clone, then python -m venv env && source env/bin/activate && pip install -r requirements.txt && python -m uvicorn main:app --reload. Host libraries hurt more here – jump to the error section when PDF export or scrapers complain.

First-time configuration

Minimum .env (or exports):

OPENAI_API_KEY=sk-your-key
TAVILY_API_KEY=tvly-your-key
# Optional
# RETRIEVER=tavily
# FAST_LLM=openai:gpt-4o-mini
# SMART_LLM=openai:gpt-4.1
# REPORT_SOURCE=web
# DOC_PATH=./my-docs
# DEEP_RESEARCH_BREADTH=4
# DEEP_RESEARCH_DEPTH=2

Hybrid mode is the quiet advantage. Drop PDFs/DOCX/CSV into my-docs (or the mounted volume), set report_source to local or hybrid, and the same run can cite internal files beside the live web. Other retrievers (Bing, SearXNG, arXiv, …) swap with RETRIEVER= plus their keys. LLM swaps follow the project llms docs – keep output token ceilings high on long-report models.

Verify the install works

1. UI loads on :3000 (compose) or the backend answers on :8000 without a blank 404 page.
2. Run a short query such as “EU CRA obligations for open-source maintainers 2026”. You want a plan step, parallel searches, then a multi-section report with citations.
3. On the pip API, call the researcher cost/source helpers after a run.
4. Confirm version: git describe --tags on a clone, or package metadata after pip – should line up with GitHub v3.7.0.

Finished in a few minutes, more than ~10 sources, export buttons alive? Ship it.

Common install errors and fixes

cannot load library ‘gobject-2.0-0’ / pango (WeasyPrint) – Research finishes; PDF/DOCX export dies. macOS: brew install pango glib gobject-introspection, and on Apple silicon often [email protected] from brew with pip3.12. Linux: sudo apt install libglib2.0-dev libpango-1.0-0. Match WeasyPrint’s own first-steps guide for your OS. This is the #1 host-library trap on bare metal.

localhost:8000 returns 404 after fresh clone or upgrade – Frontend route or static assets missing (lightweight vs full UI paths split across recent releases; see discussions around GitHub issue #1510-class failures). Prefer docker compose so the UI lands on :3000. Bare metal: make sure the full tree is present and any Next/static build step actually ran.

Docker empty reports, “no module named tavily”, or 0 pages scraped – Stale image or scraper settings. Rebuild from the current Dockerfile, check that tavily-python is in the image, set LOGGING_LEVEL=DEBUG, fix volume permissions (Windows bind mounts especially), pull current tags.

Model access / 403 / “does not exist” – The exact FAST_LLM/SMART_LLM string is not enabled on your key. Point both at models the project can call, or switch Azure/Ollama paths.

Chrome/chromedriver mismatch (Selenium scraper) – Align Chrome to a driver build, or set SCRAPER=bs and accept simpler HTML extraction.

Upgrade and uninstall

Pip path: pip install -U gpt-researcher. Full stack: git fetch && git checkout v3.7.0 (or main), recreate the venv, or docker compose build --no-cache && docker compose up.

The catch is .env drift – new context-filter vars and the Python 3.12 floor showed up with 3.7.x. Most configs need no migration script; Deep Research and multi-agent paths simply got sharper across 3.x.

Uninstall pip: pip uninstall gpt-researcher. Clone: deactivate, rm -rf gpt-researcher. Docker: docker compose down -v, docker rmi the images, drop named volumes and host mounts. Wipe outputs/ and .env when keys must leave the disk. Nothing registers as a system service by default.

Is a full workstation overkill if you only need three reports a month? Sometimes. The first time hybrid local+web citations spare you a rewrite on an internal brief, the UI stops feeling optional.

FAQ

Does GPT Researcher v3.7.0 still need embeddings?

No. Default is Jev (TYPESAFE_API_KEY) or local keyword ranking. Embeddings only if you set CONTEXT_FILTER=embeddings.

Pip package or Docker – which should I pick for a team demo?

Docker compose. One command gives port 3000 UI, pinned Chromium, and isolated system libs – nobody debugs pango on a shared laptop five minutes before the meeting. Pip fits when you already own the front end and only import GPTResearcher inside FastAPI or LangGraph.

Why did my Deep Research report cost more or time out?

About five minutes and ~$0.40 per run with o3-mini high is the ballpark the Deep Research docs quote at default breadth 4 / depth 2 / concurrency 4. Those knobs multiply branch count fast. Raise concurrency only when RAM and API tier keep up; drop breadth/depth for cheap scouting. Tavily or OpenAI rate limits show up as thin source lists or hollow sections – backoff, or swap RETRIEVER. Call get_costs() after a run if the invoice surprise is what you are debugging.