Local long-context DNA scoring without waiting on hosted NIM queues. That is the job. Evo 2 package evo2 0.6.0 (PyPI, released June 19, 2026) is the Arc Institute open stack with Vortex inference and a light path that does not need Hopper FP8 for the 7B family.
Deploy guide only – no paper tour. Sources: Arc Institute evo2 README, PyPI evo2, v0.5.0 20B release notes, and the Nature Evo 2 paper.
Hardware reality check before you touch pip
Match the checkpoint to the GPU you actually have. Official FP8 table plus the v0.5.0 notes:
| Checkpoint | FP8 / TE | Typical GPU footprint | Notes |
|---|---|---|---|
| evo2_7b / 7b_base / 7b_262k | No | ~14-15 GB weight class | Light install path |
| evo2_20b | Yes (Hopper) | 38 GB weights; ~46 GB peak @ 8 kb on 1× H100 | Sweet spot on one Hopper card |
| evo2_40b / 40b_base | Yes | ~80 GB weights; usually 2× H100 | Vortex shards across visible CUDA devices |
| evo2_1b_base | Yes | Small but still TE/FP8 | Docs discourage it for quality |
Linux preferred (WSL2 limited). README floor as of 0.6.0: CUDA 12.1+, cuDNN 9.3+, GCC 9+ or Clang 10+ (C++17), Python 3.11 or 3.12, Torch 2.6.x or 2.7.x. Disk is the silent killer – plan HF cache space up front (7B ~14-15 GB; 20B weights 38 GB; 40B ~80 GB).
Plenty of clusters still hand you a 20 GB home quota and a shiny GPU. The model does not care about your quota. The cache does. Set the cache path before the first download or you will learn this the expensive way.
Install Evo 2 0.6.0 (light path first)
Clean env. Do not mix a random system CUDA toolkit with conda packages and hope.
conda create -n evo2 python=3.12 -y
conda activate evo2
# example CUDA 12.8 wheel - match your driver
pip install torch==2.7.1 --index-url https://download.pytorch.org/whl/cu128
pip install flash-attn==2.8.0.post2 --no-build-isolation
pip install evo2==0.6.0
Enough for evo2_7b, evo2_7b_base, and evo2_7b_262k. When TE is missing and the name looks like a 7B model, the library falls back to bf16 projections – that is the whole point of starting light.
Full path when you need 20B / 40B FP8 numerics:
conda install -c nvidia cuda-nvcc cuda-cudart-dev -y
conda install -c conda-forge transformer-engine-torch=2.3.0 -y
pip install flash-attn==2.8.0.post2 --no-build-isolation
pip install evo2==0.6.0
From source only if you want main tip:
git clone https://github.com/ArcInstitute/evo2
cd evo2
pip install -e .
Pro tip:
export HF_HOME=/scratch/$USER/hf(any large volume) before the firstEvo2(...)call. First download lands in the Hugging Face cache. Pip never owns those files.
Docker when conda fights you
docker build -t evo2 .
docker run -it --rm --gpus '"device=0"'
-v $PWD/huggingface:/root/.cache/huggingface
evo2 bash
# inside:
python -m evo2.test.test_evo2_generation --model_name evo2_7b
Apptainer/Singularity: bind the same cache and pass --nv. Community SIFs exist; the official Dockerfile tracks 0.6.x pins more closely.
Bare-bones load
No YAML server file. The checkpoint name is the config.
from evo2 import Evo2
model = Evo2("evo2_7b")
# optional: Evo2("evo2_7b", use_kernels=True) # needs vtx>=1.1.0
out = model.generate(
prompt_seqs=["ACGT"],
n_tokens=100,
temperature=0.7,
top_k=4,
)
print(out.sequences[0])
Hopper single card: Evo2("evo2_20b"). 40B: leave multiple GPUs visible and let Vortex shard – skip manual .to(device) on multi-GPU loads. (Kernels path showed up around the 0.6 timeframe; if use_kernels=True explodes, check vtx first.)
Verify
python -m evo2.test.test_evo2_generation --model_name evo2_7b
# or: --model_name evo2_20b / evo2_40b
# optional: --use_kernels
python -c "import evo2; print(evo2.__version__)" # expect 0.6.0
Short DNA string + clean exit = you are done with install theater.
Common install errors and fixes
- RuntimeError: Found empty
transformer-enginemeta package – meta package without framework extras (GitHub #201 pattern). Preferconda-forge transformer-engine-torch=2.3.0beforepip install evo2; fresh env if the broken meta is already baked in. - ImportError: FP8 / TE required even for evo2_7b – name/config missed the 7B fallback, or leftover vortex bits (issue #208 territory). Reinstall light path only, call exactly
Evo2('evo2_7b'), purge half-dead TE site-packages. - flash-attn build failures / version warnings – Torch first, then
flash-attn==2.8.0.post2 --no-build-isolation, working nvcc. Docs still pin 2.8.0.post2 for 0.6.x; some stacks warn about older supported ranges (issue #220). Run the generation test anyway. - flash-attn runtime “operation not supported” on older Ampere – you drifted off the light-path sweet spot. Stay on 7B bf16 with a driver/CUDA pair the README actually lists.
Weird question worth asking once: are you debugging the install, or debugging a GPU that was never going to run FP8 numerics? Those two tickets feel identical at 1 a.m.
Upgrade and uninstall
pip install -U evo2==0.6.0
# or
pip install -U git+https://github.com/ArcInstitute/evo2.git
Weights stay in HF cache. Pip upgrade does not refresh them.
pip uninstall evo2 flash-attn -y
# optional conda TE cleanup
rm -rf $HF_HOME/hub/models--arcinstitute* # only if you are sure
Docker: docker rmi evo2 and drop the bound cache volume if you mounted one.
Run the evo2_7b generation test on your box first. Open the repo BRCA1 notebook only after that smoke test passes.
FAQ
Do I need an H100 to use Evo 2 at all?
No. 7B runs without FP8/TE on supported CUDA GPUs. Hopper is for 20B/40B FP8 numerics.
Light install worked, but evo2_20b crashes – why?
You skipped Transformer Engine on purpose. Full path: conda transformer-engine-torch=2.3.0, confirm nvidia-smi shows Hopper-class FP8 hardware, then Evo2("evo2_20b"). Same machine that loved 7B will hard-fail 20B without that stack – that is expected, not a corrupt wheel.
Where did tens of GB of disk go after the first run?
Hugging Face hub cache under HF_HOME (default ~/.cache/huggingface). Official Docker guidance is to volume-mount that path for a reason; Apptainer users bind it the same way. Pip uninstall leaves models--arcinstitute--* folders alone. Point HF_HOME at scratch before first load if home quotas are tight – cleanup is manual deletion of those hub dirs, not another pip command.