Skip to content

Deploy CUAD v1 for Contract Clause Extraction

Install CUAD v1 end-to-end: system specs, Zenodo/GitHub downloads, pinned deps, first extract run, and real GitHub install fixes.

8 min readIntermediate

Contract clause extraction with CUAD v1 (Contract Understanding Atticus Dataset) replaces rereading every NDA for the same 41 clause types. You get expert-labeled spans – 510 commercial contracts, 13,000+ labels, CC BY 4.0 – and research scripts/checkpoints you can actually run.

This guide targets the public release that is still CUAD v1 (NeurIPS 2021 paper; no later major dataset tag as of late 2025). Clone code, pull data and weights from the real hosts, pin a stack that still executes train.py, smoke-test one extract, then fix the install failures that show up in the Atticus issue tracker.

System requirements

Authors tested Python 3.8, PyTorch 1.7, and Hugging Face Transformers 4.3/4.4 (repo readme). Treat that trio as the compatibility target. Fresh global Transformers will fight you.

Resource Minimum (practical estimate) Recommended (as of late 2025)
OS Linux or macOS (Windows via WSL2) Ubuntu-class Linux
CPU Several cores 8+ cores for snappier preprocessing
RAM ~16 GB class machine 32 GB+ if you keep PDFs, caches, and large checkpoints
GPU None for slow RoBERTa-base inference 16 GB+ VRAM for train/eval; more headroom for DeBERTa-xlarge (~900M params)
Disk Code + small data.zip Tens of GB if you keep full CUAD_v1 PDFs/TXTs, multiple checkpoints, HF caches
Python 3.8+ 3.8-3.10 in a fresh venv

Published checkpoint classes (readme/Zenodo): RoBERTa-base ~100M params, RoBERTa-large ~300M, DeBERTa-xlarge ~900M. Hardware pain scales with that list – not with the JSON file size.

Official download sources

  • Code:https://github.com/The-Atticus-Project/cuad
  • Training/eval JSON (lightweight):data.zip in the repo (~18 MB; HF builders hit the same raw URL)
  • Full CUAD v1 bundle (PDFs, TXTs, CSV, Excels):Zenodo record 4595826
  • Fine-tuned checkpoints:Zenodo record 4599830
  • Paper:arXiv:2103.06268
  • Optional HF path:load_dataset("theatticusproject/cuad-qa") or legacy "cuad" (train 22 450 / test 4 182 examples per the dataset card)

License is CC BY 4.0 – commercial and research use with attribution.

Two download lanes exist on purpose: the tiny repo data.zip is enough to train/eval in SQuAD-style JSON; Zenodo’s full zip is for when you care about original PDF/TXT layout. Mixing them up is how people “install CUAD” and still have no contracts on disk.

Install CUAD v1 step by step

Prefer a clean virtualenv. Do not bolt this onto a random global Transformers install.

# 1) Clone code
git clone https://github.com/The-Atticus-Project/cuad.git
cd cuad

# 2) Environment
python3.8 -m venv .venv
source .venv/bin/activate # Windows: .venvScriptsactivate
pip install -U pip wheel

# 3) Pin a stack close to what the authors tested
# Adjust the torch index URL for your CUDA (or use CPU wheels)
pip install torch==1.7.1 torchvision==0.8.2 -f https://download.pytorch.org/whl/torch_stable.html
pip install "transformers==4.4.2" "tokenizers<0.11" scikit-learn pandas numpy tqdm tensorboard

# 4) Unpack training data shipped with the repo
unzip data.zip -d data
# Expect train_separate_questions.json and test.json under data/

# 5) (Optional) Full corpus with PDFs/TXTs
# Download CUAD_v1.zip from https://zenodo.org/records/4595826 and unzip beside the repo

Inference-only? Skip the heavy train extras. Pull a checkpoint zip from Zenodo 4599830 and extract to e.g. ./trained_models/roberta-base.

Pro tip: Pin transformers==4.4.2 (or another 4.3-4.9 line still exporting the old optimizer symbol) before you run train.py. Details under Common install errors.

First-time configuration

Smallest folder layout that works after unzip:

cuad/
 train.py evaluate.py utils.py run.sh
 data/
 train_separate_questions.json
 test.json
 trained_models/ # create this
 roberta-base/ # checkpoint files from Zenodo

Stock run.sh on main is the reference config. Example pass with a safer train batch for a single mid-size GPU:

CUDA_VISIBLE_DEVICES=0 python train.py 
 --output_dir ./trained_models/roberta-base 
 --model_type roberta 
 --model_name_or_path roberta-base 
 --train_file ./data/train_separate_questions.json 
 --predict_file ./data/test.json 
 --do_train --do_eval 
 --version_2_with_negative 
 --learning_rate 1e-4 
 --num_train_epochs 4 
 --per_gpu_eval_batch_size 8 
 --per_gpu_train_batch_size 8 
 --max_seq_length 512 
 --max_answer_length 512 
 --doc_stride 256 
 --save_steps 1000 
 --n_best_size 20 
 --overwrite_output_dir

Keep --version_2_with_negative. Lots of category×contract pairs have no span – same unanswerable setup as SQuAD 2.0. Official run.sh also sets max_seq_length 512 and doc_stride 256; production inference must honor that windowing story (see errors).

Inference-only path after checkpoint download:

from transformers import AutoTokenizer, AutoModelForQuestionAnswering
import torch

path = "./trained_models/roberta-base"
tokenizer = AutoTokenizer.from_pretrained(path, use_fast=False)
model = AutoModelForQuestionAnswering.from_pretrained(path)
model.eval()

question = (
 'Highlight the parts (if any) of this contract related to "Governing Law" '
 "that should be reviewed by a lawyer."
)
# context = full contract text (or a sliding window ≤512 tokens total with question)
inputs = tokenizer(question, context, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
 out = model(**inputs)
start = int(out.start_logits.argmax())
end = int(out.end_logits.argmax())
print(tokenizer.decode(inputs["input_ids"][0][start : end + 1]))

Fast tokenizers and this QA head historically fought overflow handling – use_fast=False matches demos that actually returned spans.

Verify the install works

Three checks, short to long:

  1. Files present:ls data/train_separate_questions.json data/test.json
  2. Model loads: run the snippet above; non-empty decode or empty string for absent clauses – not a stack trace.
  3. Official metrics path: after train.py --do_eval (or copying n-best outputs into the checkpoint folder), run:
python evaluate.py
# default paths inside evaluate.py point at ./data/test.json and
# ./trained_models/roberta-base - edit if your layout differs

AUPR and precision at high recall beat vanilla F1 for this task. Per the paper (arXiv:2103.06268), the strong DeBERTa-class line sat near 47.8% AUPR, 44.0% precision @ 80% recall, and 17.8% precision @ 90% recall – sparse spans, not clean SQuAD-easy answers.

HF shortcut without cloning train code:

from datasets import load_dataset
ds = load_dataset("theatticusproject/cuad-qa")
print(ds)
print(ds["train"][0]["question"][:120], ds["train"][0]["answers"])

Common install errors and fixes

The catch is these are the failures people file – GitHub issues #18, #21, #22 and HF cuad-qa threads – not hypothetical lint.

ImportError: cannot import name 'AdamW' from 'transformers'
Turns out modern Transformers dropped that re-export. Pin transformers==4.4.2 (or a 4.3-4.9 line that still exports it), or edit train.py to from torch.optim import AdamW.

Training dies with CUDA OOM
Stock per_gpu_train_batch_size=40 in run.sh is aggressive on 16 GB cards. Drop to 4-8 (as in the command above). DeBERTa-xlarge (~900M) is a different hardware class than RoBERTa-base (~100M).

FileNotFoundError for nbest_predictions_.json
evaluate.py expects that file under the model directory after eval (issue #21 pattern: paths like ./trained_models/roberta-base/...). Run with --do_eval, confirm output_dir, then align the hard-coded paths at the bottom of evaluate.py.

IndexError on answer["answer_start"][0] in custom preprocessing
Empty lists are normal: many labels have no span. Guard with if len(answers["answer_start"]) == 0: and treat as unanswerable (CLS start/end).

Late-page clauses never show up
max_seq_length 512 plus question tokens truncates long SEC filings. Training already windows with doc_stride 256; if inference feeds only the first window, you silently miss answers further down.

Upgrade, migration, uninstall

No CUAD v2 dataset drop from Atticus as of this writing – v1 stays the reference. “Upgrade” usually means git pull, replace Zenodo checkpoints, or port training to a current Trainer + torch.optim.AdamW while keeping the same JSON labels.

deactivate
rm -rf .venv trained_models data __pycache__
# optional: rm -rf ~/.cache/huggingface/datasets/*cuad*
cd .. && rm -rf cuad

Adjacent pieces: MAUD (merger agreements) from the same org, or HF evaluate.load("cuad") for the scoring wrapper.

Long contracts still feel like searching a warehouse with a flashlight. The dataset doesn’t shrink the warehouse – it labels the shelves so your model knows what a target clause looks like when it finally walks past it.

FAQ

Is CUAD a Docker app I can compose up?

No. Dataset + research train/eval scripts + published QA checkpoints. You bring Python and PyTorch.

Git clone alone isn’t enough – what else do I need for a first extract?

Unzip data.zip (or load HF). Drop a Zenodo checkpoint into trained_models/. Pin Transformers. Run the AutoModelForQuestionAnswering snippet on a real contract window. Want metrics? train.py --do_eval or n-best JSON into evaluate.py. Need raw layout? Full PDFs sit in Zenodo CUAD_v1.zip (4595826), not in the tiny repo zip.

Why pin ancient Transformers instead of latest?

train.py is a lightly adapted 2021 SQuAD fine-tune script. It still imports AdamW from transformers and assumes that processor surface. Latest releases removed the import and moved helpers. Pinning 4.4.x is the known-good path in minutes. Porting to Trainer + torch.optim.AdamW is the better long-term fix – that’s a code change, not an install flag. Community tickets against the repo are still about this mismatch.

Next action: create the venv, unzip data.zip, download RoBERTa-base from Zenodo 4599830, and run one clause query on a contract from test.json before you touch multi-GPU training.