Contract clause extraction with CUAD v1 (Contract Understanding Atticus Dataset) replaces rereading every NDA for the same 41 clause types. You get expert-labeled spans – 510 commercial contracts, 13,000+ labels, CC BY 4.0 – and research scripts/checkpoints you can actually run.
This guide targets the public release that is still CUAD v1 (NeurIPS 2021 paper; no later major dataset tag as of late 2025). Clone code, pull data and weights from the real hosts, pin a stack that still executes train.py, smoke-test one extract, then fix the install failures that show up in the Atticus issue tracker.
System requirements
Authors tested Python 3.8, PyTorch 1.7, and Hugging Face Transformers 4.3/4.4 (repo readme). Treat that trio as the compatibility target. Fresh global Transformers will fight you.
| Resource | Minimum (practical estimate) | Recommended (as of late 2025) |
|---|---|---|
| OS | Linux or macOS (Windows via WSL2) | Ubuntu-class Linux |
| CPU | Several cores | 8+ cores for snappier preprocessing |
| RAM | ~16 GB class machine | 32 GB+ if you keep PDFs, caches, and large checkpoints |
| GPU | None for slow RoBERTa-base inference | 16 GB+ VRAM for train/eval; more headroom for DeBERTa-xlarge (~900M params) |
| Disk | Code + small data.zip |
Tens of GB if you keep full CUAD_v1 PDFs/TXTs, multiple checkpoints, HF caches |
| Python | 3.8+ | 3.8-3.10 in a fresh venv |
Published checkpoint classes (readme/Zenodo): RoBERTa-base ~100M params, RoBERTa-large ~300M, DeBERTa-xlarge ~900M. Hardware pain scales with that list – not with the JSON file size.
Official download sources
- Code:
https://github.com/The-Atticus-Project/cuad - Training/eval JSON (lightweight):
data.zipin the repo (~18 MB; HF builders hit the same raw URL) - Full CUAD v1 bundle (PDFs, TXTs, CSV, Excels):Zenodo record 4595826
- Fine-tuned checkpoints:Zenodo record 4599830
- Paper:arXiv:2103.06268
- Optional HF path:
load_dataset("theatticusproject/cuad-qa")or legacy"cuad"(train 22 450 / test 4 182 examples per the dataset card)
License is CC BY 4.0 – commercial and research use with attribution.
Two download lanes exist on purpose: the tiny repo data.zip is enough to train/eval in SQuAD-style JSON; Zenodo’s full zip is for when you care about original PDF/TXT layout. Mixing them up is how people “install CUAD” and still have no contracts on disk.
Install CUAD v1 step by step
Prefer a clean virtualenv. Do not bolt this onto a random global Transformers install.
# 1) Clone code
git clone https://github.com/The-Atticus-Project/cuad.git
cd cuad
# 2) Environment
python3.8 -m venv .venv
source .venv/bin/activate # Windows: .venvScriptsactivate
pip install -U pip wheel
# 3) Pin a stack close to what the authors tested
# Adjust the torch index URL for your CUDA (or use CPU wheels)
pip install torch==1.7.1 torchvision==0.8.2 -f https://download.pytorch.org/whl/torch_stable.html
pip install "transformers==4.4.2" "tokenizers<0.11" scikit-learn pandas numpy tqdm tensorboard
# 4) Unpack training data shipped with the repo
unzip data.zip -d data
# Expect train_separate_questions.json and test.json under data/
# 5) (Optional) Full corpus with PDFs/TXTs
# Download CUAD_v1.zip from https://zenodo.org/records/4595826 and unzip beside the repo
Inference-only? Skip the heavy train extras. Pull a checkpoint zip from Zenodo 4599830 and extract to e.g. ./trained_models/roberta-base.
Pro tip: Pin
transformers==4.4.2(or another 4.3-4.9 line still exporting the old optimizer symbol) before you runtrain.py. Details under Common install errors.
First-time configuration
Smallest folder layout that works after unzip:
cuad/
train.py evaluate.py utils.py run.sh
data/
train_separate_questions.json
test.json
trained_models/ # create this
roberta-base/ # checkpoint files from Zenodo
Stock run.sh on main is the reference config. Example pass with a safer train batch for a single mid-size GPU:
CUDA_VISIBLE_DEVICES=0 python train.py
--output_dir ./trained_models/roberta-base
--model_type roberta
--model_name_or_path roberta-base
--train_file ./data/train_separate_questions.json
--predict_file ./data/test.json
--do_train --do_eval
--version_2_with_negative
--learning_rate 1e-4
--num_train_epochs 4
--per_gpu_eval_batch_size 8
--per_gpu_train_batch_size 8
--max_seq_length 512
--max_answer_length 512
--doc_stride 256
--save_steps 1000
--n_best_size 20
--overwrite_output_dir
Keep --version_2_with_negative. Lots of category×contract pairs have no span – same unanswerable setup as SQuAD 2.0. Official run.sh also sets max_seq_length 512 and doc_stride 256; production inference must honor that windowing story (see errors).
Inference-only path after checkpoint download:
from transformers import AutoTokenizer, AutoModelForQuestionAnswering
import torch
path = "./trained_models/roberta-base"
tokenizer = AutoTokenizer.from_pretrained(path, use_fast=False)
model = AutoModelForQuestionAnswering.from_pretrained(path)
model.eval()
question = (
'Highlight the parts (if any) of this contract related to "Governing Law" '
"that should be reviewed by a lawyer."
)
# context = full contract text (or a sliding window ≤512 tokens total with question)
inputs = tokenizer(question, context, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
out = model(**inputs)
start = int(out.start_logits.argmax())
end = int(out.end_logits.argmax())
print(tokenizer.decode(inputs["input_ids"][0][start : end + 1]))
Fast tokenizers and this QA head historically fought overflow handling – use_fast=False matches demos that actually returned spans.
Verify the install works
Three checks, short to long:
- Files present:
ls data/train_separate_questions.json data/test.json - Model loads: run the snippet above; non-empty decode or empty string for absent clauses – not a stack trace.
- Official metrics path: after
train.py --do_eval(or copying n-best outputs into the checkpoint folder), run:
python evaluate.py
# default paths inside evaluate.py point at ./data/test.json and
# ./trained_models/roberta-base - edit if your layout differs
AUPR and precision at high recall beat vanilla F1 for this task. Per the paper (arXiv:2103.06268), the strong DeBERTa-class line sat near 47.8% AUPR, 44.0% precision @ 80% recall, and 17.8% precision @ 90% recall – sparse spans, not clean SQuAD-easy answers.
HF shortcut without cloning train code:
from datasets import load_dataset
ds = load_dataset("theatticusproject/cuad-qa")
print(ds)
print(ds["train"][0]["question"][:120], ds["train"][0]["answers"])
Common install errors and fixes
The catch is these are the failures people file – GitHub issues #18, #21, #22 and HF cuad-qa threads – not hypothetical lint.
ImportError: cannot import name 'AdamW' from 'transformers'
Turns out modern Transformers dropped that re-export. Pin transformers==4.4.2 (or a 4.3-4.9 line that still exports it), or edit train.py to from torch.optim import AdamW.
Training dies with CUDA OOM
Stock per_gpu_train_batch_size=40 in run.sh is aggressive on 16 GB cards. Drop to 4-8 (as in the command above). DeBERTa-xlarge (~900M) is a different hardware class than RoBERTa-base (~100M).
FileNotFoundError for nbest_predictions_.jsonevaluate.py expects that file under the model directory after eval (issue #21 pattern: paths like ./trained_models/roberta-base/...). Run with --do_eval, confirm output_dir, then align the hard-coded paths at the bottom of evaluate.py.
IndexError on answer["answer_start"][0] in custom preprocessing
Empty lists are normal: many labels have no span. Guard with if len(answers["answer_start"]) == 0: and treat as unanswerable (CLS start/end).
Late-page clauses never show upmax_seq_length 512 plus question tokens truncates long SEC filings. Training already windows with doc_stride 256; if inference feeds only the first window, you silently miss answers further down.
Upgrade, migration, uninstall
No CUAD v2 dataset drop from Atticus as of this writing – v1 stays the reference. “Upgrade” usually means git pull, replace Zenodo checkpoints, or port training to a current Trainer + torch.optim.AdamW while keeping the same JSON labels.
deactivate
rm -rf .venv trained_models data __pycache__
# optional: rm -rf ~/.cache/huggingface/datasets/*cuad*
cd .. && rm -rf cuad
Adjacent pieces: MAUD (merger agreements) from the same org, or HF evaluate.load("cuad") for the scoring wrapper.
Long contracts still feel like searching a warehouse with a flashlight. The dataset doesn’t shrink the warehouse – it labels the shelves so your model knows what a target clause looks like when it finally walks past it.
FAQ
Is CUAD a Docker app I can compose up?
No. Dataset + research train/eval scripts + published QA checkpoints. You bring Python and PyTorch.
Git clone alone isn’t enough – what else do I need for a first extract?
Unzip data.zip (or load HF). Drop a Zenodo checkpoint into trained_models/. Pin Transformers. Run the AutoModelForQuestionAnswering snippet on a real contract window. Want metrics? train.py --do_eval or n-best JSON into evaluate.py. Need raw layout? Full PDFs sit in Zenodo CUAD_v1.zip (4595826), not in the tiny repo zip.
Why pin ancient Transformers instead of latest?
train.py is a lightly adapted 2021 SQuAD fine-tune script. It still imports AdamW from transformers and assumes that processor surface. Latest releases removed the import and moved helpers. Pinning 4.4.x is the known-good path in minutes. Porting to Trainer + torch.optim.AdamW is the better long-term fix – that’s a code change, not an install flag. Community tickets against the repo are still about this mismatch.
Next action: create the venv, unzip data.zip, download RoBERTa-base from Zenodo 4599830, and run one clause query on a contract from test.json before you touch multi-GPU training.