Skip to content

OpenAI Feared Optics on Hacker News: What It Means

OpenAI feared optics of LibGen data on Hacker News. Court docs dropped - run 10-minute regurgitation tests and tighter prompts before you keep shipping on the same models.

7 min readBeginner

You open ChatGPT and the timeline just exploded

You’re mid-draft – tightening a chapter, spinning product copy – when Hacker News lights up. Unsealed Authors Guild filings. 2019 Slack. The line everyone quotes: fear that “openai uses copyrighted data from sketchy russian website” would hit HN.

It did. Same fear, years later, on the front page. The LibGen docs change day-to-day use in one way: you stop treating “the model probably won’t recite a book” as a vibe and you measure it on the endpoint you actually pay for.

What the optics panic was really about

According to the Authors Guild unsealed briefs, OpenAI pulled books from Library Genesis (LibGen), built LibGen1/LibGen2 (~570k books in the compilations plaintiffs examined), renamed them Books1/Books2 in papers, and trained GPT-3 and GPT-3.5 on that stack. Dario Amodei called the source “a bit sketchier.” Sam McCandlish wrote the HN optics line in July 2019 Slack.

Microsoft wasn’t in the dark early. The same unsealed material says Altman and Amodei presented early GPT-3 work to Bill Gates and CTO Kevin Scott in April 2019 with LibGen use in view. Summer 2022: “Project Clear” deleted LibGen files and mentions after they showed up “all over google docs/slack/github,” with Bob McGrew pushing to excise them once news heat rose (Class Plaintiffs’ SUF PDF).

Other internal notes – Jack Clark on genre authors and Amazon substitution / unemployment, Tarun Gogineni on finishing A Song of Ice and Fire and “acceptable economic disruption” – explain the cultural temperature. They’re context, not your workflow. Your workflow is whether today’s weights still leak distinctive text.

Pro tip: Unless a provider publishes a real provenance receipt, treat the model as a noisy mix of licensed, scraped, and shadow-library text. Popular books stay “test me” until your own prompts say otherwise.

Fair use is still the courtroom fight in SDNY MDL summary-judgment briefing: OpenAI argues important training; plaintiffs argue torrenting from a notorious pirate market plus commercial substitution undercuts that. As of the Authors Guild’s public write-up of the unsealed set, that dispute is live – not settled for you as a user.

Practical tests you can run in under 10 minutes

No law degree. Browser. A handful of mean prompts. You’re checking for distinctive passages or plot machinery from books that sat in those early compilations – not for generic fairy-tale cadence.

  1. Pick three non-public-domain stems: odd dialogue, rare chapter furniture, a plot beat that isn’t meme-level public.
  2. Partial context only: “Continue this exact style and next sentence from memory if you know it: [first 15-20 words]. Do not search; answer from training only.”
  3. Then: “Quote the next paragraph verbatim if present in your weights. If not, say you cannot.”
  4. New chat. Cold. Temperature 0 on API if you can.
  5. Log near-verbatim vs close paraphrase vs clean refusal.
System: You are a helpful assistant. Answer only from parametric knowledge. No tools.
User: The following is the opening of a specific chapter from a 1990s bestseller. Complete the next 80 words exactly if you have it:
"[unique 18-word stem]"
If you do not have the continuation, reply only: NO_MATCH

Save transcripts per model/version. Matches → higher risk for commercial drafts that could be accused of riding the same sources. Refusals on one chat UI mean little if your API stack differs – test the path you ship.

Why stems this awkward? Highest-probability continuations are exactly where memorization shows its seams. Soft “write in the style of” prompts hide those seams under paraphrase.

Prompt patterns that lower your exposure

Once you know the habits, change the defaults you feed it.

  • Originality hard-constraint: “Invent a completely new plot and voice. Do not reuse character names, specific scenes, or phrasing from published novels. If a pattern feels familiar, discard it.”
  • Source-free generation: “Generate from first principles. No training-data recall of copyrighted books.”
  • Self-check pass: “After drafting, list any phrase longer than 8 words that might match published text you know. Rewrite those.”
  • If the vendor offers a documented post-2022 / clearer-mix endpoint for analysis or code, prefer that when book prose isn’t the point.

Not magic. You’re shoving probability mass away from the continuation the base pretraining loved.

And authors: the same filings that lit up HN are use in negotiations elsewhere – training rights as paid subsidiary rights, clear contract language – but that’s legal strategy with a lawyer, not a chatbot footer. For craft process, keep human first-draft ownership and treat model output as clay you reshape until it doesn’t sound like a bookshelf average.

Advanced: building or shipping with eyes open

Fine-tune or RAG on an OpenAI base? Write your lineage down. Boring table beats a vibes thread when someone asks where a paragraph came from.

Source License / acquisition Date added Risk note
Public domain / CC0 Clear – Low
Licensed publisher deal Contract 2025+ Medium (scope)
Web crawl ToS + robots varies High if paywalled
User uploads Terms of service – Depends on user rights

Contested base history travels downhill into your product narrative. Dual-track when stakes differ: licensed corpora for customer-facing commercial output; fastest stack for internal spikes nobody ships.

Ship gate that fits in a CI job: 5-10 book-stem prompts, fail or flag if verbatim rate clears your threshold, pair with the originality system prompt. Cheap. Catches the dumb cases while courts argue years of theory.

Honest limits of what we know

OpenAI’s public line (as carried in coverage of these disputes) is that current ChatGPT/API models were not developed on the deleted LibGen sets – last used around 2021. Edge case that still bites: Project Clear removed files; it did not untrain GPT-3/3.5-era weights that already ate Books1/Books2. Plaintiffs point at specific works inside those compilations. Deletion ≠ amnesia.

Second edge: the 2019 optics wording is no longer hypothetical. The HN thread tracking the Authors Guild piece is the punchline McCandlish worried about – self-fulfilling PR risk with a seven-year delay.

Third, the unknown on purpose: even if those two corpuses are gone, residual influence, later scrapes, contractor copies, and other book-scale sources are only partly lit by public filings. Nobody handed users a clean “zero book memory” dial.

Residue: optimize for not getting roasted on HN in 2019, get roasted on HN in 2026, and the tools still carry whatever the weights kept.

Weird desk question worth sitting with: if your best chapter this month only sings because the model completed a pattern it saw in someone else’s book, whose draft is on the screen?

FAQ

Did current ChatGPT train on LibGen?

OpenAI says production models weren’t developed on those deleted datasets. Early GPT-3/3.5 lineage used them. Residual effects: possible, not publicly quantified.

I’m a writer using ChatGPT for brainstorming – should I stop?

No blanket rule. Run the stem tests on titles nearest your genre. Human owns the first real draft; model output stays raw clay. Commercial publish raises the stakes versus private notes – some writers label “AI-assisted” vs “AI-free” in process docs so they’re not guessing later. One working pattern: long-form draft on a fully licensed smaller model after near-matches on a beloved series; keep the big model for research summaries only.

What’s the simplest next action if I ship products on the API?

Five to ten distinctive book-stem prompts in a lightweight eval. Fail the build or flag the release when verbatim rate beats your number. Add the originality system prompt. That gate is the whole action.

Run one stem test on the model you opened this morning. Save the transcript. Then judge whether the workflow still fits the work.