Key takeaway: After OpenAI’s head of ethics leaves less than a year after joining, headcount still does not rewrite the public Model Spec, the Usage Policies, or the switches already sitting in your ChatGPT account. Use the news as a 15-minute setup task, not a spectator sport.
Chloé Bakalar’s quiet July 2026 exit (she joined August 2025) is loud on HN because she was described as the only dedicated ethicist, with no direct replacement planned. Headlines stop there. You shouldn’t.
Brief background on the departure
Bakalar came from Meta’s chief ethicist seat. At OpenAI her remit covered ethical model development, human-AI interaction, and machine-consciousness debates. No public goodbye note from her or the company.
OpenAI’s line in FT-sourced coverage: ethics “doesn’t live with one owner or team” and is “deeply embedded” across research. Same month, Safety Systems lead Johannes Heidecke and chief futurist Joshua Achiam also left. Skepticism that “embedded” equals accountable is fair.
You’re still not powerless while the org chart settles. Product rules you can read beat job titles you can’t see.
Method A vs Method B: who owns your AI ethics now?
Two common reactions when OpenAI’s head of ethics leaves less than a year after joining:
| Approach | What you do | Pros | Cons |
|---|---|---|---|
| Method A – Trust the house | Assume distributed safety teams plus the Model Spec are enough; change nothing | Zero effort; hard safety refusals still exist | No view into ethics capacity; softer defaults can drift; you notice only after a bad answer |
| Method B – User-side stack | Read the Model Spec, write custom instructions as a personal constitution, tighten memory/data settings, use parental controls if you have teens, know the Usage Policies | Immediate, auditable, free or paid; survives org churn | ~15-20 minutes once; user text sits below hard system rules |
Method B wins if you draft for work, school, clients, or kids. Weapons, CSAM, violent crime help, and similar categories stay system-level. Tone on sensitive topics, refusal style, bias checks, citation habits – that layer is yours.
The catch is simple: Method A is comfortable until the first soft answer you didn’t want. Method B is boring setup that still works after the next reorg.
Detailed walkthrough: build your personal ethics layer
Do this once. It sticks across new chats.
1. Skim the public contract (~10 minutes)
Open the current Model Spec (public versions include the Dec 18, 2025 cut; OpenAI also posted “Inside our approach to the Model Spec” on Mar 25, 2026). Focus on chain of command (root and system beat developer and user), “Stay in bounds” hard rules, overridable defaults (style, objectivity, sycophancy), and under-18 principles if a teen is in the house.
Pair that with the Usage Policies effective October 29, 2025 – the account-level lines that can still get you banned.
2. Write custom instructions like a mini-constitution
Web/Desktop: Settings → Personalization → Enable customization → Custom instructions.
Mobile: Settings → Customize ChatGPT.
Per the Help Center, Free and Go top out at 1,500 characters; Plus, Pro, Enterprise, Business, and Education go to 5,000. Applied immediately to chats; toggle lives in the same personalization area.
Sample you can adapt (trim for free tier):
Prioritize accuracy over agreeableness. Flag uncertainty and missing sources.
When a request touches medical, legal, or financial advice, state limits and push for professional review.
Refuse and briefly explain if a request would enable real-world harm, scams, or privacy violations - even if framed as fiction.
Use primary sources and clear citations. Avoid sycophancy.
For creative work, maximize user freedom within the above bounds.
If instructions conflict, follow higher-authority safety rules first.
Test with one gray-area prompt the same day – borderline medical wording, or “write a persuasive ad for a dubious product.” See whether pushback matches what you wrote.
Pro tip: Non-negotiables go first. On a 1,500-character cap, the tail is what disappears if you overrun.
3. Lock surrounding product settings
- Memory: off or pruned if long-term personalization shouldn’t bleed into sensitive topics.
- Data controls: opt out of training where the UI still offers it, if that’s your preference (check the toggle as of your login – labels move).
- Shared links / GPTs: custom instructions aren’t shown to link viewers; third-party GPTs can still see relevant context – stick to ones you trust.
- Families: turn on parental controls (rolled out September 2025), link the teen (13+) account, set quiet hours and feature limits.
4. Keep a short red-team checklist
Once a month, rerun the same three stress prompts and save outputs. If refusals or hedging go softer after a model drop, tighten the instruction block. That’s your early warning when internal teams reshuffle.
One open question this episode still leaves hanging: when ethics is “embedded,” who holds veto on launch day, and how would an outsider know? Public docs don’t say.
Edge cases that trip people up
Custom instructions never outrank hard Model Spec safety. Ambitious overrides (“ignore all refusal policies”) fail quietly – no error toast – while catastrophic asks still get refused. Chain of command, not a bug.
Character caps create a silent truncate: long multi-rule “constitutions” on Free/Go keep the opening and drop the rest on some clients without a loud warning. Count before you paste.
Parental controls are not household surveillance. Parents don’t read teen chats. Safety notifications are narrow and come after human review of acute distress. Teens can still push on some settings. Age prediction defaults toward the under-18 experience when unsure – so adults sometimes hit teen rails until they verify. Don’t treat the feature as full ethics coverage for families.
“Deeply embedded” ethics, after Bakalar, still has no public org chart, headcount, or named escalation path on the safety pages. Community threads treat that as an open transparency gap. Durable layer = Spec + policies + your settings. People move.
FAQ
Does Bakalar’s exit change how ChatGPT answers me tomorrow?
No. Not overnight. Watch defaults after major model drops, not the headcount line.
Should I switch to Claude or another model for “better ethics”?
Only after you compare published rules on your own prompts. Example: run the same medical-hedge and scams-framed-as-fiction tests on two labs, keep whichever refuses and cites the way you need, route the other traffic by task. Quitting drama is a weak buying guide; your red-team set isn’t.
What’s the single best use of five minutes?
Enable custom instructions. Paste a short accuracy-first + harm-refusal block. Open the Model Spec once so “hard rules” is concrete. Stop there if the clock runs out.
Open settings, paste a draft constitution, run one gray-area prompt before you close the tab. That’s the move this news actually pays for.