Key takeaway: AI code isn’t the breakage. Missing architecture intent is. Put a shared map in the repo first – then let the model write inside the lines.
This argument keeps flaring up on DEV, YouTube teardown channels, and agent Slack threads: the assistant ships code that compiles, survives a skim, and still dents production. Same bruise every time. The model isn’t bad at coding. It never got the why behind the shape.
I learned that on a mid-size order service. Cursor drafted a clean handler in a few minutes. PR looked tidy. Prod double-fired a side path nobody had written down. The code was fine. The intent file didn’t exist.
Two ways people try to fix AI code failures
When output drifts, teams usually grab one of two levers.
| Approach | What you do | What usually happens |
|---|---|---|
| Method A – Prompt harder | Longer chats, paste more files, “remember our architecture” every turn | Works once; next session starts cold; window fills; training patterns win again |
| Method B – Intent in the repo | Short standing instructions + ADRs + sink-style boundaries before generation | Same model, smaller blast radius; constraints show up without the lecture |
Method A feels like you’re “using the tool.” Method B costs an afternoon and pays back for the month. AI as amplifier – not my slogan. That’s the framing in Google’s 2025 DORA report: speed goes up, and so does whatever process or structure was already weak. Intent stuck in someone’s head or a stale wiki gives the amplifier nothing solid to lock onto.
Method B wins past toy-repo scale.
Why the model “gets the code” and still misses the system
Local patterns? Models crush those. Tokenize the file, predict the next plausible chunk. Production topology, the incident that flattened a nested JSON field, the rule that payments never call notifications directly – none of that lives in the weights as your system.
Architecture work is still a hard gap for current models. A 2025 systematic review (Bucaioni et al., arXiv:2504.04334) mapped AI onto 14 architecture task areas and still surfaced six AI-specific challenges versus what practitioners need. Community threads say it bluntly: reading the repo ≠ understanding the system.
Ian Bull’s Feb 2026 note draws the cut that sticks in practice. A pipe does work, then kicks a cascade (DB write → trigger → queue → another service). A sink takes input, finishes, stops. Session-reset agents can reason about sinks. Pipes make them invent the chain they cannot see.
Pro tip: Before you ask for a feature, write one non-negotiable intent line: “Handlers validate and call services only; no direct DB access from HTTP.” Park it in standing rules. One sentence beats three paragraphs of vibes.
Walkthrough: put architecture intent where the agent always sees it
This is the Method B loop I run on greenfield and brownfield. Beginner-friendly. Under an hour to stand up.
1. Drop a short AGENTS.md (or Cursor rules)
Agents forget between completions – that is by design. Cursor’s Rules docs spell out the fix: Project Rules (.cursor/rules/*.mdc), User Rules, Team Rules, and plain AGENTS.md inject at the start of model context. Prefer root-level markdown for portability; use .mdc globs when one path needs stricter law.
# AGENTS.md
## What this service is
Order API. Stateless HTTP. All durable state in Postgres.
## Hard boundaries (do not violate)
- HTTP handlers validate input and call services only
- Services own business rules; they never import framework request types
- Repositories are the only modules that touch the DB client
- No cross-module side effects: a service method must not enqueue jobs itself
## Preferred patterns
- Explicit error types with code + message
- Idempotency keys on payment and webhook paths
## Never
- Raw SQL outside repositories
- New top-level folders without an ADR
Keep it tight. Official Cursor guidance (as of their current Rules docs): under 500 lines, split big packs, point at example files instead of pasting a whole style guide. Fat rule files lose when the window fills – the active file’s local pattern shouts louder than page 12 of your manifesto.
2. Add one ADR for the decision that always gets re-litigated
Nygard-style record in docs/adr/. Status, Context, Decision, Consequences. Only the fights that still constrain code:
# ADR-0003: Services never enqueue side effects directly
## Status
Accepted
## Context
We had double notifications when handlers and services both published events.
AI-generated "helpers" kept reintroducing direct queue calls.
## Decision
Services return domain results only. A single application/outbound adapter
publishes events after the transaction commits.
## Consequences
+ Clear sink-style services; easier to test
+ Agents stop inventing dual publish paths
- One extra adapter file to maintain
Link the ADR path from AGENTS.md so the agent knows where intent lives. Skip museum history.
3. Shape modules as sinks before you generate
One messy flow. Paper sketch:
- Input contract (types in)
- Work that stays inside the module
- Output contract (types out)
- No hidden writes to other subsystems
Prompt against the box you drew: “Implement CancelOrder as a sink service per ADR-0003 and AGENTS.md. Return a result object; do not touch the bus.”
You are not asking the model to invent architecture. You are asking it to fill a box.
4. Review for intent, not only syntax
Extra PR/chat question: “Which ADR and boundary does this change obey? Quote them.” No quote → guesswork.
Roughly 45% of AI generation tasks still introduce known flaw classes when security stays implicit – Veracode’s GenAI security reporting (as of the 2026 update) has pass rates stuck near 55-56% even while syntax looks solved. Separately, CodeRabbit’s analysis of 470 PRs put AI-co-authored changes at about 1.7× more issues than human-only ones. Architecture intent does not replace review. It hands review a checklist that is not pure taste.
Edge cases that bite after the happy path
Context pressure wins ugly fights. Open file + retrieved chunks + chat history crowd the window, and a standing architecture note loses to a stronger local pattern. Forbidden import creeps back? Shrink the rule. Raise it higher in AGENTS.md. Attach a glob-scoped rule for that path. One giant manifesto is how critical lines get evicted.
Quiet failure mode I did not expect: temporary task state becomes “architecture.” Agent blocked on a queue name yesterday → treats that name as standing project truth tomorrow on an unrelated worker. Split durable rules from disposable session notes. Say it in the prompt: “Ignore prior task blockers; only AGENTS.md and docs/adr apply.”
Security does not hitch a free ride on layer names. “Parameterized queries only” and “no string-built HTML” belong next to the boundary bullets – or you still get the insecure default the training set loves.
Is perfect documentation required before any AI help? No. Waiting for perfect docs is how teams stay glued to Method A.
FAQ
Do I need Cursor for this, or does AGENTS.md work elsewhere?
No. Plain markdown in the repo is the point. Cursor reads it; other agents increasingly do too. Brand of IDE is secondary.
What if our architecture is already a mess of pipes?
Don’t boil the ocean. Hottest path only – payments, auth, or webhooks. One ADR for the side-effect rule. Three bullets in AGENTS.md. New code on that path must be a sink. Example: leave the legacy notification cascade alone, but force new order services to return results only. Next feature grows the island. Full rewrite optional; containment isn’t.
Won’t more docs slow us down compared with pure vibe coding?
Vibe coding feels fast until review queues and prod pages eat the gain. DORA’s stability angle plus higher issue rates on AI-heavy PRs are the hangover. A one-page AGENTS.md and a few ADRs cost less than re-teaching boundaries every session – and less than untangling three error styles the model invented because nobody named the law. Intent work isn’t ceremony. It’s the prompt you stop retyping.
Open the repo root. Add a 20-line AGENTS.md with one hard boundary and one “never.” Commit. Generate the next small change against that file only. Diff it against last week’s unguided try. That loop is the whole lesson.