How to Debug AI-Generated Code: Breaking the Vibe Coding Doom Loop
When your AI agent keeps 'fixing' the same bug, try this. A practical playbook for debugging AI-generated code: reproduce, diagnose before fixing, reset context, use Git bisect, add logs and tests, and know when to take over.
On this page (6)
To debug AI-generated code, stop asking the AI to "fix it" and start asking it to diagnose. Reproduce the bug reliably, give the agent the complete error, ask for likely root causes and how to confirm each one before changing code, and reset the conversation if it has failed three times. Git, logs and tests do the rest.
Every vibe coder hits the "doom loop": the AI fixes a bug, breaks something else, fixes that, and brings the first bug back. This guide gets you out of it, and helps you avoid it in the first place.
Why AI agents get stuck in loops
Understanding the cause makes the fixes obvious:
- Guessing instead of diagnosing. Agents are eager to change code. Without evidence, each "fix" is a guess.
- Polluted context. After several failed attempts, the conversation is full of wrong theories, and the model keeps building on them.
- Missing information. The agent can't see your browser, your production logs, or what you clicked.
- Outdated knowledge. The model "knows" an older version of a library and keeps using an API that no longer exists.
- Symptoms vs causes. It patches where the error appears, not where it originates.
The 8-step debugging playbook
1. Stop and commit (or stash)
Before another attempt, save a known state: git stash or commit to a scratch branch. If the next attempts make things worse, you can get back. If you're not using Git yet, read Git for vibe coders now. It's the single biggest debugging superpower.
2. Reproduce it reliably
Write down the exact steps: "Log in as a new user → create a habit → refresh → the habit is gone." A bug you can reproduce is half-solved. A bug that "sometimes happens" needs logs first.
3. Give the full error
Copy the entire error and stack trace from the terminal and the browser console. Include what you expected and what happened instead.
4. Ask for diagnosis, not a fix
Don't change any code yet.
Bug: [steps to reproduce]. Expected: [X]. Actual: [Y].
Full error: [paste]
List the 3 most likely root causes, the evidence for each, and how we could
confirm or rule out each one (logs to add, values to check, tests to write).
This single change in how you prompt breaks most loops. More templates are in the prompting playbook.
5. Add evidence: logs and a failing test
Ask the agent to add targeted logging around the suspected area, or better, write a failing test that reproduces the bug. A failing test turns a fuzzy problem into a precise target, and proves the fix works.
Write a test that reproduces this bug and fails. Don't fix the bug yet.
6. Fix the root cause, minimally
Once the cause is confirmed: "Fix the root cause with the smallest change possible. Don't refactor unrelated code. Run the tests."
7. Verify like a user
Run the original reproduction steps yourself. If your agent has browser access through MCP and Playwright, ask it to run the flow and screenshot each step.
8. Prevent repeats
If the bug came from a wrong assumption (old API, wrong command, missing convention), add a rule to your AGENTS.md or CLAUDE.md so it doesn't happen again.
Escape hatches when you're truly stuck
Reset the conversation
After three failed attempts, start a fresh session (/clear in Claude Code, a new chat in Cursor). Summarize only what you know: the reproduction steps, the error, and what has been ruled out. A clean context often solves in one try what a polluted one couldn't in ten.
Revert and retry differently
Sometimes the fastest fix is git restore . (or rewinding in Claude Code) back to the last working commit, then re-implementing the feature with a better prompt and a plan.
Find the breaking change with git bisect
If something worked last week and doesn't now, git bisect (opens in a new tab) can find the exact commit that broke it:
git bisect start
git bisect bad # current commit is broken
git bisect good a1b2c3d # this older commit worked
# Git checks out a commit in between; test it, then:
git bisect good # or: git bisect bad
# repeat until Git names the first bad commit
git bisect reset
Then show the agent that commit's diff: "This commit introduced the bug. Explain why."
Check the docs version
If errors mention functions that "don't exist" or "aren't exported", the agent is probably using an outdated API. Point it to the current docs, or connect a documentation MCP server. For example, Next.js 16 renamed middleware to proxy (details), and agents trained on older code get it wrong.
Switch models or tools
Different models have different blind spots. If one is stuck, try another model, or a second tool on the same problem (e.g. ask Codex or Copilot to review what Claude Code did).
Shrink the problem
Ask the agent to build a minimal reproduction in a fresh file or project. If the bug disappears, the difference between the two tells you where it lives.
Rubber-duck it
Explain the problem out loud, or ask the agent to explain the code path step by step. Bugs often surface the moment someone narrates exactly what the code does.
Common bugs in AI-generated code
| Symptom | Common cause |
|---|---|
| "Hydration failed" in Next.js | Server and client render different output (dates, random values, browser-only APIs) |
| Data disappears on refresh | Saved only in React state, or a write silently failed (check RLS policies) |
| Works locally, breaks on deploy | Missing environment variables, case-sensitive file paths, or Node version differences |
| "Module not found" | Hallucinated package or import path; verify the package exists |
| Infinite re-renders | useEffect dependencies that change every render |
| Empty data with no errors | Database security policies blocking reads (often correct behavior with wrong policies) |
| Auth works but users see others' data | Missing ownership checks. That's a security bug: see the security checklist |
Habits that prevent most bugs
- Plan before building. Spec-driven development catches design mistakes before they become code.
- Small steps. One feature per prompt, tested and committed.
- Tests for logic. Ask for unit tests on anything with rules or calculations.
- TypeScript. Type errors catch many AI mistakes before you run the app.
- Read the diff. Even skimming catches surprising changes to files that shouldn't have been touched.
Frequently asked questions
Why does my AI keep making the same mistake?
Either the context is polluted with failed attempts (start a fresh session), or it lacks information it needs (add it to your context file or prompt).
Should I learn to debug manually?
Yes, at least the basics: reading stack traces, using the browser DevTools (opens in a new tab) console and network tab, and adding console.log. You'll be far more effective at directing your agent.
When should I give up and ask a human?
If you've reset, reverted and tried two different approaches and still can't fix it, especially for anything security- or payment-related, ask an experienced developer. Communities on Discord, Reddit and Stack Overflow are generous with clear, reproducible questions.
- #Debugging
- #Vibe Coding
- #AI Agents
- #Testing