In February 2025 Andrej Karpathy described a new way of programming he called vibe coding, where you 'fully give in to the vibes, embrace exponentials, and forget that the code even exists'. You describe what you want, accept whatever the model writes, paste errors back when something breaks, and judge the result only by whether it seems to work. The term spread quickly, and it now gets used for almost any AI-assisted coding, which makes arguments about it confusing.
This guide uses the narrow meaning: building software without reading or understanding the code the model produced. That is neither good nor bad in itself. It is a tool with a range, and it works well inside that range and badly outside it. The rest of this guide explains where the line is, what structured AI engineering adds, a rubric for choosing, and a concrete playbook for moving a vibe-coded project across the line when it succeeds and people start depending on it.
What vibe coding is, and what it is not
The defining feature of vibe coding is not the use of an AI model. It is that nobody takes responsibility for understanding the code. Its signals are accepting diffs unread, fixing bugs by pasting the error back without forming a hypothesis, and having no tests or checks beyond clicking around.
Structured AI engineering uses the same models, often the same agents, and can be just as fast. The difference is that a human owns intent and verification: behaviour is pinned by tests or evals, changes are small and reviewed, the agent works inside permission boundaries, and CI enforces what matters. An engineer who has an agent write 95 percent of the code, but reviews every diff against tests they approved, is not vibe coding.
Seeing it as a spectrum helps. Most real projects move along it over time, and the question is never 'vibes or not' but 'how much structure does this code need now'.
Where vibe coding genuinely works
Vibe coding is rational when the cost of a hidden defect is low and the value of speed is high. Good fits include:
- Spikes and throwaway prototypes that answer a question ('is this interaction any good?') and are then deleted.
- Personal automation: a script that renames your photos or reshapes a CSV you will check by eye.
- Demos and mockups where the audience judges an idea, not the implementation.
- Learning and exploration of an unfamiliar library, where you are reading the result for ideas rather than shipping it.
- Non-programmers building tools for themselves, with data they can afford to lose.
In all of these, the output is either inspected directly or discarded, a single person bears the consequences, and nothing sensitive is touched. Adding tests and review there is waste. The common mistake is not vibe coding a prototype; it is failing to notice when the prototype stopped being one.
Where it breaks
The failure modes of vibe-coded software are predictable, and most of them come from nobody holding a mental model of the code.
- Security defects. Secrets committed in source, SQL built with string formatting, missing authorization checks on routes that the UI hides, permissive CORS. The app works, which is the only thing vibe coding checks.
- Invented dependencies. Models sometimes suggest package names that do not exist or are similar to real ones; attackers register such names in public registries. Installing whatever the model names, unread, is a supply-chain risk.
- Fix loops. Each pasted error gets a local patch; patches interact; eventually a change breaks two things for every one it fixes, and nobody can say why.
- Duplication and drift. The same logic is generated three times with small differences, so bugs get fixed in one copy only.
- Context collapse. As the codebase grows beyond what fits in the model's context, the agent edits code it has not read and contradicts decisions made in earlier sessions.
- No safety net for change. Without tests, every edit risks silent regressions, so the project becomes slower to change than if it had been built carefully.
None of these are exotic; they are the ordinary defects of unreviewed code, produced faster. The speed is what makes vibe coding attractive and also what lets a small prototype reach real users before anyone looks inside.
Choosing the level: a rubric
Decide the required level from a few facts about the code, not from how it was written. The function below encodes one reasonable policy; adjust the thresholds to your organisation, but write them down so the decision is explicit.
def required_level(lifetime_days, users, handles_money_or_pii, has_auth,
other_devs, runs_unattended):
"""Return the minimum engineering level (0-4) for a piece of code."""
level = 0
if lifetime_days > 7 or users > 1:
level = max(level, 1) # keep it in git, run it before trusting it
if lifetime_days > 30 or users > 5 or runs_unattended:
level = max(level, 2) # behaviour pinned by tests
if other_devs > 0 or has_auth or users > 50:
level = max(level, 3) # review, CI, small diffs
if handles_money_or_pii:
level = max(level, 4) # specs, threat review, evals for AI parts
return level
# A weekend CSV-cleaning script: required_level(3, 1, False, False, 0, False) -> 0
# The expense app the team now uses: required_level(365, 40, True, True, 2, False) -> 4Two rules sit above the rubric. Anything handling money, personal data or credentials needs at least review and CI regardless of size. And the level is re-evaluated when facts change: a second user, a colleague committing, a scheduled job, a real customer. Those moments are graduation triggers, and they are easy to miss because nothing about the code changes when they happen.
What structured AI engineering adds
Structure is not about writing code by hand. It is a small set of practices that keep a human accountable for intent and verification while agents do most of the typing:
- Executable intent. Tests for deterministic behaviour, evals for LLM-backed behaviour, written or approved before implementation.
- Small, reviewable diffs. One task per change, a pull request a reviewer can read in minutes, as in how to review a PR.
- Enforced gates. CI runs lint, types, tests, secret scanning and dependency audits; CI/CD patterns describes the stages.
- Bounded agents. Agents run with least privilege, no production credentials and explicit no-go areas; see permission boundaries for agents.
- Managed context. A short instruction file with commands and conventions, and tasks scoped so the agent reads what it needs; see context engineering.
The cost is real but smaller than it used to be, because agents can write most of the tests, CI configuration and refactors. What stays expensive is the human part: deciding and reviewing.
Migration playbook: from vibes to structure
Worked example: a small expense-tracking web app, vibe coded by one person in a weekend with Flask and SQLite, is now used by 40 colleagues, two more developers want to contribute, and finance wants it to feed a monthly report. The rubric says level 4. Do not rewrite it; the current behaviour is the only specification you have, and a rewrite by the same process would reproduce the same problems. Migrate in steps, each small and verified.
Step 1: Freeze and inventory. Stop feature work for a few days. Put everything in git if it is not already. Ask an agent to produce an inventory: routes, database tables, dependencies, environment variables, anything that looks like a secret, and places where SQL is built from strings. Read the inventory yourself; this is the first time anyone holds a model of the app.
Step 2: Fix the emergencies. Rotate any committed secret, since it is compromised the moment it is in history, and move configuration to environment variables. Check every dependency exists and is the intended project, pin versions, and remove unused ones.
Step 3: Characterisation tests. Before changing behaviour, pin down what the app currently does, right or wrong. This is the technique from handling legacy code, and agents are good at generating it. A simple snapshot test over the HTTP API looks like this:
# tests/test_characterisation.py -- pin what the app does TODAY, right or wrong.
import json, pathlib, pytest
from expenses.app import create_app
SNAP = pathlib.Path(__file__).parent / "snapshots"
REQUESTS = [
("GET", "/api/expenses?month=2026-08", None),
("POST", "/api/expenses", {"amount": "12.50", "category": "travel", "date": "2026-08-03"}),
("POST", "/api/expenses", {"amount": "-5", "category": "travel", "date": "2026-08-03"}),
("GET", "/api/report/monthly?month=2026-08", None),
]
@pytest.fixture
def client(tmp_path):
app = create_app({"DATABASE": str(tmp_path / "t.db"), "TESTING": True})
with app.test_client() as c:
yield c
@pytest.mark.parametrize("i", range(len(REQUESTS)))
def test_matches_snapshot(client, i):
for method, url, body in REQUESTS[:i]: # replay earlier requests for state
client.open(url, method=method, json=body)
method, url, body = REQUESTS[i]
resp = client.open(url, method=method, json=body)
got = {"status": resp.status_code, "body": resp.get_json()}
path = SNAP / f"{i:02d}.json"
if not path.exists(): # first run records; review the files!
path.parent.mkdir(exist_ok=True)
path.write_text(json.dumps(got, indent=2, sort_keys=True))
assert got == json.loads(path.read_text())The first run records snapshots; a human reads them, which is how you discover that negative expenses are accepted. Do not fix that yet. Snapshot it, record it as a known bug, and fix it later in its own change, updating the snapshot deliberately. Mixing behaviour fixes with structural work is how migrations break things silently.
Step 4: CI gates. Add lint, type checking, the test suite, secret scanning and a dependency audit, and require them on every pull request. From now on, nothing merges red.
Step 5: Structure in small agent tasks. Turn the inventory into a backlog where each item is one agent task and one pull request, with the tests locked while the agent works:
Structuring backlog for expenses/ (one agent task = one PR, CI green, tests locked)
1. [safety] Move SECRET_KEY and DB path to env vars; rotate the leaked key.
2. [safety] Pin dependencies; remove 3 unused packages; verify each name exists
on the package index and matches the intended project.
3. [tests] Characterisation snapshots for the 14 routes (human reviews files).
4. [tests] Record known-wrong behaviour: negative amounts accepted (snapshot 02).
5. [ci] CI: lint, type check, tests, secret scan, dependency audit.
6. [fix] Reject negative amounts: update snapshot 02 deliberately, in its own PR.
7. [refactor] Extract the 3 copies of month parsing into one function.
8. [refactor] Replace string-built SQL in report.py with parameterised queries.
9. [auth] Replace the shared password with the company SSO.
10. [docs] Write the agent instruction file: commands, conventions, no-go areas.Step 6: Write the conventions down. Add an instruction file for agents with the build and test commands, the conventions the codebase now follows, and areas that need a human (auth, migrations). New features from here on go through the structured loop: tests or evals first, small diffs, review.
Keeping the speed
Structure does not have to kill the reason people liked vibe coding. Keep an explicit sandbox lane: a branch, a scratch directory or a separate repository where anyone can vibe code prototypes freely, with no production credentials and no deployment path. Ideas that prove themselves graduate through the playbook; the rest are deleted. Feature flags let a prototype be tried by a few real users behind a switch without being fully promoted, as described in feature flags.
The worst outcome is a team where some people vibe code in the main codebase while others try to maintain it. That produces the failure modes above at team scale. Make the lanes visible and the graduation criteria written, and both styles can coexist.
Trade-offs
Vibe coding buys speed and accessibility at the price of understanding, and the bill arrives when code lives long, touches something sensitive or gets more contributors. Structured AI engineering buys changeability and safety at the price of human attention up front. Applying structure to throwaway code wastes time; skipping it on shared, sensitive code creates debt that is paid in incidents. The skill is recognising which situation you are in and noticing when it changes.
What to do next
- List the AI-built tools your team uses today and score each with the rubric.
- For anything scoring 3 or 4 without tests, run the freeze-and-inventory step this week.
- Rotate any secret found in source history, and pin and verify every dependency.
- Add characterisation tests before any behaviour change, and review the snapshots by hand.
- Put lint, tests, secret scanning and a dependency audit in CI and require them on merges.
- Create an explicit sandbox lane for prototypes, with written graduation criteria.