Vibe coding can be useful for low-risk experiments. It is not a safe shortcut for every business system. The right level of human control depends on the data, users and harm involved.
A private mock-up is one thing. Bespoke software development with logins and payment data is another. Treating them the same creates avoidable risk.
The National Cyber Security Centre now describes AI-assisted coding as a spectrum. Its advice is simple. Different code needs different levels of oversight. That gives business owners a better question to ask. Do not ask whether AI wrote the code. Ask what checks the code received before it went live.
What vibe coding means
Vibe coding means directing an AI tool with plain-language prompts while it writes much of the software. At the fullest end, the AI may choose the structure, write the code and create the tests. The person guides the result by asking for changes.
There are many steps between hand-written code and full vibe coding. A developer might use AI to finish a line. They might ask it to draft one function. They might let it build a whole feature, then inspect every change.
AI researcher Andrej Karpathy gave the approach its name in 2025. In full vibe coding, a person describes an app in natural language. An AI agent then writes and changes the application from those prompts. Other AI coding tools give developers much less help. They may only suggest the next line while a developer keeps control of the wider program.
Those methods should not share one label. They give the AI different amounts of control. They also leave the human with different amounts of knowledge.
The key issue is ownership. Someone still needs to understand what the software does. They need to know how it can fail. They also need to fix it when the result is wrong.
What changed in 2026
On 18 June 2026, the NCSC published a practical view of AI-assisted software work. It called this the vibe coding spectrum.
The NCSC did not ban vibe coding. It said the level of care should match the risk. A proof of concept may suit broad AI use. Public login code or software that handles sensitive data needs much more rigour.
This matters because AI output can look convincing. Code may run and pass a simple test while still holding a security flaw. A tidy screen is not evidence of a safe system.
One research paper shows why care is needed. The study used 200 feature requests from real open-source projects. It tested one setup using SWE-Agent with Claude 4 Sonnet. In that setup, 61% of the solutions were functionally correct. Only 10.5% were secure.
The paper was revised in February 2026. Its early attempts to add vulnerability hints did not remove the security problem. Better prompts alone were not enough in that test. The full paper is available through arXiv.
The NCSC risk spectrum
The NCSC spectrum starts with human coding supported by tools such as autocomplete. It ends with AI taking control of structure, code and tests. Most real work sits somewhere in the middle.
Risk should decide where a project sits. The amount of AI use is only one part. The quality of review matters just as much.
| Use | Risk | Sensible AI role | Evidence before launch |
|---|---|---|---|
| Throwaway mock-up with fake data | Low | AI can build most of it | Confirm it is private, uses no real data and will not be reused without review |
| Internal tool with limited access | Low to medium | AI can draft features | Code review, access checks, tested calculations and a clear owner |
| Public feature with no account data | Medium | AI can assist under review | Security, accessibility, browser and failure tests |
| Integration moving customer or order data | High | AI can support an experienced developer | Threat model, access design, error recovery, logs and negative tests |
| Login, payments, health, safety or sensitive data | Very high | AI has a limited, supervised role | Senior review, security testing, monitoring, rollback and incident plans |
This is a decision aid, not a formal security rating. A small internal tool can still cause harm. A public feature can be low risk if it stores nothing and controls nothing. Classify the real use before choosing the process.
Three realistic business scenarios
1. A private stock calculator
A warehouse manager wants to test a better reorder rule. The first version uses made-up stock figures. It has no customer details and cannot place an order.
Full vibe coding may be reasonable for that first mock-up. The manager can test the idea without treating the result as finished software.
The risk changes when the tool connects to live stock. It rises again if the tool can create purchase orders. At that point, a developer should inspect the code and calculations. They should test bad data, duplicate actions and failed connections. The prototype has become an operational system.
2. A customer portal login
A company wants customers to see documents, invoices and support requests. The visible page may look simple. The work behind it is not.
The system must prove who the user is. It must limit each account to the right records. Password reset, session expiry and failed login controls all matter. Logs must show suspicious access without exposing private data.
This is not a good place for unchecked vibe coding. AI may help draft a small part. An experienced developer still needs to own the design. The build needs access tests and clear recovery steps. Our guide to customer portal software explains the wider system choices.
3. An order link between two systems
A wholesaler wants online orders copied into its stock system. AI can help map fields or create a first draft of the connection. That can save routine typing during development.
The hard questions remain human ones. What happens when one service is down? Can the same order arrive twice? Which system owns the final status? Who can change the connection key?
A sound API integration needs safe access, useful logs and tested retries. It also needs a plan for changes made by either supplier. Working once is not enough.
What proper human oversight looks like
Human oversight is more than reading the finished screen. The reviewer needs enough skill to challenge the design and trace important paths through the code.
First, the team should map the data and users. That reveals where access checks belong. It also shows what an attacker or simple mistake could expose. A clear software specification records those rules before code is written.
Next, the reviewer should inspect the software structure. AI can create extra packages, hidden dependencies or repeated code. Each dependency needs a reason to exist. Secrets such as passwords and keys must stay outside the code.
Tests then need to cover failure as well as success. A login test should try the wrong user and expired session. An order link should test duplicate messages and missing fields. A file upload should reject unsafe types and oversized files. Our UAT guide shows how business users can record the result.
High-risk work may also need an independent security test. The person who built a feature can miss the same assumption twice. A fresh reviewer has a different view.
The NCSC's wider secure AI system guidance covers design, development, deployment and operation. That full life cycle matters. Safe code can become unsafe through weak setup or poor maintenance.
Vibe coding myths and facts
Myth: AI-written code is always insecure
Fact: AI can produce useful and secure code. The output still needs checks suited to its use. The tool name does not prove safety either way.
Myth: Passing tests means the software is safe
Fact: Tests only prove what they check. AI may write tests that confirm the happy path while missing misuse, bad data and access faults.
Myth: The AI provider owns the result
Fact: The business and software supplier still need clear ownership. Someone must approve changes, respond to faults and keep the system patched.
Myth: Any human review is enough
Fact: Review quality depends on skill and context. A quick glance cannot test access rules, data handling or failure recovery.
What to ask a software supplier
Asking whether a supplier uses AI will not tell you much. Most coding tools now include some form of AI help. Ask how the supplier controls it.
Start with the work itself. Which parts did AI write? Did it suggest a few lines, draft whole features or choose the system structure? The answer should match the risk of the feature.
Then ask who reviewed the output. A named technical owner is better than a vague promise that the team checked it. That person should be able to explain the code without returning to the AI tool.
Ask to see the test approach. A list of passed tests is useful, but it needs context. Did the team test failed logins and wrong permissions? Did it try missing fields, duplicate requests and lost connections? These cases often reveal more than the normal path.
Find out what was sent to the AI provider. Live customer records, passwords and private business code should not be pasted into a tool without clear controls. The supplier should know the tool's data terms and account settings.
Maintenance matters too. Ask whether another developer can work on the code. Check that the project has useful notes, version history and a list of dependencies. A fast first build loses its value if nobody can safely change it.
Good answers will be specific. They will name the reviewer, tests and release controls. A supplier should not treat the AI brand as proof of quality.
A launch checklist for AI-assisted software
Use this list before a tool reaches real users or live data. A high-risk system will need more evidence.
- Name the person who owns the code and launch decision.
- List the data used, stored and shared.
- Mark any personal, financial or commercially sensitive data.
- Map every user role and the records each role can access.
- Record which code and tests came from AI.
- Check packages, licences and third-party services.
- Remove passwords, keys and tokens from the source code.
- Review the design against likely threats and mistakes.
- Test wrong inputs, failed services and duplicate actions.
- Test backup, restore and rollback steps.
- Add useful logs and alerts without logging private data.
- Agree who will patch dependencies and fix later faults.
- Arrange independent review where the harm could be serious.
- Keep a human approval point before each live release.
If a supplier cannot show this evidence, a polished demo should not settle the decision. Ask who understands the finished code. Ask how they tested access and failure. Ask what happens after launch.
Should your business use vibe coding?
Use it where speed of learning matters and the cost of failure is low. A disposable mock-up can answer a useful question. It should not quietly become the live system.
For business software, AI works best as a capable assistant under clear ownership. It can draft routine code and help explore options. It cannot accept responsibility for customer data, downtime or a bad release.
The best process is risk-led. Start with the users, data and possible harm. Then choose how much AI freedom makes sense. Build the review and test plan at the same time.
If you are weighing an AI-built prototype against a supported system, see how we approach bespoke software development. You can also show us the current process and the result you need.
Source notes
- NCSC, The vibe coding spectrum approach to AI-assisted software development, published 18 June 2026.
- Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-generated Code in Real-world Tasks, first published 2 December 2025 and revised 16 February 2026.
- NCSC, Guidelines for secure AI system development.