Securing code written by AI assistants: what to check before it ships

Code from an AI assistant is not secure or insecure because of where it came from. It is a first draft from a contributor who writes quickly, sounds confident and answers for nothing. Make it pass the same review and automated gates as any other change, and add the few checks that assistant-specific mistakes require.

What are the security risks of AI-generated code?

The failures are rarely exotic. Assistants reproduce the patterns they have seen most, and plenty of public code is written for a tutorial, not for production. The volume is what changes: more code, reviewed faster, by people who did not write it.

The myth to drop is that passing tests means the code is safe. Tests check what the author thought to check, and when the assistant writes the tests as well, they share its blind spots.

What should you check before AI-generated code merges?

  1. Packages that exist, and are the right ones. Assistants sometimes suggest libraries that do not exist. Under LLM09 Misinformation, the OWASP Top 10 for LLM Applications 2025 gives the attack as an example: find the package names coding assistants commonly hallucinate, then publish malicious packages under those names. The pattern is widely nicknamed slopsquatting. Check that every new dependency exists, is the package you meant, and has a history you would accept from a human suggestion. A name nobody on the team has heard of is the moment to stop, not to install. Pin versions in a lockfile and review lockfile changes.
  2. Insecure defaults. Look for disabled certificate verification, permissive CORS, queries built by string concatenation, debug modes left on, weak or home-made cryptography, and broad cloud permissions. These are the copy-paste defaults of example code.
  3. Secrets. Generated code often includes placeholder keys that become real ones, and developers paste real credentials into prompts. Run secret scanning on every push, and set a rule about what may go into a prompt.
  4. Gates that cannot be skipped. Static analysis, dependency and secret scanning on every pull request, with branch protection so a failed check blocks the merge. A gate that only warns is a suggestion.
  5. Provenance. Know which assistants are approved, how they are configured, and what they send off your network. Keep a dependency inventory so you can answer where a package came from when it turns out to be bad.

Where should human review effort go?

Reviewers cannot read every generated line with the same care, so spend attention by risk. Authentication, authorization, input handling, anything touching money or personal information, and every new dependency get a full human review. Boilerplate and tests can lean on the automated gates.

Should AI-assisted commits be labelled? Labelling helps less than people expect. What matters is that every change, whoever or whatever drafted it, meets the same review and pipeline gates before it merges. The trade-off is speed: a real gate slows some merges. That cost is smaller than cleaning up one malicious package in production.

What changes when the assistant is an agent?

An agent that opens pull requests, runs commands or edits pipelines is a second risk, separate from the code it writes. Give it the permissions of a new contractor: a working branch, no merge rights, no production credentials. That is covered in AI-generated code security review, with the dependency side in AI supply chain security and the wider list in the OWASP LLM Top 10 for buyers.

How the work is bounded

The scope is agreed in writing before work starts, and the engagement is quoted in writing with it.

SecHB does not issue certifications, attestations or audit opinions: those come from accredited certification bodies, CPA firms and QSAs. The work here is what an organization does to be ready for them.

Nothing here is legal advice. Where a question turns on the law, the work is done alongside the client’s counsel, not instead of them.

Questions we are asked

Is AI-generated code safe to ship?

It is neither safe nor unsafe by origin. It is unreviewed code from a fast contributor with no accountability, and it is as secure as the review and the automated checks it passes through before merge.

How do we catch hallucinated packages?

Check that every new dependency exists, is the package you meant, and has a history you would accept from a human suggestion. A name nobody on the team has heard of is the moment to stop, not to install.

Should AI-assisted commits be labelled?

Labelling helps less than people expect. What matters is that every change, whoever or whatever drafted it, meets the same review and pipeline gates before it merges.

Should we ban coding assistants?

Banning assistants rarely holds, and it tends to push use onto personal accounts where you can see nothing. Approve a tool, configure it, and put the effort into the gates.

Checking your own pipeline

Send a description of your pipeline and the assistants in use. The reply says which gates are missing and which ones only warn. More of our writing is at writing.

Discuss a scope