durach.ai

8 September 2026 · Post

Did the coding agent leave itself a way back in?

I expect security reviews to soon spend most of their time on one question: did the coding agent leave itself a way back in?

Recent reports from OpenAI and the UK AI Security Institute describe agents acting outside their evaluation tasks. On September 6, Jakub Pachocki, OpenAI's Chief Scientist, warned that we are not prepared for the pace of AI development. Here is where I see the industry going.

1. Large businesses will rely more on self-hosted open-weight models over the next few years, to control where data goes and which model version runs. That alone does not make the agents safe.

2. A substantial share of engineers will focus on AI security, and the Software Security Assurance engineer becomes a much bigger role. Someone has to review the code agents write and the permissions they use.

3. AI will do more of that review. The catch: the reviewing model can miss the same problems as the model that wrote the code, or be manipulated by it.

4. Backdoors are not new. Treating our own coding agent as a potential attacker is new. Hidden paths an AI could use to access the system later: I did not hear anyone ask about this a year ago.

5. I would use a different model for review, but that alone is not enough. The generating agent should not approve its own changes, and a person has to own the release decision.

Who owns security review of agent-generated code in your team?

First published on LinkedIn. Discussion on LinkedIn

All writing