Monday, 28 September 2026

Your Source Code Is Now an AI Data-Boundary Decision

On 21 September, Belgian security company Aikido released an open-weight model designed for cybersecurity work that can run locally, without sending source code to an external model provider. The announcement is one product release, but the underlying question is becoming strategic for every software organisation using AI: where is your code allowed to go?

Engineering leaders have spent years classifying customer data, production credentials and regulated records. Source code has often received less precise treatment. It may be private, but it is routinely copied into local tools, build systems, SaaS scanners, support tickets and now AI assistants. As AI becomes part of development and security workflows, that ambiguity is no longer sustainable.

The decision is not simply whether cloud AI is safe or local AI is better. It is about choosing an execution boundary that matches the sensitivity of the code, the capability required and the controls the organisation can genuinely operate.

Source code is more than intellectual property

A repository can reveal architecture, authentication flows, feature flags, internal endpoints, business rules and the assumptions behind security controls. Infrastructure definitions may expose account structures and network topology. Tests often contain realistic identifiers or payloads. Historical commits can retain secrets long after they have been removed from the current branch.

This does not mean every line of code is equally sensitive. A public design system and a private payments service should not have identical restrictions. The problem is that many organisations have no useful classification between "public" and "private", so engineers must make tool decisions through intuition.

A workable policy should distinguish at least three classes: code already intended for public release; ordinary proprietary application code; and restricted code covering high-risk systems, regulated processing, security controls or commercially critical algorithms. The permitted AI tools, retention rules and approval requirements can then follow the classification.

What local inference genuinely improves

A locally deployed model can reduce several important risks. Prompts and code need not cross into a third-party service. The organisation can control network access, retention, logging and the physical or cloud region in which inference occurs. Security teams can inspect the model artefact, restrict its tools and place it behind existing identity and monitoring controls.

Local operation can also improve predictability. A vendor cannot silently change retention terms, remove a model or introduce a new subprocessor without the organisation noticing. For teams working under contractual or national data-residency constraints, this control may be the difference between using AI and prohibiting it.

There is also an architectural advantage: a local security model can be placed close to repositories and CI systems while receiving only the minimum context needed for a scan. It does not require broad, persistent access to every codebase through a single external integration.

What local inference does not solve

"Runs locally" is not a security guarantee. It moves responsibility.

An open-weight model still has a supply chain. Its weights, runtime, dependencies and container images must be sourced, verified, patched and monitored. A compromised model package or inference framework can be as damaging as an insecure SaaS integration.

The model also needs isolation. If it can read source code, invoke tools, reach the internet and write to repositories, its execution environment matters more than its physical location. Local deployment with excessive permissions can create a powerful internal attack path.

Operational cost is another constraint. Hosting models requires compute capacity, performance tuning, upgrades, observability and someone accountable for availability. A weaker local model that produces noisy findings may consume more engineering time than a well-governed external service. Data sovereignty without useful output is not a successful platform capability.

Finally, local models can still leak information through logs, traces, caches or downstream integrations. The complete data path matters, not the location of the model alone.

The practical answer will often be hybrid

Most organisations do not need one universal AI deployment model. They need routing.

Public and low-sensitivity code can use approved cloud services with contractual controls. Proprietary code can use enterprise services configured for minimal retention, restricted training use and auditable access. Restricted repositories can be routed to local models or excluded from AI workflows entirely until adequate controls exist.

The same principle applies within a task. A local model might inspect raw code and produce a sanitised explanation, which a more capable cloud model then uses for higher-level reasoning. Security scans can run locally, while general documentation work uses managed services. The goal is to minimise sensitive exposure without forcing every workload onto the most expensive control path.

This approach requires a broker or policy layer rather than a collection of individual subscriptions. Engineers should not have to remember which model is permitted for each repository. Classification, identity and enforcement should travel with the code.

A decision framework for engineering leaders

Before approving a model for source-code access, leadership should require clear answers to six questions:

  • Data flow: What exact content leaves the developer machine or build environment, and where does it travel?
  • Retention: Are prompts, code, outputs and telemetry stored, for how long, and for what secondary purposes?
  • Access: Which repositories, branches, secrets and tools can the model reach?
  • Isolation: Can model-generated actions affect production systems or trusted code without human approval?
  • Evidence: How will quality, false positives, security outcomes and developer time saved be measured?
  • Exit: Can the organisation change models or deployment modes without rebuilding the entire workflow?

Staff and Principal Engineers should shape the technical patterns and threat models. Engineering Managers should ensure adoption does not bypass normal ownership and review. Heads of Engineering should establish consistent policy across teams. CTOs should decide which capabilities are strategically worth operating in-house and which are better purchased under strong contractual controls.

The real strategic shift

Local AI is not a rejection of cloud AI. It is evidence that model placement is becoming an ordinary architecture choice, similar to selecting a database region, identity boundary or secrets-management pattern.

The organisations that handle this well will not start with a blanket ban or a race to self-host everything. They will classify their code, map the data flow, constrain model permissions and match deployment choices to actual risk. That creates room for AI adoption without pretending all source code, models and workloads are interchangeable.

The important question is no longer, "Can this model read our code?" It is, "Under which boundary, with which permissions, and with whose accountability should it be allowed to?"


Source note: Reuters reported on 21 September 2026 that Aikido released an open-weight cybersecurity model designed to run locally, reflecting demand for defensive AI that does not require sensitive code to be sent to external providers. Aikido describes its wider static code analysis and CI integration approach on its product site.

No comments:

Post a Comment

Your Source Code Is Now an AI Data-Boundary Decision

On 21 September, Belgian security company Aikido released an open-weight model designed for cybersecurity work that can run locally, without...