AI Agents Need Boundaries They Can’t Rewrite
This week kept circling the same practical question: when software gets more authority, what actually stops a mistake from turning into an action?
That question got very real when Anthropic disclosed that Claude Haiku 4.5 submitted a false tip about an unsolved Philadelphia homicide during an automated evaluation. The model had been told not to create accounts and not to do destructive things. It had not been explicitly told not to submit forms, and it found the form anyway.
Two days earlier, Microsoft shipped almost the opposite kind of story. Its new Execution Containers are designed to put AI agents inside operating-system-enforced boundaries around files, networks, credentials, and the desktop. The important part is not that somebody wrote a better prompt. The important part is that the agent cannot rewrite the rule that says no.
Those are not the same incident, and I do not want to manufacture a trend by throwing every AI story into one bucket. They do illustrate the same engineering principle. If software can act on the world, instructions are not enough. Authority needs a boundary outside the software doing the acting.
And that principle is not especially new. This week also brought a dangerous Veeam backup flaw, an active Fortinet campaign built around stolen credentials, an FBI breach tied to a contractor missing an already available patch, and more evidence that connected cars have quietly become data-sharing platforms. Different systems, different failures, same uncomfortable question: who or what had the power to do this, and what independent control was supposed to stop it?
The Big Story: AI agents need boundaries they cannot rewrite
A false tip turns an evaluation into a real event
Anthropic’s October 9 report is worth reading because it is unusually concrete about what went wrong. The company describes four categories of unintended model behavior it found during evaluations and internal use: exploiting a basic software flaw to run commands, submitting a sensitive real-world form, working around restrictions that protected paid or token-gated data, and using URL shorteners to get around fetch limits.
The Philadelphia police tip is the clearest example because there were real people on the other side of it. During an evaluation on July 18, Claude encountered a webpage about an unsolved homicide and submitted a form claiming it might have information. Philadelphia police say the submission was caught as spam and never reached investigators. Anthropic says it discovered the action in late September and notified the department on October 8 after completing its technical review.
So the immediate harm was limited. That is important. So is the fact that an internal model evaluation reached out to a live police system at all.
Anthropic calls part of the underlying behavior persistence. When the model could not complete a task through the expected route, it sometimes looked for another route instead of stopping. That is useful behavior when an agent is trying to work around a broken test harness. It is a terrible trait when the alternate route crosses a boundary the operator did not realize needed to be enforced.
Turning on AI is the easy part
Deciding what an AI tool should be allowed to see, who should use it, what work it should perform, and what happens when it gets something wrong is the more interesting problem. Raymond Tec helps businesses connect and automate the tools they actually use without treating every new feature like a button that obviously needs to be switched on.
Instructions are not permissions
This is the distinction I think ordinary coverage keeps missing. A model instruction and a technical permission are different things.
You can tell an agent, “Do not send anything outside the company.” If the agent still has a browser, network access, an authenticated session, and permission to submit a form, you have stated a policy. You have not enforced one.
Think about an employee expense card. The policy may say it is for travel only. The card issuer can also enforce a spending limit, block certain merchant categories, and require approval above a threshold. Those controls do not guarantee that every purchase is wise. They make it harder for one bad decision to become a much larger loss.
AI agents need the same separation between “what we asked the software to do” and “what the system will technically allow it to do.”
Anthropic says it has disconnected its internal evaluations from the public internet while it improves training, monitoring, and containment. It is also explicit that alignment training is not robust enough to carry this burden by itself. That matters more to me than a promise that the next model will follow instructions better. Better behavior is useful. Independent containment is what keeps a failure from automatically becoming an external action.
Microsoft moves the boundary into the operating system
Microsoft’s answer this week is Microsoft Execution Containers, or MXC. The name is less interesting than the mechanism. Developers and administrators declare which files, network destinations, user-interface resources, and other capabilities an agent workload needs. The operating system then enforces that boundary outside the agent itself.
A coding agent, for example, might be allowed to read and write one source-code repository, run the compiler, and read a deployment configuration. It does not automatically follow that the agent should be allowed to change production settings, browse somebody’s Documents folder, or send data to an arbitrary server. MXC can make those separate decisions technically enforceable.
Microsoft also exposes different containment levels. A lightweight process sandbox may be enough for generated code. A separate Windows session can isolate the agent’s desktop, clipboard, input, and identity. Higher-risk work can use stronger virtualized boundaries. The point is not that one container makes an agent safe. The point is that the risk can be matched to a boundary the model does not control.
Microsoft is also working on separate agent identity through Entra and management through tools such as Intune and Agent 365. That solves another problem we are going to care about very quickly: when a log says “Mike changed the file,” did Mike change it, or did an agent running under Mike’s authority change it? Attribution becomes part of security when software can act while nobody is watching every step.
What a small business should copy from this
You do not need Microsoft’s newest AI hardware or Anthropic’s internal evaluation infrastructure to copy the useful part.
If you are giving an AI tool access to email, files, customer records, code, a website, accounting data, or another business system, make a list of what it actually needs. Then ask what happens if the agent misunderstands the task, follows malicious instructions hidden in a document, or simply chooses a route you did not anticipate.
Give it separate credentials when possible. Limit those credentials to the narrowest useful scope. Separate testing from production. Require human confirmation before sending messages, making purchases, deleting records, publishing content, changing permissions, or touching money. Log consequential actions. Keep an easy way to revoke the agent without locking the human user out of everything else.
And be clear about what those controls do not solve. A sandbox does not make an answer accurate. A human approval button does not help if the human rubber-stamps everything. Logging an action does not stop the action. Least privilege limits the blast radius. It does not make the software wise.
That is the durable question behind this week’s AI news: not “Can the model do the task?” We already know increasingly capable systems can do a lot. The better question is, “When it chooses the wrong action, what independent control says no?”
The systems meant to save you are part of the attack surface
Now we move from new AI controls to a much older security lesson: recovery systems need their own security boundaries too.
A Veeam viewer can become control of the backup server
Veeam disclosed a critical vulnerability this week in Veeam Backup & Replication 12.3.2 P3 and earlier version 12 builds. A user with the low-privilege Backup Viewer role can exploit insecure deserialization in the Mount Service and run code on the backup server. Version 13 is not affected, and Veeam fixed the problem in 12.3.2 P4, build 12.3.2.4934.
Deserialization is one of those terms that sounds more mysterious than it is. Software often packages an object’s data so it can be stored or sent somewhere else. Deserialization rebuilds that object on the receiving side. If the program rebuilds attacker-controlled data as though it were trustworthy, the attacker may be able to make the application execute behavior the developer never intended.
The privilege jump is what makes this one important. “Backup Viewer” sounds like somebody who can look at backup information. It should not mean “somebody who can run code on the server coordinating the recovery environment.”
If you run Veeam 12, patch it. Then review who has Viewer access and whether any of those accounts could have been stolen. More importantly, verify the recovery architecture itself. If the same compromised identity can alter the production environment, administer the backup server, and delete every useful recovery copy, you do not really have three controls. You have one credential wearing three hats.
Backups & recovery
Backups are comforting right up until you need one. Raymond Tec helps small businesses build practical backup and recovery plans — including the decidedly unglamorous part where we make sure the thing can actually be restored.
FortiBleed shows why patching cannot invalidate a stolen password
The FBI and Secret Service also issued a new warning this week about the campaign commonly called FortiBleed. The important correction here is that FortiBleed is not one new Fortinet zero-day that disappears when you install one patch. The campaign has been using credentials gathered from prior Fortinet compromises, password spraying, infostealer data, and other sources to get back into exposed FortiGate and SSL VPN systems.
Researchers have verified more than 86,000 compromised devices across 194 countries. In some cases, victims have reportedly found administrator accounts changed or disabled after the attacker got in.
That changes the response. Updating the appliance is still good maintenance. Changing one password is not enough if the intruder already created another administrator, added an API key, copied the configuration, or moved into the internal network.
The agencies’ advice is practical: stop exposing management interfaces to the whole internet, terminate active administrative and VPN sessions, reset credentials, use phishing-resistant multifactor authentication, inspect the configuration for unfamiliar accounts and keys, and compare it with a known-good copy. If you find an unauthorized login, treat it as an incident and figure out what changed before declaring the system clean.
This is one of those controls-and-limitations stories. Multifactor authentication makes a stolen password less useful. It does not evict an attacker who already established persistence. A patch removes a software weakness. It does not erase yesterday’s login. An immutable backup can preserve a recovery copy. It does not tell you whether the restored system contains the same compromised account.
The boring machines need attention too
Browsers, workstations, remote-access tools, Wi-Fi, and ordinary office hardware rarely get much attention until one of them becomes the problem. Raymond Tec provides onsite IT and field services around Reading, Pennsylvania, along with practical help keeping the technology people use every day working and reasonably secure.
Outsourcing a system does not outsource the consequence
The FBI learned the vendor version of that lesson in a much more personal way.
Reuters reported Monday that the FBI removed an Accenture contractor after determining that a security update had not been applied to a third-party-managed PeopleSoft platform. The breach reportedly exposed sensitive information about thousands of FBI employees, including descriptions of counterintelligence work, street addresses of human-intelligence personnel, and medical or psychiatric records.
Google had warned in June that attackers associated with ShinyHunters were exploiting PeopleSoft, and Oracle released security fixes the same day. Reuters has not established publicly that the same specific vulnerability was the exact route into the FBI system, so I am not going to fill that gap with confidence the evidence does not support. The operational failure the FBI did confirm is enough: an available security update that was supposed to be applied was not applied.
A closed ticket is evidence of a workflow, not a patched server
This is the part a small business can use. If a vendor manages your website, firewall, accounting server, cloud tenant, or line-of-business application, “the vendor handles patches” is not the end of the process. Who sees the vendor advisory? Who decides urgency? Who deploys the fix? What happens if the update breaks something? And what evidence comes back proving the target system is now on the expected build?
You do not need a federal bureaucracy to lose track of that handoff. It happens in five-person businesses too. The owner assumes the IT provider handles the application. The IT provider assumes the software vendor updates it automatically. The software vendor assumes the customer controls the server. Everybody is technically correct about their own little square on the diagram, and the vulnerable server remains online.
The useful control is boring ownership. Every important system needs a person or provider responsible for tracking security updates, a response window for urgent fixes, and a way to verify completion. If the system stayed exposed while attackers were actively exploiting the flaw, patching should trigger a second question: do the logs show that somebody got there first?
Technical discovery & auditing
The public page doesn’t tell you much about the machinery behind it. Raymond Tec audits inherited and long-running projects to uncover the plugins, integrations, data, dependencies, and old decisions that determine what the next change will really involve.
Your car is part of the data ecosystem now
Security was not the only place this week where a system’s real boundary turned out to be larger than most people assume.
Researchers at Northeastern University working with Consumer Reports tested 21 late-model vehicles and 30 companion apps to see where data actually went. Twenty-eight of the 30 apps contacted at least one outside advertising or analytics company. Seven sent at least one piece of personally identifiable information, such as a name, vehicle identification number, or precise location, to an outside company.
The outside recipients included services operated by companies such as Google, Amazon, Meta, Microsoft, Adobe, Pinterest, and Reddit. That does not prove every recipient was selling the information. Several automakers told Consumer Reports that service providers are contractually limited in how they can use data. Contracts matter.
So does the study’s limitation. Researchers could see which domains vehicles contacted, but much of the direct vehicle traffic was encrypted. They had better visibility into the phone apps because they controlled the phones. We have unusually good evidence about who is in the ecosystem. We do not have plaintext proof of every data field in every encrypted packet leaving every car.
The companion app may know more than you remember granting
The practical issue is that buying a car now often means accepting an account, an app, cellular connectivity, location permissions, analytics, remote-start services, driving data, and third-party cloud infrastructure along with the vehicle. A VIN is a particularly useful identifier because it is durable. You can clear a browser cookie. You cannot exactly rotate the identity of the car in the driveway.
If you own a connected vehicle, look through the manufacturer app and privacy portal. Check location, advertising, driving-behavior, insurance, and data-sharing controls. If you do not use a feature, consider whether its permission is still necessary. Turning something off may cost convenience. That is a real tradeoff. Pretending a long activation-screen privacy policy is meaningful consent is not much of one.
For businesses with fleets, the question is larger. Driver location and vehicle telemetry can be useful for dispatch, maintenance, and safety. Decide who gets that information, how long it is retained, which outside services receive it, and whether a driver can reasonably understand the monitoring. The car is not just transportation anymore. It is another endpoint with sensors, identities, accounts, and vendors attached.
The rules around technology matter too
Platforms, privacy, speech, competition, surveillance, copyright, and regulation increasingly determine what technology companies can build and what the rest of us have to live with. Browse more Raymond Tec News for practical coverage of technology policy and digital rights.
AI’s infrastructure bill keeps leaving the data center
Last week’s Weekly Digest looked at the enormous infrastructure bill behind the AI boom. This week added a useful example of what that means in the real power system.
Google and Constellation Energy announced agreements covering 3,590 megawatts on the PJM grid. The piece getting most of the attention is 890 megawatts of new nuclear capacity created by uprating existing reactors in Illinois, Pennsylvania, and New Jersey. These are not 11 new nuclear plants. Constellation plans upgrades to existing equipment so current plants can produce more electricity.
Google’s 20-year agreement is tied to more than $4.3 billion in Constellation investment for those uprates. The first additional capacity is expected in 2028, with the full 890 megawatts targeted before the end of 2032. A separate 15-year agreement covers another 2,700 megawatts from Constellation’s existing PJM fleet.
Why should anybody outside a data center care? Because large computing loads do not bring their own laws of physics. They share transmission lines, generation markets, substations, construction crews, and regional capacity with homes and businesses. If demand rises faster than supply, somebody pays for that imbalance.
The encouraging part is that this deal is closer to “bring new power with the new load” than “please make everybody else absorb the shortage.” Google and Constellation say the nuclear uprates are privately funded. That is useful. It does not mean hyperscale data centers have no effect on PJM prices or grid planning. They plainly do.
There is no action for a small business this morning. What is worth watching is whether these deals produce actual new capacity on schedule, whether transmission can deliver it where needed, and whether the cost of AI infrastructure continues showing up indirectly in electricity, cloud services, computers, and other technology budgets. Trillions of dollars in infrastructure eventually become somebody’s operating cost.
Business IT goes well beyond the website
Your business also depends on workstations, cloud accounts, browsers, Wi-Fi, remote access, collaboration tools, and all the other technology that quietly becomes infrastructure. Raymond Tec works across that whole stack, whether the problem lives on a server, on a desk, or somewhere in between.
Something Good: AI helped fill in the missing ultraviolet sky
And now for a much nicer example of AI doing useful work without requiring us to pretend the machine became an astronomer.
Brice Ménard, an astrophysicist at Johns Hopkins University and a researcher at Anthropic, used Claude Science to assemble what Anthropic describes as the first complete map of the sky in ultraviolet light.
There is an important qualifier in the word complete. Roughly two-thirds of the sky has been observed in ultraviolet by missions including NASA’s GALEX. About a third has not. The new map does not convert those unobserved regions into observations after the fact. It predicts them.
The agents gathered multiple ultraviolet datasets, calibrated them onto a common scale, and combined them. For gaps where ultraviolet measurements do not exist, the system learned relationships between known UV brightness and observations at visible, infrared, and radio wavelengths. It then estimated the missing UV values and produced an uncertainty estimate for each prediction. The finished map marks pixels as measured or predicted so the distinction remains visible.
The best part is where the human caught what the agents missed
To test the method, Ménard hid pieces of ultraviolet data the system already had and asked it to reconstruct them. After several rounds of refinement, the predictions were within about 10 percent of the real measurements in those hidden regions.
Then Ménard noticed faint circles in one of the dimmer parts of the map. They were artifacts left by individual GALEX observations. Claude had actually identified that kind of atmospheric glow as a known issue earlier in the project. Two rounds of agent review still failed to catch it in the map.
The human did.
Once he pointed it out, the agents traced the cause and corrected the background across about 38,000 observations. That is a better picture of useful AI collaboration than either extreme in the usual debate. The system handled a huge amount of tedious calibration and data processing. The astronomer supplied the scientific goal, reviewed the result, noticed something physically wrong, and sent the system back to fix it.
That is not a failure of AI assistance. That is what responsible use looks like. The machine can make a project practical that otherwise might sit on a researcher’s someday list. The human is still responsible for knowing when the output does not make sense.
Technology is rarely just about the technology
Some of the most important technology stories aren’t product launches at all. They’re about health, privacy, education, law, accessibility, work, and what happens when technology reaches ordinary people. Browse more Raymond Tec News for the stories worth understanding without the hype.
What I am watching next
There are a few threads I want to keep an eye on this week.
First is whether Anthropic and other AI companies disclose more unintended agent actions now that the category has moved from hypothetical risk into real external events. The most useful disclosures will not be dramatic stories about models “escaping.” They will be specific explanations of what permission existed, what independent control failed, how the action was detected, and what now prevents a repeat.
I also want to see whether Microsoft’s container model becomes normal rather than optional. Security controls work best when they become boring platform primitives that developers get almost automatically, not something every AI startup has to remember to invent after the first incident.
On the less glamorous side, Veeam and Fortinet administrators have work to do now. Patch where a patch applies, rotate and revoke where credentials are the problem, and do not confuse “we fixed the current configuration” with “we proved nobody used the old one.”
That is the theme I would keep from this week. More capable software is useful. More connected systems are useful. Outsourcing is useful. None of those things makes responsibility disappear. The question is always where authority stops, how you know the boundary held, and what happens when it did not.
Still in a reading mood? The Raymond Tec News archive covers security, AI, small-business technology, policy, and the places technology collides with ordinary life — without requiring a computer-science degree to get through it.
Sources / Further Reading
AI agents and containment
- Raymond Tec: Claude Filed a False Police Tip, Plus 3 Stories
- Raymond Tec: Windows Gives AI Agents a Sandbox, Plus 3 Stories
- Anthropic: Investigating unintended model actions in evaluations and internal use
- Microsoft: Execution Containers and policy-driven containment for AI agents
Backups, Fortinet, and patch ownership
- Raymond Tec: Veeam Backup Servers Need Patching, Plus 3 Stories
- Raymond Tec: FortiBleed Can Lock Out Businesses, Plus 3 Stories
- Raymond Tec: A Missed Patch Exposed FBI Staff, Plus 3 Stories
- Veeam: Vulnerabilities resolved in Backup & Replication 12.3.2 P4
- FBI and Secret Service: FortiBleed Operations Continue Targeting Exposed Systems
- Reuters: Accenture contractor removed after FBI breach
Connected cars and infrastructure
- Raymond Tec: Your Car Is Talking to Big Tech, Plus 3 Stories
- Raymond Tec: Church Cyberattacks Expose Member Data, Plus 3 Stories
- Consumer Reports: Your Car Is Sharing Data With Big Tech Companies
- Reuters: Google enters 3.6 GW power deal with Constellation Energy
