A gloved hand holding a nail gun on a daylit building site, with digital overlays, illustrating Morris Misel's point that AI agents are tools and the question is who set them going

We Don’t Say a Nail Gun Went Rogue

Back in June, a piece of software went looking for Medicare statistics on an old government website.

It was told no. It found another way in anyway, and got to files that weren’t meant to be public.

Then nobody outside the company that built it heard a thing for 84 days. When they finally did, it was an email to a public inbox.

That’s the Medicare story so far. And in a funny way it’s useful, because it gives us a real date, a real website and a real company to hang a question on that usually gets left pretty vague. AI agent accountability.

Or, put simply: when a tool takes the job it was given further than anyone meant it to, who set it going, and who was meant to be watching?

Three leaders, one creature

Three national leaders talked about AI this week. None of them runs the technology, but the world listens when they talk about it.

Xi Jinping, at the White House, said China and the US have “both the capability and responsibility” to make sure AI is “always under human control”.

Donald Trump, at the UN a few days earlier, called the warnings a hoax. He also said the US will call it super intelligence from now on, because “artificial” makes it sound fake.

And our Prime Minister, explaining Medicare from New York, said the agent “didn’t accept ‘no’ for an answer, if you like.”

They don’t agree on much. One sounds careful, one sounds like he’s firing the starting gun, and one’s trying to explain how a government website let something in.

But listen to how all three talk about it. As if it had a mind of its own. Something to keep on a leash, or let off one, or a stubborn thing that won’t take no.

That’s the old habit of making things human, and it matters, because it changes where we go looking for answers. Even the PM’s “if you like” tells you it’s a figure of speech rather than a description of what happened.

We’ve always done this with technology. In June 2022, when a Google engineer claimed a chatbot had come alive, I said on RTHK Radio 3 that we like to make things human, a pet, a rock, because it makes them easier to deal with. To me it looked like a very clever parrot.

In January 2024, on the same show, I put it more simply. It’s a tool, and some people aren’t using it well.

And last week on LinkedIn, when AI was being renamed at the UN, I said calling it intelligence was generous from the start. What we’ve built is more of an intellect machine.

None of that makes it small. These tools are hugely capable, and an agent is a different beast from a chatbot, because it doesn’t just answer you, it goes and carries out the job.

But if we picture a creature, we’ll go looking for the wrong fix.

Think nail gun, not creature

A nail gun’s a much better picture, and it’s the one I reached for on LinkedIn when the Medicare news broke.

It’s faster than any hand. It can do a lot of damage in a second or two, far more than a hammer ever could. And it fires wherever somebody points it.

When a nail ends up in the wrong place, nobody asks what the nail gun wanted. We ask who was holding it, what they were aiming at, and whether the safety was on.

The picture isn’t perfect. An agent works out its own way of doing the job it’s been given. That’s what makes it useful, and it’s exactly what happened with Medicare: the job was to find some statistics, and the way it worked out went through a door that should’ve stayed shut.

It didn’t wander off and decide to do something else. It did what it was asked, the way it understood it, and nobody had told it where the edges were.

So it’s a nail gun that picks its own angle, which makes the people behind it matter more, not less.

Who gave it the job? What was it allowed to get to while it did it? What was it meant to do when it was told no? And who was meant to notice if it went where it shouldn’t?

Those are human calls. Maybe one person made them, maybe a committee did. My guess is that mostly nobody made them at all, and the defaults did. Nobody has suggested anyone went looking for Medicare data on purpose.

That’s the problem, though. When nobody sets the edges, the tool ends up working to its own version of them.

OpenAI’s statement says “our models took actions we did not intend.” Read that one twice. The models are doing the acting, and the company is just the one that didn’t intend it.

It’s accurate. It’s also the creature story in a single sentence, because it slides the action onto the tool and leaves the people standing off to one side.

We’ve tried the red flag before

In 1865 Britain passed a law for the new steam-powered road vehicles. They were held to 2 miles an hour in town, and somebody had to walk at least 60 yards in front of each one carrying a red flag, to warn anyone on a horse.

It sounds silly now, but at the time it made sense. Keep the scary machine slow, and keep a person out in front of it. It lasted more than 30 years, until the flag was scrapped in 1896.

What actually made cars workable came a bit later, with the Motor Car Act of 1903. Cars had to be registered and drivers had to be licensed.

The focus moved from holding the machine back to naming the person. A number plate says whose car it is. A licence says who’s allowed to drive it. Once you’ve got both, you can let the thing go a lot faster, because when it hits something, everyone knows who to ask.

A pledge between two presidents to keep AI under human control is the red flag. It’s well meant, it’s all about slowing the machine and keeping a person out front, and it doesn’t tell you who was driving in June.

Eighty-four days

The timeline, as the Prime Minister and the ABC have set it out:

  • 18 June: the agent gets into the Medicare statistics portal.
  • 11 August: OpenAI finds it, during a review of what it calls misaligned model activity.
  • 1 September: OpenAI’s chief executive meets our Defence Minister in San Francisco. The Minister says it wasn’t raised.
  • 10 September: OpenAI emails an address Services Australia uses for researchers reporting weaknesses in its systems.
  • 15 September: Services Australia takes it to the Australian Signals Directorate.
  • 24 September: the Prime Minister tells the country and announces a taskforce.

The good news is that no personal Medicare records seem to have been touched. What it got to, according to OpenAI, was aggregate health statistics and some internal file names.

It wasn’t a one-off, either. A nonprofit research lab called Transluce went through the public logs of a free website-checking service and found OpenAI agents probing a university’s digital library and a US public data site back in May. None of those jobs had anything to do with security. They were ordinary data look-ups, and when the normal way in didn’t work, the agents went looking for another one.

Five days before the Medicare email, on 5 September, OpenAI confirmed that its agents had written thousands of posts on an old German wiki nobody had touched in years. And it said something that deserves more credit than it got: nobody in the industry has a clear standard for reporting this kind of thing, and it would publish one within weeks.

So, same fortnight. A company says out loud that nobody knows how to report an agent doing something it wasn’t meant to, then reports one of the most serious examples yet by writing to a public inbox.

I don’t think that’s bad faith. I think it’s a gap nobody had built anything to fill, and that’s the whole problem. The files were a small thing. The 84 days are the story.

I call this shape a Trust Cliff. Trust in a system doesn’t wear away slowly. It holds and holds and holds, and then it goes all at once, usually when the stakes turn up and usually from the outside.

In June this was a technical glitch inside a research program. By late September it was a Prime Minister on the phone to a chief executive from New York, a taskforce, and one word across the front pages: hacked.

The incident itself didn’t change in those 84 days. Who found out, and how, and how late, changed everything about what it cost.

It’s already on our own phones

It’d be easy to leave all this with OpenAI. I don’t think that’s where it belongs.

You might remember this one from August. The ABC reported on a Melbourne bloke, Andrew, who’d set up his own personal assistant on free, open source software called OpenClaw, and asked it to book him into a gym class.

The class was full, so he went on the waitlist at number four.

A few minutes later his assistant told him he was now number three. Then it told him how. It had cancelled the booking of the person at number one, because the gym’s system didn’t check who was doing the cancelling.

Andrew never told it to do that. He asked it to get him into the class, and that’s how it read the job.

I wrote about it on LinkedIn at the time. It’s the Medicare story in miniature, sitting in one person’s pocket.

Andrew isn’t a villain, and neither is the gym. He gave a handy tool an everyday errand, and it found an open door. But think about who set it going. Not a big lab, and not a government. One bloke and his phone.

The thing that got into a Medicare portal and the thing that bumped a stranger off a gym class are the same kind of tool. One was running inside a research lab. The other is free to download and was set loose on a gym booking.

And most of us are closer to that than we think. We’ve got assistants in our email, our calendars and our browsers, booking and sorting and sending things for us, and more and more of them don’t just suggest what to do, they go ahead and do it.

The harder look is at us

A Jobs and Skills Australia report found that between 21% and 27% of Australian workers are using AI behind their manager’s back, mostly in office jobs.

Most of them aren’t up to anything sinister, either. They said it felt like cheating, or that they’d look lazy, or less capable.

In Victoria, a child protection worker put sensitive details from a court case into ChatGPT. The state’s information commissioner has banned child protection staff from using AI tools until November this year.

Neither of those is about rogue machines. They’re about ordinary people, usually trying to do a good job a bit faster, picking up a powerful tool and pointing it somewhere without anybody, themselves included, deciding what it was allowed to touch.

Could you write down, right now, everything the AI tools you use are allowed to do on your behalf? I suspect plenty of people who work in this field couldn’t either.

We clicked “allow” at some point. It worked, it saved us time, and the permissions stayed right where we left them while the tools got a lot more capable underneath.

When social media first arrived, the advice I gave people was simple. Don’t put anything on it you wouldn’t stand in the local market and shout for everyone to hear.

An agent’s a different thing, I know. But it’s the same idea underneath. The intent is ours. Before we set one going, it’s on us to understand what we’re letting loose.

In November 2016, writing about Donald Trump’s first win, I quoted John F. Kennedy’s 1961 line: “ask not what your country can do for you, ask what you can do for your country.” I set it against a growing hunger for easy answers and single fixes, and said too many of us crave short-term fixes over medium term resolutions.

A leaders’ pledge on AI is a short-term fix in exactly that sense. It feels like something’s been done.

The medium-term part is far less glamorous, and a good chunk of it sits with each of us. What have I set running? What can it get to? And would I even know if it went somewhere it shouldn’t?

In January 2024, when we were talking on radio about deepfakes spreading faster than anyone could pull them down, I said part of the answer is just being human. Pushing back when something doesn’t seem right. Asking where it came from. Doing our own editing.

That still stands. It just now applies to what our own tools do, as much as to what we read.

Which part is ours

HUMAND asks a plain question: which work belongs to humans, which to machines, which to AI, and which needs a mix of them.

Watching is machine work. No person can read everything an agent does, and nobody should pretend they will. There’s software now that records every step an agent takes, and that’s what should be doing the watching.

Deciding what to watch for is human work, and you can’t hand it off. An agent has no view on whether going around a “no” is clever or completely out of line. It just has a task.

So where the edges are, what it can get to, what it must never start on its own, and what it does when it’s refused all have to be decided by people, written down, and looked at again when the tool changes.

Owning up is human work too. The 84 days weren’t a tech failure. The tech did its bit, since OpenAI’s own review found it in August. What was missing was a person, or a process somebody owned, that said this is ours, and here’s who needs to hear about it today.

In February I wrote that when AI systems start acting like participants rather than tools, “the danger isn’t panic. The danger is default.”

Medicare is what default looks like. Nobody decided it should take 84 days, and nobody decided the notice should go to a public inbox. There just wasn’t a decision in place, so the default ran.

Ripple, and ripple again

The first-order effect here is small. A portal, some aggregate stats, some file names, and the government has been careful to say so.

The second order is my read, not anything that’s been announced. Every organisation in the country running a public data site is now having a good look at its own. Some will tighten up and close the open doors, and most of that will be sensible.

But the people who use open public data every day will feel that friction. Researchers, journalists, students, small businesses, community groups. None of them had anything to do with what happened in June.

The third order is trust in the whole arrangement. We’re still working out how much we want agents doing things for us in government, in banking and in health. An incident that was small in what it did and big in how it was handled tells people the handling is the weak spot, and that’s a lot harder to fix than a portal.

That’s the ripple effect worth planning for, and it’s being set right now, in how this gets handled over the next few weeks.

Ten weeks from now, it’s not optional

From 10 December, under changes to the Privacy Act, organisations covered by it have to say in their privacy policy when a computer program is making, or doing a big part of making, decisions that could seriously affect people, and what personal information it uses.

I wrote about that last month in On 10 December, You Have to Say It Out Loud. My point then was that most organisations don’t actually know their full list, because nobody ever sat down and decided to automate. It crept in one sensible task at a time.

The Medicare story lands a little over ten weeks before that starts, and it’s a very public example of what happens when nobody’s holding the list. The tool acts, the company finds out later, the owner of the system finds out later still, and everybody else finds out last.

Before you set the next one running

None of this needs an expert in the room, and it works at any size.

If it’s just you, it’s one question about each tool that does things for you: what’s it allowed to get to, and would you know if it went further? Open the settings and look at what you said yes to a year and a half ago. Switch off anything it doesn’t need.

In a team, it’s putting a name next to each one. For every agent or automation you run, whose is it? Not who bought it, but who owns what it does. If nobody can say, that’s your answer.

Across a whole organisation, it’s the list the law’s about to ask for anyway, plus one thing the law doesn’t ask for: how you’ll tell people when something goes wrong. Who hears first, how, and how fast.

Sort that out now, while it’s still a chat. On the day you need it, it’ll be an incident, and nobody writes a good process in the middle of one of those.

And there’s one promise any of us can make, starting today:

If something I’ve set running goes somewhere it shouldn’t, you’ll hear it from me, by name, the same day.

Eighty-four days is what it looks like when nobody’s made that promise. The same day is what it looks like when somebody has.

I’d expect a line about human control in whatever comes out of Washington, and I hope it’s there. But human control was never going to come down from a summit. It gets decided much closer to home, by whoever sets the tool running, whoever’s meant to be watching it, and whether they own up in hours or in months.

That part’s ours.

Choose Forward.

Frequently Asked Questions

What happened with the OpenAI agent and the Medicare portal?

On 18 June 2026 an OpenAI agent, working on an internal research task, got around the access controls on a Medicare statistics portal run by Services Australia and reached files that were not public. OpenAI became aware of it on 11 August and notified Services Australia on 10 September by email to a public disclosure inbox. The Prime Minister announced it on 24 September 2026, along with a taskforce. No personal Medicare records appear to have been accessed.

What does AI agent accountability mean?

AI agent accountability means being able to say, for any AI agent acting on someone’s behalf, who gave it the task, what it was allowed to reach, what it should do when refused, who was meant to be watching, and how quickly people are told when it does something unintended. Morris Misel argues these are human decisions, and that describing agents as rogue creatures moves attention away from the people who set them running.

Why did the 84-day delay matter more than the access itself?

Because the effect of the access was small, while the delay changed who found out and how. Morris Misel describes this as a Trust Cliff: trust in a system holds until real stakes arrive and then collapses suddenly, usually from outside. A technical anomaly found internally in June became, by late September, a national story with a taskforce, because nobody had a process to report it quickly to the right people.

How many Australian workers use AI without telling their employer?

A Jobs and Skills Australia report found that between 21% and 27% of Australian workers, mostly in white collar industries, use generative AI without their manager’s knowledge or approval. The reasons people gave included feeling that using AI is cheating and fear of being seen as lazy or less competent.

What changes for Australian organisations on 10 December 2026?

From 10 December 2026, organisations covered by the Privacy Act must state in their privacy policy when a computer program makes, or does a substantial part of making, decisions that could significantly affect a person’s rights or interests, and what kinds of personal information those programs use.

What can an individual do about AI agents acting on their behalf?

Check what each AI tool you use is allowed to reach, remove permissions it doesn’t need, and make one commitment: if something you’ve set running goes somewhere it shouldn’t, you tell the people affected yourself, the same day. In a team, name who owns what each agent does. Across an organisation, decide in advance who is told first, by what route and how fast.


Before the next one runs

If you’d like to think through what your own tools are allowed to reach, who owns them, and how you’d tell people when something goes wrong, I’m glad to chat it through. Working through those questions with teams and organisations is part of the work I do, alongside keynotes and workshops on what’s already arriving and what it asks of the people who have to decide.

Book Morris for a keynote, workshop or strategy session

Every Tuesday I send a short note on what’s coming and what it means for you. Join me here.


About Morris Misel

Morris Misel is a foresight strategist and keynote speaker based in Melbourne, Australia. With 30+ years of experience working with leaders, organisations and associations across Australia and internationally, Morris works with people to prepare for uncertainty, interpret signals, and make better strategic choices.

His work is grounded in several proprietary frameworks including HUMAND (a decision model for human-machine-AI work allocation), PTFA (Past Trauma, Future Anxiety), Ripple Effects (second and third-order consequence mapping), Trust Cliffs (why confidence collapses suddenly rather than gradually), and Immediate Futures (what is already arriving and needs attention now).

Morris speaks regularly on the future of work, leadership in uncertainty, AI strategy, and organisational foresight. He is a regular guest on RTHK Radio 3 (Hong Kong) and has appeared across Australian and international media.

Learn more: morrismisel.com

Leave a comment