"article": "We assumed the perimeter was digital. For two decades, security budgets flowed into firewalls, endpoint detection, and the careful segmentation of network trust. Then Brinks Home disclosed that 4.9 million records had been exfiltrated, and the post-mortem pointed at something far older than the internet: a telephone call. The attacker called, borrowed the voice of an internal employee or a trusted vendor, and talked a credential loose from a human who was simply doing their job. Mandiant's 2025 incident data now ranks vishing — voice-solicited phishing — above email as the primary initial access vector in enterprise environments. CrowdStrike's separate telemetry reports a 442 percent surge in voice-based intrusion activity. The code is law, but the humans are the bug — and the humans are the ones answering the phone.\n\nThe timing is not coincidental. In the same quarter those numbers landed, Google began the broader rollout of \"Let Google Call,\" an AI agent that places real commercial calls, states plainly that it is automated, and expects the business on the other end of the line to engage. We have seen this pattern before: a technology product posing as a technical milestone that is, in reality, a mass deployment of a social experiment. Only this time, the test subjects are not the users who opted in. They are every person who answers a phone, in every business Google's agent has ever dialed.\n\nGoogle Duplex first appeared at I/O in 2018, and for seven years the underlying components have been quietly engineered: text-to-speech that carries prosody, hesitation, and the small imperfections that make a voice feel human; dialogue-state tracking that survives interruptions and mid-sentence corrections; intent recognition that parses the chaos of a live commercial conversation. \"Let Google Call\" is not an architectural breakthrough. It is a productization decision — the moment a laboratory capability is released into the wild as an active social actor. The agent calls a local business, announces it is automated, and proceeds with the task: checking opening hours, confirming a price, scheduling an appointment. By all measurable standards, it is extremely good at being a caller.\n\nThe problem is what the call does to the person on the other end. Every completed AI call is a behavioral data point in a societal experiment whose outcomes the experimenters have not modeled. The receiver engages, the task completes, no harm follows. Repeat that across millions of calls, and the receiver learns a lesson that no security training can easily unteach: an automated voice asking for reasonable things is normal, harmless, and worth complying with.\n\nThe adversary has been studying the same lesson. Microsoft attributes the ShinyHunters campaign to operations against more than one thousand organizations, compromising roughly 1.5 billion records across multiple industries. The campaign's signature is not elegant. It is repetitive, industrial, and effective. Vishing calls were used to funnel targets into OAuth abuse, which produced legitimate-looking access to Salesforce environments at Brinks Home, ADT, and EY. The attackers did not need a technical vulnerability in each victim. They needed to find one person in each organization who, when a confident voice requested a session token or a password reset, did what came naturally.\n\nConsider what this attack pattern actually reveals about the target surface. The Salesforce environments were not penetrated; they were walked into. The OAuth flows that granted access are designed to be easy for authorized users, and that ease is precisely what the attackers abused. Every organization in the ShinyHunters chain was one API token away from compromise, and the token was surrendered, not stolen. This is the defining property of the vishing era: the technical breach is downstream, and the upstream event is always a human being who was talked into an action.\n\nConsumer survey data from the same period sketches the depth of the trust problem. Klaviyo's 2026 AI Trends report found that only 13 percent of consumers fully trust AI; broader platform-trust research puts the share of American consumers who distrust the major technology platforms at 64 percent. And yet the phone keeps ringing, and people keep picking it up. Distrust of platforms does not seem to translate into suspicion of individual callers — because suspicion is a cognitive cost, and the voice on the line is designed to make that cost feel unnecessary.\n\n## The Shared Trust Stack\n\nHere is what makes the Brinks Home breach worth studying, beyond the headlines. The intrusion was not a novel exploit. It was the inevitable yield of a system in which everyone trusts the sound of a voice, and no one verifies who is generating it. And the AI industry, by deploying voice agents into that system, is not a bystander to the problem. It is an active participant in its manufacture — a contribution made, as far as the public record shows, without an equivalent contribution to the infrastructure of verification.\n\nI have spent the past several years inside decentralized finance protocols, where \"don't trust, verify\" is not a slogan but a cryptographic requirement. The phrase has an uncomfortable translation for the voice channel. When I diagram the anatomy of a vishing attack against the anatomy of a legitimate AI voice agent, the overlap is almost total. Four layers, identical in both.\n\nThe first layer is natural-sounding speech. Attackers can rent voice-cloning pipelines that synthesize a specific individual's voice from a few seconds of public audio; the legitimate agent uses the same class of text-to-speech models to sound fluent and calm. The second layer is contextually plausible scripting. A vishing script knows the business is open until six and has a new hire in accounts payable; a commercial AI agent knows the same store's hours and its refund policy. Both draw on harvested or public context to make their questions feel routine. The third layer is an appeal to urgency or normality — a password must be reset before the end of the day; an appointment must be confirmed before the slot expires. The fourth layer is the request for a specific action: send a code, share a file, confirm a change.\n\nListen to the two calls side by side and the symmetry is unnerving. The legitimate agent says: \"I'm an automated assistant calling on behalf of a customer to confirm your opening hours. Could you verify today's closing time?\" The vishing script says: \"I'm from IT. There's been an anomaly with your account. I need you to verify the code sent to your phone so we can secure it before the audit.\" Both are polite. Both are scripted. Both put the receiver in a position where declining feels obstructive. The difference is not in the voice, the grammar, or even the level of urgency. It is in the authorization behind the request — and the authorization is precisely the thing the telephone network does not carry. A listener on a phone has no way to verify a caller's identity beyond the caller's own words, and those words are the one thing both the legitimate agent and the attacker can control perfectly.\n\nThere is an additional, quieter effect that the industry does not measure: call completion rates capture whether the agent got its answer, but they do not capture what the agent left behind. Each successful AI call is an implicit lesson in the receiver's environment — \"machine callers are benign and cooperative\" — and that lesson accumulates in exactly the population the attackers are targeting. The AI industry is performing, at no cost to itself, the market education that vishing operations would otherwise have to fund. Attackers are not running this conditioning themselves; they are harvesting the surplus that legitimate deployments already produced.\n\n## The Disclosure Paradox\n\nThere is a second, darker wrinkle. Google's agent says it is automated, and this is presented as transparency. In some narrow ethical sense, it is — but transparency without verifiability is theater. The best disguise is a half-truth, and once a population is conditioned to hear \"I am a computer\" as the legitimate preamble to a reasonable commercial call, the declaration becomes a costume available to anyone who can synthesize a voice. Nothing binds those words to the entity actually placing the call. The statement lives at the text layer, with zero protocol-layer assurance. An attacker can run a vishing script while claiming to be an AI agent; the receiver, already calibrated to accept automated callers, is no more protected than they were against a human impersonating an IT administrator.\n\nEmail went through the same maturation, and the parallels are instructive. Spam filters and authentication standards — SPF, DKIM, DMARC — were built because the text layer alone could not be trusted. The first generation of phishing emails was blocked by the sheer improbability of a Nigerian prince writing to a random accountant; the second generation was blocked by signature standards. Voice communication is today where email was in the late 1990s: an open channel in which the cost of impersonation is falling toward zero, and no protocol-layer mechanism exists to distinguish the authorized from the malicious. Without a signature standard for AI callers, the phrase \"I am automated\" is not a security control. It is a decoy.\n\nIn the governance systems I help design, this would be laughable. No agent gets to claim authority by declaration. A proposal, a vote, a transfer: each carries a signature, a provenance chain, and a credential resolvable to a known identity. The telephone network has no analogue. STIR/SHAKEN authenticates the originating number at the carrier level — a meaningful countermeasure against number spoofing — but it says nothing about whether an AI is authorized to act on behalf of a business or a user. The identity vacuum at the application layer is the hole through which the vishing industry has climbed.\n\nThe standards gap will not close itself. The telecommunications industry has a template in STIR/SHAKEN, but no one has extended it to the application layer where the agent lives. The Chinese market, which has dealt with high volumes of AI outbound calls for years, imposes labeling requirements on AI calls — but a label, like a spoken disclosure, is only as good as its verification. A regulatory mandate for disclosure without a protocol for verifying the disclosure multiplies the theater. The carrier layer sees the signaling; it authenticates numbers; it has not been asked to authenticate agents. Until that question is settled, the disclosure paradox remains structural rather than incidental.\n\n## Human Conditioning\n\nI want to linger on the training that is happening, because the industry names it incorrectly. The phrase in strategy decks is \"model training.\" The more accurate description, seen from the receiving end, is human conditioning. Every successful AI call is a positive reinforcement event in the receiver's behavioral loop: the phone rings, the voice is calm, the task completes, no harm follows. Extrapolate to a million calls and the reflex becomes reliable — answer, listen, comply. That is precisely the behavioral sequence a vishing attack requires, and the attacker does not need to run a single training epoch on this pattern. The platforms have already executed the conditioning on their behalf.\n\nThe conditioning compounds with a second vulnerability. The weakest link in the authentication chain has always been the voice-based factor. SMS and voice-call multi-factor authentication were designed as secondary proof, but vishing turns the factor itself into the attack surface: the attacker asks the target to \"verify\" the code that the target just received, and the target obliges. OAuth abuse, the mechanism that opened the Salesforce environments at Brinks and EY, depends on this confusion. The industry's response — pushing organizations toward FIDO2 and passkeys — is correct but incomplete, because the weak factor persists wherever legacy systems remain. Every AI call that normalizes the phrase \"please confirm your code\" makes the legacy factor marginally more dangerous.\n\nA painful early lesson from my audit of Curve's governance comes back to me here. I had spent months inside simulation data — hundreds of thousands of lines of voting records — trying to understand how a system that looked democratic could produce outcomes that reinforced a small number of entrenched whales. The answer was that the people with the most power had the least incentive to verify what they voted on, and the mechanism quietly optimized for their convenience. The voice channel runs on the same skew. The parties who benefit most from the call-completion reflex — the platforms deploying agents, the enterprises automating outbound campaigns — have the weakest incentive to build the verification layer that would blunt it. Silence, in governance and in security, is the only consensus that never forks, but ours is a
