The chief financial officer appears on screen. The same voice, the same glasses, the same way of demanding a quick decision. Around him, several colleagues seem to agree. Should you make the transfer? This scene captures the trap posed by deepfakes in video calls: turning a familiar audiovisual presence into proof of identity. Looking ahead to September 2026, the challenge is not just to spot a manipulated image, but to prevent a convincing imitation from being enough to trigger an irreversible action.
Fraud that primarily exploits trust
The case that has become emblematic dates back to 2024. In Hong Kong, an employee made transfers totaling around HK$200 million, or roughly US$25 million, after a video call featuring fake executives and colleagues. Engineering group Arup subsequently confirmed that it was the company targeted. According to details released by police, the meeting used digital imitations of people the employee knew.
This case does not demonstrate that any scammer can fabricate a perfect meeting. It shows something more important: having more faces on screen does not constitute independent verification. If the attacker controls the entire scene, signs of approval are merely window dressing. The known facts establish this risk; the scenarios considered here for September 2026 are forward-looking assessments, not a verified record of incidents as of that date.
The attack generally combines several ingredients: public information about executives, voice samples, videos, a credible pretext and time pressure. An organizational chart, a conference posted online and working habits exposed on social media can provide the raw material. Deepfakes do not replace social engineering: they give it a face.
Why looking more closely is not enough
An unstable mouth outline, imperfect lip synchronization, distorted glasses: these anomalies can raise suspicions. But a poor connection, video compression or a beauty filter can also produce flaws. Conversely, an imitation can look convincing in a small window, especially during a short conversation. The absence of visible anomalies is therefore not proof of authenticity.
Asking someone to turn their head or pass a hand in front of their face can disrupt some systems. However, this is not a universal test. A more capable system can handle these movements; a fraudster can also blame a faulty camera. Even a spontaneous response mainly proves that someone is reacting live, not that they are who they claim to be.
Voice carries the same ambiguity. Accent, rhythm and intonation are familiar clues, but they are not secrets. A personal question can help, provided its answer is neither public nor accessible in a compromised mailbox. Verification must therefore shift to something the attacker does not control.
The right question: what action needs to be authorized?
Not every meeting warrants the same level of scrutiny. A sales presentation and a change of bank details do not expose a company to the same risk. The starting point is to identify sensitive actions: payments, sharing confidential files, resetting access, adding an administrator or changing supplier details.
For these operations, video calls must remain a discussion channel and never become an authorization mechanism on their own. A useful procedure comes down to a few simple rules:
- Call back through an established channel, using a number from the internal directory or supplier record, never from the request received.
- Require confirmation of the transaction in the business application, displaying its amount, its beneficiary or the exact nature of the permissions being granted.
- Require a second, independent approval for high-impact operations.
- Put any urgent exception on hold pending verification, even when it appears to come from management.
A callback is no magic solution: a phone line can also be hijacked, and a voice imitated. Its strength comes from the channel’s independence and its combination with other controls. Two approvals obtained in the same suspicious meeting are not equivalent to two approvals given separately through a controlled workflow.
Authenticate the account, then secure the decision
Multifactor authentication protects against some account takeovers. Security keys and passkeys based on FIDO2/WebAuthn notably offer phishing resistance that manually entered codes do not provide. For sensitive accounts, this protection reduces the risk of a fraudster joining a meeting using a legitimate professional identity.
But three things must be distinguished: the logged-in account, the person using it and the action they authorize. A session can be hijacked; a workstation can be compromised. An authenticated participant badge therefore does not certify that the face being transmitted is real. Final authorization must remain tied to an explicit, traceable and verifiable operation.
Platform settings matter too: limit guests, control admissions, restrict screen sharing and verify unexpected participants. These measures close off entry points without solving audiovisual impersonation on their own. They work better alongside managed devices, regular updates and monitoring for unusual logins.
Detectors: an alarm, not a verdict
Detection tools look for inconsistencies between sound and image, traces of generation or unusual signal characteristics, among other indicators. Their usefulness depends heavily on testing conditions. Performance measured on prepared videos does not guarantee the same results on a compressed, dark or unstable stream, or one produced by an unknown model.
Before buying, a company should ask what types of manipulation are tested, how false positives are measured and what happens when the tool is uncertain. A suspicion score should trigger verification, not an accusation. Conversely, a reassuring result must never allow anyone to bypass dual approval.
Deployment also raises privacy questions: where do the streams go, are they stored, and who can access them? Analyzing faces and voices requires the involvement of security, legal and data protection teams. Cryptographically signed provenance mechanisms can provide clues about the origin of content, but do not automatically prove the identity or sincerity of the person filmed.
Learning to interrupt without accusing
The decisive reflex is organizational: allowing everyone to say, “I’ll verify this through our usual procedure.” Exercises should train people to take that step rather than turn employees into pixel experts. When in doubt, pause the action, contact security through a known channel and preserve the available evidence in accordance with internal rules. Management must accept these checks for themselves.
What now? The plausible scenario for September 2026 is one of more accessible imitations and less reliable visual clues, without every fraud becoming undetectable. The best-prepared companies will not be those that promise to recognize every fake face. They will be those where a video, however convincing, is never enough to move money, disclose a secret or grant critical access.


