ailiteracynepal 🇳🇵
Text size

Chapter 05 · Section I · 15 min read

Deepfakes and synthetic content

Cheap, fast, convincing — synthetic video, audio and images are no longer a frontier problem, and Nepal will feel the impact at street level before any institution is ready.

For most of the history of photography, a moving image of a person saying something was reasonably strong evidence that the person had said it. That assumption — quietly, in the space of about three years — has stopped being safe. In 2026 a competent synthetic video can be produced from a single reference photo and a thirty-second voice sample, on a mid-range laptop, by someone with no special training. The barrier is no longer technical skill or expensive equipment; it is whether the person making the fake can be bothered. Increasingly, they can.

Where the technology actually is

Three capabilities have matured in parallel, and it is the combination that matters.

Video. Open-source diffusion models can now generate short clips of a recognisable real person — politician, celebrity, your neighbour — performing actions and saying words of the maker’s choosing. Quality is uneven. A skilled forger working for an hour can produce something most viewers, scrolling on a phone, will not flag. The tells (a flicker on a tooth, a hand with the wrong number of fingers, lighting that does not match the room) are visible if you look for them, but the design of social-media feeds is that nobody looks.

Audio. Voice cloning is now the most mature of the three. Roughly thirty seconds of clean audio of a target voice — a podcast clip, a phone-call recording, a Facebook video — is enough to produce arbitrary new sentences in that voice with intonation a casual listener will accept. Phone-line audio compression actively helps the forgery, because it hides the small artefacts.

Image. Static images have been forgeable for years; what has changed is the speed and the scale. A small operation can now generate hundreds of plausibly-real photographs in an afternoon — fake protesters, fake crime scenes, fake screenshots of nonexistent news articles.

Three risk patterns for Nepal

The global discourse on deepfakes tends to focus on Hollywood-style scenarios. Nepal will see something more pedestrian and more damaging.

1. Political smears in narrow election windows. Imagine a fifteen-second clip surfacing on TikTok at 9 p.m. the night before voting: a candidate, in their own voice and likeness, appearing to confess to a payoff or to insult a particular caste or community. By the time fact-checkers wake up, the clip has been forwarded through ten thousand Viber groups. By the time the candidate issues a denial, voting is underway. The denial reaches a fraction of the audience the fake reached, and reaches them less viscerally. The Election Commission has no procedure for this. Most district-level journalists do not have the tools to authenticate the clip even if they wanted to.

2. Non-consensual intimate imagery, weaponised against women. This is already happening at scale globally, and Nepal is not protected by language or distance. A college student’s photograph from Facebook is enough raw material. The image is generated, shared in a WhatsApp group, then on Telegram, then back into the woman’s own social circle. The harm is not abstract — it is loss of jobs, loss of marriages, withdrawal from public life, and in the worst cases, suicide. The Cybercrime Bureau receives a complaint; the image is already irrecoverable. Existing law was written for a world where producing such images was hard.

3. Voice-clone scam calls. A parent in Birgunj receives a phone call. The voice on the other end is their son, who works in Qatar. He is in trouble — an accident, a hospital, the police, send money now to this account. The voice is correct. The fear is correct. The money is sent, and afterwards the real son calls from Qatar to ask why his mother is crying. The audio sample for the clone came from a thirty-second voice note the son had posted to a public Facebook group two years earlier.

These three patterns share a structure: the fake exploits an emotional reaction (outrage, shame, fear) that is engineered to short-circuit verification. The technology matters, but the social engineering matters more.

Defensive habits worth building

The instinct to demand a technical solution — “an app that detects deepfakes” — is understandable and largely wrong. Detection tools exist, they are improving, and they will always lag behind generation tools, because the generator can train against the detector but the detector has to wait to see new fakes. Anyone selling you confident “AI deepfake detection” is selling you a probability with a generous round-up.

The defences that actually work are mostly behavioural.

Cross-source verification. If a clip is real and consequential, it will be reported by more than one trusted source within hours. Wait for the second source. The cost of waiting is small. The cost of forwarding a fake is larger than you think.

Original-context lookup. Reverse-image search the screenshot. Search the quote. Check whether the alleged location and date appear in news from that day. A surprising fraction of viral fakes are debunked by a five-minute search.

Default skepticism for the explosive. The more emotionally activating the content, the more likely it was designed to be activating. A clip that makes you want to share it immediately is exactly the clip to sit on for a day.

Out-of-band verification for the personal. If a relative calls in distress asking for money, hang up and call them back on a number you already know. If they cannot be reached, call a third family member. A voice on a phone is no longer proof of who is speaking.

What institutions should be doing — and mostly are not

Newsrooms in Nepal need verification desks, even small ones, with training in basic forensic techniques. The Election Commission needs published procedures for handling alleged synthetic media during election periods, with a fast-track review and a public record of decisions. Platforms — TikTok, Facebook, Viber — should be pressured (by the Press Council, by civil society, by government) to label suspected synthetic media and slow its propagation during sensitive windows. None of these are happening at scale yet. The technology is moving faster than the response.

In the meantime, the burden falls on individuals. That is unfair, but it is the situation.

Check your understanding

Quick check

A short, explosive video clip of a public figure appears on your TikTok feed late at night, with no source link and no other outlet reporting it. What is the most defensible first reaction?

Quick check

Which statement most accurately describes the state of AI-based deepfake detection tools in 2026?

What comes next

If synthetic content is going to flood the information supply, the institutions whose job it is to separate signal from noise — newsrooms — are under unusual pressure. The next section looks at how Nepali journalism is using AI itself, where that helps, where it erodes trust, and what a workflow that preserves credibility actually looks like.