A man sits in front of his laptop and asks an AI to generate a video of himself speaking Mandarin. He has never spoken Mandarin. Within seconds, the screen shows his face, his gestures, his mannerisms, forming words he does not know in a voice that sounds almost like his own. He laughs. It feels like a party trick.
Google’s Gemini model now converts anything into anything. Text to video. Image to audio. A still photograph to a moving, speaking figure. The demonstrations are, by design, playful. A sketch of a dog becomes an animated creature. A napkin doodle becomes a polished rendering. But tucked inside every playful demo is a quieter capability: you can make a person appear to say or do something they never said or did. And the distance between imagining that and executing it has collapsed to a single prompt.
The conversation about deepfakes has focused, reasonably, on the worst cases. Election interference. Revenge pornography. Fraud. But those scenarios assume malice, a person who sets out to deceive or harm. What’s less examined is the vast middle territory where fabrication happens without a clear villain. The person who generates a funny video of a friend “singing” at karaoke. The teenager who makes a clip of a classmate “confessing” a crush as a joke. The employee who tests what their boss would look like delivering a ridiculous speech. None of these people think of themselves as doing something wrong. That is precisely what makes it worth examining.
The Ethics of Almost
Most ethical frameworks are built around actions. You did something, or you didn’t. You caused harm, or you didn’t. But frictionless fabrication introduces a new category: the thing you almost did. The video you generated but didn’t share. The fake voice clip you made and then deleted. The synthetic image you created of someone you know, looked at for ten seconds, and closed the tab.
These moments feel private. They feel harmless. No one was hurt because no one saw it. But something happened in that moment anyway. You discovered that you could represent another person without their knowledge. You experienced what it felt like to control their image, their voice, their apparent behavior. And you learned how easy it was.
This matters because ease changes behavior. Research in behavioral economics has shown repeatedly that reducing friction increases action. Make a form shorter and more people fill it out. Put candy at eye level and more people take it. When the barrier between “I wonder what that would look like” and “now I’m looking at it” drops to zero, the wondering itself becomes a kind of doing. The thought experiment becomes an experiment.
The question is not whether generating a synthetic video of someone without their consent is wrong. In most cases, the answer is obviously yes. The question is what it means that millions of people will soon be able to do it in the time it takes to type a sentence, and that most of them will not think of it as a decision at all.
Consent in the Age of Instant Representation
Consent has always been complicated in photography and video. Street photographers capture strangers without asking. Security cameras record everyone. Social media posts include bystanders who never agreed to appear. But in each of these cases, the representation is at least tethered to something that actually happened. The person was there. The moment was real. The image, however unauthorized, documents reality.
Synthetic media severs that tether entirely. A generated video of you speaking words you never spoke is not a record of anything. It is a construction. And it carries your face, your voice, the markers that other people use to identify you as you. When someone sees it, they process it the same way they process a real video. The parts of the brain that recognize faces and interpret speech do not pause to check metadata.
This means consent is no longer just about being present when a camera is pointed at you. It extends to every use of your likeness, in any context, by anyone with access to a photograph and a text box. Your face becomes raw material. Your voice becomes a template. And the person using them may genuinely believe they are doing something lighthearted.
The framing of these tools as creative is not accidental. “Imagine the possibilities” is the default marketing language. Creativity implies self-expression, art, play. But creation and fabrication share a boundary that gets thinner every time the output quality improves. When a generated video is indistinguishable from a real one, calling it “creative” is a choice about language, not a description of what it is.
The Quiet Recalibration
What happens to a society where synthetic media is ambient? Not a crisis. Not a sudden collapse of trust. Something slower. A gradual recalibration of what counts as evidence, what counts as real, what triggers skepticism and what doesn’t.
Early internet users learned not to trust email from strangers. A generation later, most people learned to question headlines shared on social media. Each wave of manipulation produced, eventually, a corresponding wave of literacy. But each wave also produced a permanent residue of doubt. You check sources now. You reverse-image-search. You read past the headline. These are healthy habits, but they are also symptoms of an environment where default trust no longer works.
Frictionless fabrication accelerates this cycle. When anyone can generate a plausible video of anyone else, the rational response is to treat all video with suspicion. This protects you from being fooled. It also means that real footage of real events carries less weight. The whistleblower’s recording. The bystander’s phone video. The documentation of something that actually happened to a real person in a real place. All of it becomes deniable, because the existence of easy fabrication gives everyone a ready-made objection: “that could be fake.”
This is not a hypothetical. It is already the standard response to inconvenient video evidence in political contexts around the world. The technical term is the “liar’s dividend.” The mere existence of deepfake technology benefits anyone who wants to deny the authenticity of real footage. And as the tools improve and spread, the dividend grows.
The people building these models know this. The demos are careful to show fun, harmless applications. A talking cartoon. A translated speech. A creative remix. But the capability does not come with a use-case filter. The same model that turns your sketch into an animation can turn your ex’s photograph into a fabricated confession. The same interface that lets you “try on” a new language lets someone else try on your face.
There is no policy solution that will prevent this. The technology is too accessible, too distributed, too useful in its legitimate applications to be contained by regulation alone. What can shift is the culture around it. The recognition that generating someone’s likeness without their knowledge is not a neutral act, even if no one else ever sees it. The understanding that “I didn’t share it” is not the same as “nothing happened.” The willingness to treat the moment before you press send as the moment that matters most.
Because that is where the ethics actually live now. Not in the sharing. Not in the consequences. In the fraction of a second where you hold a fabricated version of a real person on your screen and decide what kind of relationship with reality you want to have.
Digital Alma explores the intersection of technology, consciousness, and what it means to be human in a digital world.


Leave a Reply