I’ve been busy pruning an overgrown laurel hedge. Apparently, the previous owner cultivated the entire leafy barrier from a single twig he had surreptitiously snipped from a hedge in a garden centre. His careful propagation from that twig—growing, dividing, and replanting—constitutes cloning. A quick AI-assisted search reveals that the verb “to clone” derives from the ancient Greek klōn, meaning “twig.”
Now, “cloning” refers to any asexual reproduction—the creation of a duplicate without mingling genetic material from two progenitors. It’s a process that has become a source of scientific, social, cultural and spatial fascination. Voice cloning enters into that category.
Private listening
Digital text-to-speech (TTS) applications such as Speechify and ElevenLabs can now recite text in convincing, human-sounding synthetic voices. I’ve long used such applications to conquer the vast quantities of text required in academic and professional life. Listening to readings of texts allows me to absorb texts while driving, walking, cooking, or exercising—exploiting our human capacity to multitask across different sensory modalities.
Listening via earbuds also has a spatial dimension. It places the listener inside an audio cocoon. Others often assume that you want to be left alone, as your attention turns inward, attuned to your own mobile “sub-architecture.”
Vocal repetition
To this isolating spatiality add the influence of vocal repetition. Elsewhere I’ve explored how repeated utterances—whether for emphasis, in chants, songs, or public announcements—shape spatial awareness. See post: Hustle, twitter, bells and banter, and my book The Tuning of Place.
The echo of one’s own voice, as in a recording, reveals the dimensions and materiality of space. As an obvious example, how quickly a sound returns provides clues about a room’s size and surfaces.
The reproduction of the human voice in audio media—tape, film, video, streaming—constitutes a form of repetition. We listen and re-listen to songs, speeches, short clips, podcasts, and endlessly reworked soundbites. These repetitions claim attention and space.
Voice cloning
Voice cloning is also a form of repetition. In cloning mode, the ElevenLabs TTS system asks you to read a sample text. It then detects your vocal features and can synthesise speech using your own voice—not just what you’ve said, but new content. You hear yourself saying something you’ve never spoken.
This ability to replicate and reassign the human voice—your own or someone else’s—raises possibilities, risks, and challenges. For some, it evokes a sense of dislocation and even theft. Like the poaching gardener who started my laurel hedge from a purloined twig, voice cloning invites questions about ownership, origin, and the ethics of reproduction.
Reference
- Coyne, Richard. The Tuning of Place: Sociable Spaces and Pervasive Digital Media. Cambridge, MA: MIT Press, 2010.
Note
- The featured image is from ChatGPT: “Please provide a picture of a pair of old fashioned rusty hedge trimmers. … Revise image to show closed tarnished metal clippers lying in a laurel hedge.” (I didn’t ask it to correct the dysfunctional pivot arrangement.)
Discover more from Reflections on Technology, Media & Culture
Subscribe to get the latest posts sent to your email.
1 Comment