AI Audiobook Narration: What Authors and Publishers Need to Know Before Choosing a Voice

AI Audiobook Narration

What Authors and Publishers Need to Know Before Choosing a Voice

AI can read words. But can it truly tell a story?

By Elliott Frisby
Founder and Director, Monkeynut Audiobooks & Sound

When there is a genuine choice, choose the human.

Let me begin by saying something clearly: I am not anti-AI. I work in audio. At Monkeynut, we use technology every day. It helps us edit, master, clean recordings, improve workflows and solve problems. Used well, technology can support the creative process brilliantly. But there is a very big difference between using technology to support a performer and using technology to replace one.

When an author or publisher asks me whether they should choose AI audiobook narration or a real human narrator, my honest answer is simple: if there is a genuine choice, choose the human.

That is not because I am frightened of what technology can do. It is because I know what a skilled narrator, producer and production team bring to a book. An audiobook is not merely written text converted into an audio file. It is a performance, and the quality of that performance shapes the listener’s relationship with the author, the story and the publisher behind it.

Audiobooks are not a secondary format

The audiobook market is growing, and that makes decisions about narration increasingly important. The Publishers Association reported that UK digital consumer audiobook revenue reached £255 million in 2025, an increase of 10% on the previous year. Digital audiobooks accounted for 10% of the consumer publishing market. This is not a novelty sitting at the edge of publishing. Audio is an established and valuable way for readers to experience books.

And this is not just a British story. In the United States, audiobook publishers’ sales revenue reached US$2.43 billion in 2025, up 9% on the previous year, according to the Audio Publishers Association. For me, that reinforces the point: audio deserves the same creative care as any other edition of a book.

For some people, the audiobook will be their only encounter with an author’s work. They may listen while driving, walking, travelling or resting. They may spend eight, ten or fifteen hours with the narrator’s voice in their ears. That voice becomes the bridge between the author and the listener.

If the bridge feels impersonal, inconsistent or emotionally false, the listener may not blame the technology. They may simply decide that they do not like the book. That is why the choice of narrator is not just a production decision. It is a creative and reputational decision.

What does AI audiobook narration actually mean?

Not every form of AI narration is the same, and authors and publishers need to understand the language being used. International guidance developed by the Audio Publishers Association and the UK Publishers Association’s Audio Publishers Group distinguishes between two main categories.

An AI Voice is a synthetic voice created from samples involving a larger group of unidentified speakers. It is not presented as a digital version of one particular performer.

An Authorised Voice Replica is created using licensed or authorised samples of a specific human voice and is intended to reproduce that person’s voice. The performer, rights holder or estate has given permission for that use.

The guidance uses voice cloning to describe unauthorised replication, where a person’s voice has been copied without permission.

Those distinctions matter. A generic synthetic voice, a properly licensed replica and an unauthorised clone raise very different creative, ethical and contractual questions. Calling all of them simply “AI narration” can hide important information from authors, publishers, performers and listeners.

Why AI narration can look attractive

I understand the appeal. AI narration may appear quicker and cheaper. It can make an audio edition seem possible for a title with a limited budget. It may offer a way to produce books that would not otherwise receive an audio release. Amazon KDP already operates an invite-only virtual-voice beta in the United States, so this is no longer a theoretical conversation.

Those are real considerations, and it would be unhelpful to pretend otherwise. But speed and price are only part of the cost of an audiobook. The more important question is whether the finished production serves the book.

An author may have spent years writing, revising and protecting every sentence. A publisher may have invested in editing, design, marketing and distribution. If the audio edition is then produced by choosing the quickest available voice, the final performance may not reflect the care already invested in the work.

Cheaper is not always wiser. Faster is not always better. The saving made at the beginning can reappear later in poor listener reviews, weak engagement, reduced trust and an audio edition that does not represent the author’s best work.

A narrator is not simply a voice

A professional audiobook narrator does far more than pronounce the words correctly. They interpret the writing. They make thousands of small decisions about pace, tone, emphasis, character, breath and silence. They understand when a sentence needs energy and when it needs restraint. They can feel when a joke should be allowed to land, when a difficult passage should not be overplayed and when a pause is doing more work than another word could do.

Good narration is full of judgement. The best performers do not force themselves between the author and the listener. They serve the writing. They make the text clearer, more immediate and more human without drawing attention away from it.

AI can produce fluent and increasingly natural speech. I do not underestimate that. But natural-sounding speech and human interpretation are not the same thing.

A system can be instructed to sound sad, warm, urgent or reassuring. A human performer can understand why the moment matters and decide how much emotion the listener needs. Sometimes the right decision is not to sound more emotional. It is to hold back. That distinction can be heard.

Some stories need particular care

At Monkeynut, we work across a wide range of genres, including many titles that explore faith, grief, trauma, illness, family, survival, calling and hope. Authors trust us with work that is often deeply personal.

Imagine an author describing the death of someone they love, their experience of abuse, a crisis of faith or the day they received a life-changing diagnosis. It is possible for a synthetic voice to say every word correctly. But correct pronunciation is not the whole job.

Does the performance understand the dignity of what is being shared? Does it know when not to make suffering theatrical? Does it recognise when the author is using humour to release tension? Does it understand that one sentence may need silence around it?

A thoughtful human narrator does not just read material like this. They protect its meaning. For memoir, narrative non-fiction, faith-based books, children’s publishing, fiction and any work built around a distinctive authorial voice, I believe that human understanding is particularly valuable.

Professional audiobook production is collaborative

There is another part of the conversation that is easily missed. A professional audiobook is not created by one voice in isolation. Behind the performance there may be a casting decision, a producer guiding the session, an editor protecting the rhythm, a proof listener checking the text and a mastering engineer preparing the finished files for distribution.

A producer can hear when a narrator has misunderstood the intention of a sentence. They can spot a shift in pace, energy or microphone position before it affects an entire chapter. They can help a performer find the right tone without taking away the narrator’s own instincts.

That collaboration is not unnecessary expense around the recording. It is part of what makes the recording professional. At Monkeynut, we assign an industry producer to audiobook sessions because the goal is not simply to capture a voice. It is to guide a performance that someone will want to keep listening to for hours.

Consent and voice ownership cannot be an afterthought

The growth of synthetic voices also raises questions that go beyond performance quality. A voice is connected to a person’s identity, reputation and livelihood. If a performer’s recordings are used to create a replica, the agreement should be specific and informed.

The performer should know where the replica may be used, for how long, on which titles and platforms, and how they will be paid. The agreement should also explain what happens to the voice model and source recordings when the licence ends.

The UK Government’s March 2026 report on copyright and artificial intelligence recognised that existing law may not give people sufficient control over digital replicas of their voice or likeness. It noted that the UK does not currently have a general personal image right and that the available protections form a patchwork which may not cover every harmful or unauthorised use.

The same report confirmed that the Government’s consultation received 11,520 responses. Most respondents rejected the originally preferred proposal for a broad data-mining exception with an opt-out. Many people in the creative industries argued that requiring creators to opt out would place an unrealistic burden on them.

This area continues to develop, so authors, publishers and performers should obtain suitable legal advice for specific contracts. But the ethical starting point does not need to be complicated. If a voice is going to be copied, ask. If a performance is going to be replicated, agree the boundaries. If commercial value is created, pay the person fairly.

Consent should be active, informed and recorded. It should never be quietly assumed.

Listeners deserve clear labelling

Transparency should be the normal standard for AI-narrated audiobooks. If an audiobook is performed by a human narrator, say so. If it uses an AI Voice, say so. If it uses an Authorised Voice Replica, make that clear as well.

The international industry guidance recommends that publishers identify AI narration in a title’s metadata. It also says publishers should consider applying the terminology where more than 10% of a voice has been created using AI tools.

That is helpful because listeners should not have to guess what they are buying. Clear labelling respects the audience, protects confidence in the audiobook market and prevents a synthetic performance from being mistaken for the work of a human narrator.

Authors should also know how their audiobook will be described before agreeing to production. The wording used by a distributor or retailer can affect how the title is found, understood and received.

Can AI have a useful role in audiobook production?

Yes. Technology can support human creativity in valuable ways. It can help production teams identify noise, organise workflows, search long recordings, prepare text, improve accessibility and carry out repetitive technical tasks. An appropriately licensed synthetic voice may also have uses in assistive communication, temporary production material or projects where a conventional release would genuinely be impossible.

My objection is not to every use of AI. It is to treating human performance as an inconvenience that should automatically be removed because a cheaper option exists. An audiobook is a creative product. The fact that software can produce a voice does not mean that voice is the best choice for the book.

Technology should help us make better work. It should not lower our expectations of what the work can be.

Questions authors and publishers should ask before choosing AI narration

Before making a decision, I would ask:

  1. What does this particular book need from its narrator? Is it primarily functional information, or does it depend on emotion, character, humour, authority or personal testimony?

  2. What will the listener be told? Will the use of an AI Voice or Authorised Voice Replica be clear in the metadata and at the point of purchase?

  3. Who owns and controls the voice technology? Are the underlying recordings licensed, and has every relevant performer given informed consent?

  4. What editorial control will we have? Who will correct pronunciation, pacing, emphasis and passages that do not sound right?

  5. Has the whole book been listened to by a person? Generating speech is not the same as quality-controlling a finished audiobook.

  6. How will this choice affect the author’s reputation? Will the audio edition feel consistent with the quality of the printed and digital editions?

  7. Are we considering long-term value or only the initial price? A cheaper production is not a saving if listeners abandon it or the book has to be produced again.

  8. Have we properly explored a human production? Before assuming professional narration is impossible, speak to an experienced audiobook studio about casting, scheduling, production options and realistic costs.

Those questions move the conversation away from “Can AI read this?” and towards the more useful question: “What is the best way to bring this book to listeners?”

My view: choose the human

AI narration will continue to improve. It will become easier to access, and it will become part of more publishing conversations. We should discuss that honestly and without panic. But when an author or publisher has a genuine choice, I believe a real human narrator should always be chosen.

A human narrator brings experience, instinct, imagination and responsibility to the microphone. They can be directed. They can question a line. They can understand context. They can surprise us. Most importantly, they can care about what the words mean. That care reaches the listener.

I am not anti-AI. I am pro-human creativity, pro-consent, pro-quality and pro-fairness. Technology can be part of the future of audio, but the future should not be built by treating human voices as disposable.

Authors spend years finding the right words. When those words become an audiobook, they deserve a real performance.

And when there is a choice, choose the human.

Elliott Frisby
Founder and Director, Monkeynut Audiobooks & Sound

About Elliott Frisby

Elliott Frisby is the Founder and Director of Monkeynut Audiobooks & Sound, a specialist UK recording and production studio working across audiobooks, voice-over, ADR, podcasts and music. With around 30 years' experience creating and producing audio, Elliott works with authors, publishers, actors and production companies to create professional audio that keeps people listening.


Get in touch

Considering an audiobook / voice-over and unsure which production route is right for your book? Speak to Monkeynut Audiobooks & Sound about professional narration, casting, recording, editing, mastering and distribution.

Next
Next

Children from Our Diocese Record Audio for Tom Wright’s First Children’s Book