Drop in the book, give every speaking character a voice, fix the names a narrator would say wrong, choose how much sound goes under it, and make it chapter by chapter.
A plain AI reading of a book is one voice from start to finish. A full-cast audiobook gives every character their own voice, puts the reader inside each scene with ambience and sound effects, and brings in music at the moments that matter. It used to take a studio. Most of it can now be done from the EPUB.
This is how it works in Inkreel, step by step.
Start a new project from the home page with Add a book and pick the file. An EPUB works best, because it carries the book's own table of contents: the chapters come out with their real titles, in order, even when one chapter is split across several files or two chapters share one. PDF, Word (DOCX), plain text, Markdown and web-page files work too.
What isn't meant to be read aloud is left out: the cover, the copyright page, the contents page, picture captions, footnote numbers and printed page numbers. The book's cover becomes the project's picture and the artwork inside the finished M4B.
A copy-protected (DRM) EPUB can't be read. Use a DRM-free copy, such as the author's own file.
Every chapter is ticked. Untick what you don't want read: a long preface, an afterword, a list of illustrations.
Then set two pauses. One second after each passage is the usual audiobook pace; up to three gives the story more room. A longer pause after each chapter lets the listener know a new one is starting. Both change a finished audiobook at once, for free.
If you (or your favourite AI) have already written the chapters as scripts, with who reads every line, import them instead of having the app cast them. Download the chapter script template: one file per chapter, every word of the chapter once and in order, each line as `Scene 3, Drake: "Navigation?"`, with an optional Scenes table and Sound plan for each scene's ambience, effects and music. A delivery note like `(over the comm)` gives a line the radio sound.
On the Chapters step, Import chapter scripts takes all the files at once and matches each to its chapter by the number in its name. Every line is then read by exactly who the script says, the character list shows the exact number of lines each person speaks, and casting those chapters costs nothing.
The room, comm and machine sounds cost nothing: they're added while the audiobook is mixed. Ambience, effects and music use voice-provider credits, and the Create step prices them.
This is where most AI audiobook tools stop short. They find a handful of speakers and the rest of the cast is read by the narrator, or by whoever the tool guessed.
There are two ways to get everyone:
Either way, each character gets a Cast sheet that keeps their voice for every book and film in the same world. Pick a voice from your ElevenLabs account (your own designed and cloned voices are listed first), play its sample, or design a new one from the description.
Two settings belong to the character, not the voice:
Any character still without a voice when you create the audiobook is read by the narrator, so nothing is silently left out. You can switch that off.
Invented names and places trip up every text-to-speech voice. The Pronunciations step lists the book's own names and unusual words, most frequent first, with the sentence each appears in. Write how each should sound, spelled the way it's said ("Aruun = ah-ROON"). A rule can hold for everyone or for one speaker only, so a character can mispronounce a word on purpose.
There's no limit and no cost: nothing is generated on this step.
The Create step shows, for every chapter, what's left to do and what it will cost: casting the chapter (a text model splits it into passages and says who reads each line, word for word), the voices, and the sound. It also shows how many credits your ElevenLabs plan has left this month and how many chapters that covers.
Then choose how fast:
Whichever you pick, chapters are made in order, and if the credits run out partway through it stops cleanly and keeps everything made so far.
Voices are the big cost, and they're charged by the character. A 100,000-word novel is about 570,000 characters, so a full reading with a top ElevenLabs model takes about 570,000 credits. The faster Flash and Turbo models take half that. A monthly plan with about 120,000 credits covers roughly two chapters of a typical novel a month; a bigger plan covers the whole book at once. Casting the chapters with a text model adds a few dollars for the whole book.
Going chapter by chapter has an advantage beyond spreading the cost: you hear the first chapter before committing to the rest, and you can change a voice or a pronunciation while it's still cheap.
Every line, effect and music cue lands on a timeline. Listen to a three-minute sample, open the editor to move or trim anything, redo a single line, then export:
Both carry the book's cover and are marked as AI-made in the file's metadata. Add a still to each passage and the same project becomes a video book.
See the pricing, or read How to keep AI characters consistent between shots.