AI audiobook narration that people finish
Modern local TTS is good enough that the voice is no longer the problem. Pacing is. A model reads punctuation literally, so an unedited manuscript produces narration that is technically clear and exhausting to listen to. The work is preparing text for the ear rather than the eye.
Rewrite for the ear before you generate anything
Long subordinate clauses work on a page and collapse in audio. This pass matters more than the voice you pick.
Build a pronunciation list early
Names and invented words need overrides, and finding them after a full render means regenerating everything.
Generate per chapter and listen to each one
Whole-book renders hide problems in the middle. Chapter-level output is small enough to actually check.
Master for loudness, not peak
Distributors reject on loudness targets, and this is the most common late failure in the whole process.
Produced with Kokoro running locally, the same stack Kevin uses rather than a paid narration service.