Documentation
Audiobook workflow
Document cleanup, narration preparation, generation, assembly, and export.
- New users
- Operators
An audiobook session starts from an uploaded, downloaded, or deliberately reused document artifact.
Normal stages
- Clean source — deterministic extraction with optional agent-assisted cleanup, producing reviewable clean text.
- Segment narration — editable generation segments controlling text boundaries and pauses.
- Optimize narration — optional, separate before-and-after text revision.
- Generate audio — reviewable narration takes. Missing preparation may be included by an exact workflow plan.
- Assemble — select takes and construct the intended audio sequence.
- Export — package assembled audio with format, metadata, and cover choices.
Whole-document speech optimization and generation-time batch optimization are alternative places to perform the same kind of transformation. Review the plan and avoid enabling both unintentionally.
Generation can send narration text to a configured TTS provider. Cleanup or optimization can send text to an LLM provider. A useful plan therefore states which provider receives which data before execution. Generated takes and final exports are different artifacts; generation does not imply final assembly.