Local-first audio workspace

Create audiobooks, subtitles, and voiceovers without hiding the workflow.

Pandrator is a local-first browser workspace for preparing source material, generating speech, reviewing intermediate artifacts, and exporting finished audio or video.

Windows · Linux · browser interface

Local models keep source files and generated media on your machine. Connected providers are optional.

Three starting points

Choose an outcome. Keep each stage reviewable.

Pandrator assembles a workflow around the result you want. Intermediate text, audio, subtitles, and revisions remain available for inspection and reruns.

Audiobooks

Import documents, clean and segment the text, generate reviewable speech, and export chaptered M4B or common audio formats.

Subtitles

Transcribe audio or video, correct and translate the result, edit timing and text, then export SRT or subtitle tracks.

Voiceovers

Create speech from subtitles, synchronize it with media, and choose original, mixed, or dubbing-only audio.

A visible workflow

Preparation and generation are separate steps.

  1. 01

    Prepare

    Import text, documents, subtitles, audio, video, or a supported public media URL.

  2. 02

    Review

    Inspect extraction, transcription, correction, translation, timing, and speech plans before moving on.

  3. 03

    Generate

    Use a local speech service, a configured commercial API, or a custom compatible endpoint.

  4. 04

    Compare and export

    Regenerate segments, compare takes and RVC variants, then assemble the required media.

Local and connected

Use the processing path that fits the material.

Pandrator does not require a cloud account. It can also use commercial services where their quality, language coverage, or convenience is useful.

Fully self-hosted

  • Local speech and transcription services managed by the installer
  • Local OpenAI-compatible LLM servers such as LM Studio
  • Files, jobs, voices, and generated artifacts under your selected data root
  • No Docker or WSL requirement for the packaged Windows workflow

Connected providers

  • OpenAI and Google Gemini speech services
  • Supported cloud LLM providers for correction and translation
  • Custom OpenAI- or Gemini-compatible speech endpoints
  • Optional bounded Jina research for uncertain terminology

External providers may receive source text, subtitles, audio, or voice samples and may charge for usage. Review their terms before enabling them.

Voices and languages

Ready-made voices, cloning, and conversion are distinct tools.

Use a pre-built voice when you do not need a reference recording. Use a cloning model for a supplied sample. RVC is a later speech-to-speech conversion step, not a TTS model.

Common inputs and outputs

Documents, SRT, common audio and video inputs; M4B and common audio formats; MP4-oriented dubbed output with selectable tracks.

Open the capability matrix

Start with the smallest useful installation.

The launcher can add, update, repair, or remove optional services later. You do not need every model to complete a workflow.

Choose a first setup