Aletheia Russian TTS corpus tools
audit_corpus.py verifies that an audiobook and EPUB contain the same text before any training data is produced.
It extracts the EPUB spine in reading order and compares local Whisper transcripts from the beginning, middle,
and end of the audiobook with the normalized book text.
Generated audio, transcripts, downloaded ASR models, and future training checkpoints belong under
.codex-temp/tts-*; they are working artifacts and must not be committed.
Example on Windows PowerShell:
$env:PYTHONPATH = 'C:\Repos\Aletheia\.codex-temp\tts-python'
python tools\tts\audit_corpus.py `
--source-dir 'C:\path\to\audiobook' `
--output-dir '.codex-temp\tts-audit' `
--model base
The audit is not a copyright or voice-consent check. Training and distributing a voice model requires separate confirmation that the recordings and narrator's voice may be used for that purpose.
After the audit passes, transcribe_corpus.py creates one resumable JSON sidecar per MP3 with segment and
word timestamps. Existing sidecars are skipped unless --force is supplied. These transcripts are alignment
anchors only; the final Piper labels must come from the exact EPUB text.
build_piper_dataset.py globally aligns all ASR tokens to the EPUB in monotonic reading order, rejects weak
matches, merges adjacent short segments, and writes 22.05 kHz mono PCM WAV files plus Piper's
filename.wav|Exact book text metadata. Use --dry-run first and inspect the reported exact-token match ratio
before producing WAV files.
train_piper.sh is the CUDA/Linux entry point for scratch training and ONNX export.
train_piper_windows.ps1 provides the equivalent native Windows path and a -SmokeTest mode that executes one
training and validation batch on CUDA before a long run. See GPU_TRAINING.md for host prerequisites and the
acceptance gate.
Vosk BookTTS compression experiment
vosk_booktts_experiment.py reproduces the Android Vosk frontend, prepares real acoustic-model feeds from
booktts_calibration_ru.txt, performs calibrated QDQ int8 quantization of convolution layers, and renders
deterministic baseline/candidate WAV pairs. Keep generated .npz, candidate ONNX, and WAV files outside tracked
source directories, for example under .codex-temp. Do not replace the application asset until the candidate
loads with ONNX Runtime, runs on the target phone, and passes listening comparison.
Use prepare --bert-model plus compare-feeds to isolate a BERT replacement while keeping the acoustic model
identical in both WAV branches. compare-feeds --resume preserves existing WAVs and atomically updates its
report after every pair; run_booktts_ab_windows.ps1 is the persistent Windows entry point.
booktts_student.py builds token-level teacher targets from the current int8 BERT and trains a four-layer,
256-dimensional Transformer student with a 768-dimensional drop-in output. Its small calibration corpus is only
for validating the training/export pipeline; a candidate for the application requires a much larger licensed
literary corpus, held-out evaluation, ONNX quantization, and listening tests. The optional
--max-training-batches and --max-validation-batches limits are for CUDA memory smoke tests only.
Each completed epoch atomically replaces checkpoint.latest.pt; pass --resume with the same output directory
to continue an interrupted run.
run_booktts_student_windows.ps1 is the native Windows CUDA entry point for a full or resumed student run.
The exported student includes the dataset's original-to-compact token map so its input_ids remain compatible
with the existing Vosk vocabulary. patch_booktts_student_token_map.py adds the same lookup to an older export
without retraining.
compare_booktts_bert.py measures teacher/student and FP32/INT8 embedding error, inference time, and unseen-token
coverage on a held-out text file.
build_booktts_prosody_corpus.py creates a deterministic balance of questions, exclamations, quotations,
ellipsis, and neutral prose. Build its teacher dataset with --reuse-token-map, then fine-tune from the prior
student.pt with --initial-model; the held-out listening phrases must be passed through --exclude.
run_booktts_prosody_dataset_windows.ps1 builds this reused-vocabulary teacher dataset persistently on Windows.
After auditing that dataset, run_booktts_prosody_finetune_windows.ps1 starts the low-learning-rate CUDA run.
extract_wikisource_corpus.py streams an official Russian Wikisource XML/BZip2 dump and writes a deduplicated,
filtered sentence corpus plus provenance metadata. The source URL and dump size must be retained. Wikisource is
not a blanket rights clearance for every included work, so review the resulting corpus and applicable source
licenses before distributing a trained model.
run_booktts_corpus_windows.ps1 completes the resumable BITS download on the CUDA host, verifies the official
dump size and SHA-1, extracts the configured number of sentences, and builds the streamed float16 teacher
dataset. It intentionally stops before training so vocabulary coverage and a held-out evaluation plan can be
reviewed first.
audit_booktts_text_corpus.py records the corpus SHA-256, sentence-length and word-count distributions,
residual wiki-markup counters, and deterministic review samples. Run it before training; its report is a
technical quality check and does not replace a rights review of the source works.
audit_booktts_dataset.py opens every generated teacher shard and verifies sample counts, tensor shapes,
token ranges, binary masks, and finite teacher values. It also prints a reproducible SHA-256 over shard bytes.