de-thorsten-1p57m/README.md
566 B · 12 lines · markdown Raw
1 # Root A raw fp16 package
2
3 This package is the current duration + acoustic + decoder student stack as raw fp16 weights.
4 It is intended for a standalone runtime implementation, not PyTorch loading.
5
6 - Parameters: 1,565,324
7 - Weight blob: 3,130,648 bytes
8 - Asset bytes with phoneme config: 3,135,467
9 - Duration length scale: 1
10 - Runtime boundary: accepts Piper phoneme ID sequences, predicts durations, predicts Piper generator latents, then decodes waveform.
11 - Not complete arbitrary-text TTS yet: the Piper/eSpeak text frontend still has to be replaced or embedded.
12