# The Livestream Pipeline
I have thousands of Facebook livestreams saved — Meta deleted the originals, I kept copies. The project, stated plainly: organize the streams and the notes I made them from, optimize, normalize the video, extract the audio, upload everything to Google, transcribe it all. Then regenerate every stream from its transcript, with my voice and my likeness. The folders are the source. The renders are disposable. The renderer is a swappable backend — crude blocks with the right voice today, indistinguishable tomorrow. Resolution is not the asset. The script corpus is.
The pipeline splits three ways, and I want all three held in the head at once. Long-range, it's the [[wiki/Digital Twin|memorial/continuity twin]]: the livestreams are the natural-personality corpus — social media captured my authored output, the streams captured the person — and without them the twin would be thin and horrible. Short-range, it's [[wiki/Digital Continuity|data rot]] management: formats decay, platforms delete, codecs rot; normalize now, keep the source canonical, re-render as formats change, and nothing is ever lost to a dead container again. And in the middle, the conversion processes themselves — video to audio to transcript to script to render — each stage a durable intermediate, each one re-runnable whenever a better backend ships.
I'm evaluating the backends now and committed to none of them, because the folder contract makes them interchangeable. Tavus: Griffin is gated — research preview, trusted testers only, verified this morning — but their CVI platform is shippable. HeyGen: public API, self-service voice cloning from a two-minute sample. MiniMax H3 Max Director via fal.ai: the sitcom engine — persistent character cards, directed scenes, native audio — though its voice is synthesized rather than cloned and it tops out at 768p. TwelveLabs: I signed in today to test the video-understanding leg, and speaker attribution on two-speaker streams is the falsifier I'm running before any bulk indexing — ten free hours a month is too thin to waste on uploads that don't answer the co-host question. And the transcript leg starts free and local: mlx-whisper on the Mac, today, no cloud, no per-minute billing.
The long range is where it gets strange in the good way. Working with an agent — Ford, or any other — the journal entries I dictate through the day become feedstock. Morning journal in, afternoon sitcom out: transcripts fed through whatever stage of my likeness the current pipeline supports, pushed out in a variety of formats, practically in real time. The livestreams stop being something I performed and become something I emit. The day, rendered.
The loop closes on the oldest instinct in the vault: pay attention, externalize it, revise it. The archive was never a mausoleum. It's a studio. Everything recorded is feedstock, and everything feedstock is re-renderable forever.
## Related topics
- [[wiki/Digital Twin|Digital Twin]] - the memorial/continuity twin this pipeline feeds; the livestreams are its personality corpus
- [[wiki/Digital Continuity|Digital Continuity]] - the anti-rot layer: canonical source folders, re-renderable across formats forever
- [[wiki/Memory Continuity|Memory Continuity]] - the archive as the substrate continuity is built on