Voice to Version 1.0
Exploring transcription workflows and getting my face out of the screen.
As my travel schedule ramps up for 2026, I’ve found myself spending a significant amount of time heading to the airport or in transit between destinations. When I’m not catching up on phone calls, it occurred to me that I could reclaim that time to generate content using basic transcription available on the iPhone.
This isn’t exactly a new frontier for me. Many years ago, back when I was still rocking a BlackBerry, I was experimenting with early application support for transcription. Back then, it was never quite clear if the audio was being processed by a computer or if a WAV file was being quietly shipped off to a third party for “Mechanical Turk” style human transcription.
Gadgets are interesting but if the best camera is the one you have, then for audio, that’s my iPhone too. Plus, I’m also looking for a simple alternative now that I’ve discontinued using Limitless AI pendant.
The Medical Model
Transcription is a staple in other industries, most notably in medicine, where doctors narrate their notes for later processing. Now, I’m certainly not a doctor nor do I play one on TV, but the logic holds.
The Apple ecosystem and onboard transcription have reached a point of high reliability. Whether the processing stays on the device or hits the cloud for speech-to-text, the quality is finally there. The real question now is: what kind of post-processing can we apply to make it actually useful?
Enriching the Stream
This is where the “wonders of containers, open source, and my Google Gemini Gem(s)” come into play. By feeding a Gem the knowledge of my prior blog posts (including the frontmatter and permalinks) I can do more than just transcribe. I can enrich.
For example, if I mention a topic like the space value chain, the workflow could automatically provide deep links to my existing content. A short Python script and some n8n glue could also grab the latest Techmeme headlines or meta-commentary from Techmeme Tech Brew Ride Home to add context.
Version 1.0
My goal is a “minimum viable” toolchain that gets a draft into version control on GitHub. From there, my webhooks are already in place to build the site and ship the post into Fudge Factor Newsletter for subscriber dissemination.
The most exciting part? It keeps my face out of a screen. We could all use a “welcome departure” from the screen-staring that constitutes modern knowledge work.
The Provocative Path
Some might consider this to be the opposite of a podcast because it is recording audio to get text. However, if I wanted to be provocative, I could take the resulting text, use a trained facsimile of my voice, and create an artificial recording of me reading the post I just spoke into existence.
I don’t think I’m quite there yet. Capturing my specific inflections and mannerisms would require significant training against my video and audio archives. It’s on the list of things to try, but let’s be clear: I haven’t uploaded my consciousness to the cloud. That trope, and that dog, won’t hunt. At least not yet.
Enjoyed this post?
Consider supporting my sponsor!
🔓 Unlock Your Best Self with Lida Coaching! 50% off regular coaching rates until July 31, 2026!
A quick message from Lida:
With 18 years building, marketing, and scaling products at startups, I work with founders and leaders stuck between where they are and where they need to be. I help you cut through the noise, make sharper decisions, and get your team executing with clarity instead of chaos.
Most founders mistake complexity for strategy. They keep adding when they should be subtracting. When I'm brought in, the team is usually doing a lot right, and the real issue is focus and direction. I diagnose what's actually going wrong across product, positioning, and execution, then reset direction and get everyone aligned around one clear path forward that will lead to growth.
A bit of what I've done: I was the first PM on a fintech product that hit 10,000+ users within months, built a national incubator that supported over $500M in startup funding, and led teams at 23 Design before it was acquired by frog (the design team behind Apple’s first products). I've spoken at Google, Microsoft, and LinkedIn, etc.
🔓 Unlock Your Best Self with Lida Coaching! 50% off regular coaching rates until July 31, 2026!