Product
Local AI Audio Transcription
Batch transcription of long form video and audio files using open models, local processing and low cost infrastructure.
- Problem
- Needed to transcribe videos exceeding two hours on a recurring basis. Using paid services would increase costs and require uploading large files, making the processing workflow more challenging.
- My contribution
- Developed a solution that splits content into short audio segments and uses faster-whisper to generate transcripts. Processing can run on CPU or be accelerated with NVIDIA GPUs via CUDA, depending on available resources.
- Outcomes and learnings
- Enabled low cost batch transcription of long form content by leveraging available hardware and keeping processing local. The project demonstrated how open models can address a recurring need without relying on paid transcription services.