logoalt Hacker News

playfultones • today at 11:23 AM • 0 replies • view on HN

Sounds like something a better model armed with ffmpeg would already be able to do. Run an analysis on loudness along the track, detect when speech begins/ends (with whisper) around loud segments, and compress or cut those parts out