I just have Hermes download the video using yt-dlp and transcribe it with a local whisper model using a local whisper server if the transcript is not available through the API, whisper base models works fairly well on CPU only servers nowadays.