VoiceStudio
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation i
Compare 217 quality-filtered Audio Tools tools using public popularity, activity, licensing and source data.
Browse quality-filtered Audio Tools tools ranked by public signals. Filter by source, pricing, or recency to find a fit faster.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation i
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
Open-source framework for conversational voice AI agents
Build voice agents with open-source models
The free and privacy-friendly screen recorder with no limits 🎥
Open Source framework for voice agents, multimodal apps, and realtime AI. Maintained by Daily and the community.
ModelScope: bring the notion of Model-as-a-Service to life.
ODS V3: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
视频转字幕、字幕翻译、AI 配音与声音克隆、字幕烧录——免费开源的一站式桌面工具。基于 Whisper / FunASR 等本地模型离线语音转文字,批量处理 + 全平台 GPU 加速,跨 Wind
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
A simple, high-quality voice conversion tool focused on ease of use and performance.
UI components and hooks for building video/audio players on the web. Robust, customizable, and accessible. Modern alternative to JW Player and Video.js.
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS,
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly op
Data manipulation and transformation for audio signal processing, powered by PyTorch
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
OpenVidu Platform: self-hosted real-time video and audio for your apps, built on LiveKit and mediasoup
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
SALMONN family: A suite of advanced multi-modal LLMs
A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning
Tools for handling multimodal data in machine learning projects.
The most advanced, fully offline client-side AI suite on Android today.
🎙️ An intelligent voice productivity assistant that turns speech into clean text, useful actions, and structured knowledge. It helps users capture ideas, commun