Feature Request: Add Support for Five New Audio Models in Backend & WebUI

#15
by Exo87 - opened

Hello audiocpp team,

I'd like to propose adding backend and WebUI support for the following five audio models. Each represents a distinct capability that would significantly broaden the scope of what audiocpp can offer, from zero-shot TTS to audio grounding, singing voice synthesis, speech enhancement, and mixed audio scene generation.

  1. Raon-OpenTTS-1B (https://huggingface.co/KRAFTON/Raon-OpenTTS-1B)
    Description: Raon-OpenTTS is an open-data, open-weight zero-shot TTS system based on an F5-TTS Diffusion Transformer with flow matching. The 1B variant achieves a word error rate (WER) of 1.78 and speaker similarity (SIM) of 0.749, performing on par with state-of-the-art closed-data models like Qwen3-TTS and CosyVoice 3.

Why it should be added: This model is fully open (both weights and training data publicly available) and supports robust zero-shot voice cloning with expressive speech. Its 1B parameter size makes it a strong candidate for a production-grade TTS backend, and its open licensing aligns perfectly with audiocpp's philosophy. Adding it would give audiocpp a competitive, fully open TTS option.

  1. SpotSound (https://huggingface.co/Loie/SpotSound)
    Description: SpotSound is a LoRA adapter built on top of NVIDIA Audio Flamingo 3 that enhances large audio-language models with fine-grained temporal grounding. It can accurately pinpoint the exact start and end timestamps of specific acoustic events within long, untrimmed audio recordings based on natural language queries. It incorporates a novel training objective specifically designed to suppress hallucinated timestamps for events absent from the input, and includes SpotSound-Bench, a challenging temporal grounding benchmark.

Why it should be added: Temporal audio grounding is a capability entirely absent from most local audio toolchains. Adding SpotSound would enable audiocpp users to query audio content by natural language and receive precise timestamped results—useful for indexing, editing, and analysis workflows. This would be a unique differentiator for the WebUI.

  1. YingMusic-Singer (https://huggingface.co/GiantAILab/YingMusic-Singer)
    Description: YingMusic-Singer is a fully diffusion-based, melody-driven singing voice synthesis (SVS) framework built on a Diffusion Transformer (DiT) architecture. It synthesizes arbitrary lyrics following any reference melody without relying on phoneme-level alignment or manual melody annotations. The model takes three inputs: an optional timbre reference, a melody-providing singing clip, and modified lyrics, achieving strong melody preservation and lyric adherence without manual alignment.

Why it should be added: Singing voice synthesis is a high-demand feature that no current audiocpp model addresses. YingMusic-Singer's zero-shot capability and annotation-free approach make it practical for real-world use. Adding it would open up music creation workflows and complement audiocpp's existing TTS offerings with a distinct creative capability.

  1. UniPASE (https://huggingface.co/Xiaobin-Rong/unipase)
    Description: UniPASE is a generative model for universal speech enhancement (USE) that restores speech signals from diverse distortions across multiple sampling rates. It extends the PASE framework with a distilled WavLM-based phonetic enhancement module and an adapter-vocoder pipeline, achieving high fidelity with minimal hallucination. It employs DeWavLM-Omni to generate enhanced phonetic representations that condition an adapter to refine acoustic representations, followed by a vocoder for waveform generation.

Why it should be added: Speech enhancement is a foundational preprocessing step for many audio workflows. UniPASE's ability to handle multiple sampling rates and diverse distortions in a single model makes it an efficient addition. Its low-hallucination design is critical for maintaining intelligibility, and it would pair well with audiocpp's existing ASR and diarization backends.

  1. MiDashengLM-Gen (https://huggingface.co/mispeech/midashenglm-gen)
    Description: MiDashengLM-Gen is an end-to-end framework that uses a pre-trained Large Language Model and audio tokenizer as the backbone, combined with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. It generates coherent 16 kHz audio scenes that simultaneously blend speech, music, sound effects, and environmental acoustics from structured text descriptions.

Why it should be added: This is described as the first end-to-end trained general text-to-audio model, with speech intelligibility approaching dedicated models. Its unified generation capability represents a fundamentally different paradigm from single-modality TTS or music models. Adding it would make audiocpp the first local inference engine to support unified audio scene generation, a compelling showcase feature.

Thank you for considering this request. I'm happy to help with testing or provide additional details on any of these models.

Best regards

Thanks for nominating these models! MiDashengLM-Gen is already in audio.cpp. It was selected to fill a gap in SFX generation support. The general integration priority is: newly released (good) models > models that fill task gaps > models that provide new building blocks > older models > variants of existing models. However, an unspoken rule is that this prioritization applies only to production-ready models from reputable teams. Academic research moves very fast, and the practical readiness of many newly released models is still questionable. (I know how it works 🙂 . I've served on conference program committees for the past five years, though in the security field.) It's unrealistic to keep up with every new model, especially given the risk of spending significant time on models that may never see real-world adoption. So, for academic research models, I generally invite the authors to integrate them into audio.cpp themselves rather than spending time on them myself, unless they've seen significant adoption. I do check HF downloads and Github issues of nominated models.

Sign up or log in to comment