12-day longest streak
-
DiffRhythm
Di♪♪Rhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Python ★ 2.3k 8mo agoExplain → -
OSUM
OSUM & OSUM-EChat, open speech understanding model and empathetic spoken chatbot based on it, open-sourced by ASLP@NPU.
Python ★ 496 8mo agoExplain → -
WenetSpeech-Yue
A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
Python ★ 345 1mo agoExplain → -
SongEval
A song aesthetic evaluation toolkit trained on SongEval.
Python ★ 315 3mo agoExplain → -
MeanVC
A Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
Python ★ 298 6mo agoExplain → -
VoiceSculptor
An instruct text-to-speech solution based on LLaSA and CosyVoice2 developed by the ASLP lab and collaborators.
Python ★ 250 5mo agoExplain → -
WenetSpeech-Chuan
Official repository for the WenetSpeech-Chuan dataset.
Python ★ 217 15d agoExplain → -
WenetSpeech-Wu-Repo
A Large-scale Wu Dialect Speech Corpus with Multi-dimensional Annotations
Python ★ 171 5mo agoExplain → -
DiffRhythm2 ⑂
Di♪♪Rhythm 2: Efficient And High Fidelity Song Generation Via Block Flow Matching
Python ★ 165 8mo agoExplain → -
SongFormer
No description.
Python ★ 164 2mo agoExplain → -
Easy-Turn
Open-Source Turn-Taking Detection Model and Dataset for Full-Duplex Spoken Dialogue Systems
Python ★ 122 6mo agoExplain → -
Speaker-Reasoner
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
Python ★ 93 2mo agoExplain → -
SenSE
Official code of SenSE.
Python ★ 90 9mo agoExplain → -
YingMusic-Singer-Plus
YingMusic-Singer-Plus: Controllable Singing Voice Synthesis with Flexible Lyric Manipulation and Annotation-free Melody Guidance
Python ★ 84 3mo agoExplain → -
FlashTTS
Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation
Python ★ 67 1mo agoExplain → -
LLaSA_Plus
Llasa Speed Up
Python ★ 64 6mo agoExplain → -
Hum-Dial
ICASSP2026 HumDial Challenge
Python ★ 51 2mo agoExplain → -
MINT-Bench
No description.
Python ★ 49 2mo agoExplain → -
LLaSE-G1 ⑂
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
Python ★ 47 1y agoExplain → -
OmniCodec
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic–Acoustic Disentanglement
Python ★ 46 3mo agoExplain → -
HumDial-FDBench
The Full-Duplex Interaction Track of the ICASSP 2026 Human-like Spoken Dialogue Systems Challenge aims to advance the evaluation of full-duplex dialogue systems by in- troducing a dual-channel dialogue dataset of real human- recorded conversations.
Python ★ 36 3mo agoExplain → -
FastTurn
No description.
★ 35 2mo agoExplain → -
OSUM-Pangu
An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs
Python ★ 33 4mo agoExplain → -
ArxivWatcher
No description.
Python ★ 32 1mo agoExplain → -
MeanVC2
No description.
★ 29 1mo agoExplain → -
FMSU-Bench
Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
Python ★ 25 2mo agoExplain → -
MSU-Bench
Open repository of "MSU-Bench: Towards Understanding the Conversational Multi-Speaker Scenarios"
Python ★ 20 23d agoExplain → -
M7-TTS
M7-TTS: A Mini-Scale Multilingual and Multi-Dialect Text-to-Speech Language Model with Mimi codec and Multi Token Prediction
★ 20 4mo agoExplain → -
Music-Semantic-VAE
No description.
★ 19 2mo agoExplain → -
Smart-Glass-Challenge
No description.
Shell ★ 18 1mo agoExplain → -
SmartGlasses
This challenge focuses on evaluating speech recognition and semantic understanding capabilities of AI glasses in complex real-world environments.
★ 18 1mo agoExplain → -
C2SER ⑂
We propose C2SER, a novel audio-language model designed to enhance the stability and accuracy of speech emotion recognition through contextual perception and chain of Thought (CoT).
★ 17 1y agoExplain → -
HumDial-Challenge
No description.
TypeScript ★ 15 1mo agoExplain → -
Automatic-Song-Aesthetics-Evaluation-Challenge
No description.
TypeScript ★ 15 7mo agoExplain → -
Stream-TN
No description.
HTML ★ 13 4mo agoExplain → -
MINT-Bench-Demo ⑂
Demo page of MINT-Bench
★ 13 2mo agoExplain → -
ChinaVoices-Challenge
No description.
Python ★ 11 1d agoExplain → -
HumDial-EIBench
No description.
Python ★ 10 2mo agoExplain → -
SoulX-Podcast ⑂
SoulX-Podcast is an inference codebase by the Soul AI team for generating high-fidelity podcasts from text.
★ 9 9mo agoExplain → -
Llasa-1B-Yue-Updated
No description.
Python ★ 9 8mo agoExplain → -
gen-se ⑂
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
★ 6 1y agoExplain → -
Speaker-Reasoner-Demo
No description.
HTML ★ 5 3mo agoExplain → -
SongEval_anonymous
No description.
Python ★ 4 10mo agoExplain → -
YingMusic-Singer-Plus-Demo
No description.
HTML ★ 3 3mo agoExplain → -
DiffRhythm2.github.io
No description.
JavaScript ★ 3 8mo agoExplain → -
flashtts_demo ⑂
demopages of flashtts
HTML ★ 2 1mo agoExplain → -
osum-echat.github.io
OSUM-EChat demo page
HTML ★ 2 11mo agoExplain → -
MLC-SLM-Baseline ⑂
The project is associated with the recently-launched INTERSPEECH 2025 Workshop on Multilingual Conversational Speech Language Model (MLC-SLM) to provide participants with baseline systems for speech recognition and speaker diarization in multilingual conversational scenario.
★ 2 1y agoExplain → -
DiffRhythm.github.io
No description.
HTML ★ 2 1y agoExplain → -
msu-bench.github.io
Online demo page for "MSU-Bench: Towards Understanding the Conversational Multi-Speaker Scenarios"
JavaScript ★ 1 23d agoExplain → -
OSUM.github.io
No description.
HTML ★ 1 1y agoExplain →
No repos match these filters.
More creators on gitmyhub
JakeWharton lucidrains rafaballerini hiteshchoudhary IDouble