Kai Li (李凯)
GitHub
108 repos
1.1k followers
Explained projects
Best-Audio-Paper-2025: The Year Audio AI Became an Omni-Model Race
A curated leaderboard that ranks the audio field three ways at once, and reveals why the strongest projects now win on research, demo quality, and open-source execution together.
8 min read · Apr 8, 2026
TFACM: The Cache That Lets Speech Separation Hear in Real Time
A causal separator from Tsinghua that swaps full hindsight for a rolling memory, and tries to erase the usual quality penalty of live audio.
12 min read · Apr 8, 2026
SonicSim: The Open-Source Simulator That Teaches AI How Moving Sound Actually Behaves
A synthetic audio stack that treats motion as the signal, not metadata, then turns that motion into a dataset and benchmark for speech enhancement and separation.
9 min read · Apr 8, 2026
speech-paper-daily-skill: The grumpy AI gatekeeper for speech research
A Markdown skill that turns arXiv monitoring into a checkpointed review pipeline, with a skeptical reviewer persona, recovery-safe caching, and doc publishing baked in.
8 min read · Apr 8, 2026
AudioTrust turns voice models into a trustworthiness test
A benchmark that scores audio systems on the failures text-only evals miss: spoofing, noise, accent bias, privacy leaks, hallucination, and safety drift.
12 min read · Apr 8, 2026
Apollo-data-preprocess: The tiny preprocessing repo that teaches Apollo what music is worth keeping
Apollo-data-preprocess is not a model repo. It is the filter that strips silence, slices tracks into learnable windows, and turns raw stems into HDF5 training fuel for high-fidelity music restoration.
11 min read · Apr 8, 2026
Swift-Net: The Speech Separator That Refuses to Cheat
A real-time audio-visual model that uses mouth motion, causal convolutions, and SRUs to pull one voice out of a crowd without peeking into the future.
8 min read · Apr 8, 2026
Gull-Codec-Training: a codec that listens in frequency first
This repo is the training scaffolding behind Gull, a generative audio codec that compresses and reconstructs sound from subbands instead of raw waveforms.
11 min read · Apr 8, 2026
TIGER Makes Speech Separation Small Enough to Matter
A look at how this Tsinghua-built repo uses time-frequency interleaving and gain extraction to chase strong separation without the usual model bloat.
10 min read · Apr 7, 2026
jusperlee.github.io: TermHub Turns a Portfolio Into an AI-Operable Content System
A terminal-style site, a Git-backed CMS, and an MCP server combine so an AI agent can help build and maintain a résumé-to-portfolio engine.
7 min read · Apr 7, 2026
Google-Scholar-Bibtex-Copy: The tiny script that chooses the right paper version
A one-file userscript turns Scholar’s citation maze into a background lookup, then hunts for the most canonical BibTeX entry before you ever open the popup.
7 min read · Apr 7, 2026
The Training-Free Context Engine: Inside JusperLee/BEAVER
How structure-aware "pages" and dual-path pooling are replacing token-level pruning to solve the LLM lost-in-the-middle problem.
7 min read · Mar 30, 2026
Yet to be explained
Dolphin
Python
188 stars
Explain
Hive
Python
174 stars
Explain
speech-paper-wechat
TypeScript
17 stars
Explain
claude-code-vscode-ui
4 stars
Explain
Referential-Dangling
Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression
Python
3 stars
Explain
Claude-Codex-Settings
Python
0 stars
Explain