Stars
Localized watermarking for AI-generated speech audios, with SOTA on robustness and very fast detector
Open and efficient video and image watermarking
A gallery that showcases on-device ML/GenAI use cases and allows people to try and use models locally.
Open-source realtime voice agent server in Go with WebRTC (WHIP), barge-in, streaming STT/LLM/TTS pipelines, plugin system, multi-language SDKs, SIP telephony, ESP32 support & fully local mode.
Single-file C++ TTS runtime for Pocket TTS with ONNX Runtime — voice cloning, streaming, HTTP server, FFI C API
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Andr…
🔥Android无障碍服务(AccessibilityService)开发框架,Android自动化脚本框架,快速开发复杂自动化任务、远程协助、监听等
Recurrent neural network for audio noise reduction
Recurrent neural network for audio noise reduction
A SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD
CLI11 is a command line parser for C++11 and beyond that provides a rich feature set with a simple and intuitive interface.
Why is this running? Trace any process, port, container, or file back to what started it - CLI + TUI.
Apport intercepts Program crashes, collects debugging information about the crash and the operating system environment, and sends it to bug trackers in a standardized form. It also offers the user …
The official implementation of GTCRN, an ultra-lightweight SE model.
C/C++ WebRTC network library featuring Data Channels, Media Transport, and WebSockets
A curated list of awesome remote jobs and resources. Inspired by https://github.com/vinta/awesome-python
实现Linux Wayland下腾讯会议屏幕共享(非虚拟相机). Hook library that enables screenshare with Tencent Wemeet on Linux Wayland, without the need of using virtual cameras.
GoogleTest - Google Testing and Mocking Framework
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
#1 PDF Application on GitHub that lets you edit PDFs on any device anywhere
Python wrapper for OpenFST and its extensions from Kaldi. Also support reading/writing ark/scp files
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.


