1 726
Subscribers
+124 hours
-37 days
+1830 days
Data loading in progress...
Similar Channels
Tags Cloud
Incoming and Outgoing Mentions
---
---
---
---
---
---
Attracting Subscribers
October '26Oct '26
October '26
+3
in 0 channels
September '26
+54
in 1 channels
Get PRO
August '26
+69
in 1 channels
Get PRO
July '26
+55
in 0 channels
Get PRO
June '26
+29
in 1 channels
Get PRO
May '26
+27
in 0 channels
Get PRO
April '26
+36
in 0 channels
Get PRO
March '26
+20
in 1 channels
Get PRO
February '26
+38
in 0 channels
Get PRO
January '26
+23
in 0 channels
Get PRO
December '25
+16
in 0 channels
Get PRO
November '25
+25
in 0 channels
Get PRO
October '25
+26
in 1 channels
Get PRO
September '25
+45
in 0 channels
Get PRO
August '25
+18
in 0 channels
Get PRO
July '25
+18
in 0 channels
Get PRO
June '25
+28
in 1 channels
Get PRO
May '25
+34
in 0 channels
Get PRO
April '25
+34
in 1 channels
Get PRO
March '25
+40
in 1 channels
Get PRO
February '25
+36
in 0 channels
Get PRO
January '25
+20
in 1 channels
Get PRO
December '24
+56
in 2 channels
Get PRO
November '24
+37
in 1 channels
Get PRO
October '24
+137
in 3 channels
Get PRO
September '24
+47
in 2 channels
Get PRO
August '24
+51
in 1 channels
Get PRO
July '24
+62
in 2 channels
Get PRO
June '24
+34
in 1 channels
Get PRO
May '24
+41
in 1 channels
Get PRO
April '24
+33
in 1 channels
Get PRO
March '24
+55
in 2 channels
Get PRO
February '24
+56
in 2 channels
Get PRO
January '24
+77
in 1 channels
Get PRO
December '23
+96
in 3 channels
Get PRO
November '23
+47
in 2 channels
Get PRO
October '23
+37
in 1 channels
Get PRO
September '23
+23
in 0 channels
Get PRO
August '23
+24
in 0 channels
Get PRO
July '23
+34
in 0 channels
Get PRO
June '23
+20
in 0 channels
Get PRO
May '23
+79
in 0 channels
Get PRO
April '23
+51
in 0 channels
Get PRO
March '23
+15
in 0 channels
Get PRO
February '23
+6
in 0 channels
Get PRO
January '23
+11
in 0 channels
Get PRO
December '22
+12
in 0 channels
Get PRO
November '22
+22
in 0 channels
Get PRO
October '22
+14
in 0 channels
Get PRO
September '22
+25
in 0 channels
Get PRO
August '22
+53
in 0 channels
Get PRO
July '22
+9
in 0 channels
Get PRO
June '22
+16
in 0 channels
Get PRO
May '22
+14
in 0 channels
Get PRO
April '22
+26
in 0 channels
Get PRO
March '22
+9
in 0 channels
Get PRO
February '22
+8
in 0 channels
Get PRO
January '22
+10
in 0 channels
Get PRO
December '21
+16
in 0 channels
Get PRO
November '21
+11
in 0 channels
Get PRO
October '21
+16
in 0 channels
Get PRO
September '21
+3
in 0 channels
Get PRO
August '21
+24
in 0 channels
Get PRO
July '21
+13
in 0 channels
Get PRO
June '21
+16
in 0 channels
Get PRO
May '21
+220
in 0 channels
| Date | Subscriber Growth | Mentions | Channels | |
| 03 October | 0 | |||
| 02 October | +1 | |||
| 01 October | +2 |
Channel Posts
You can inject directly into KV cache ;)
Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
https://arxiv.org/abs/2609.30784
Popular thing in modern LLMs, Mostik is doing similar research, also
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
https://arxiv.org/abs/2510.03215
| 2 | https://kyutai.org/blog/2026-09-28-pocket-tts-drifting/ | 453 |
| 3 | Good Paper from Interspeech and dataset for game developers
https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset
https://ncai-official.github.io/speech/publications/designed-vocalizations-dataset/index.html
Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations such as monster growls and robotic voices underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, created by applying professional vocal effects processing to diverse vocal sources to produce paired original and effect-modified audio. We further provide a standardized test set with explicit seen/unseen splits over source types and preset styles to assess generalization under controlled conditions, together with baseline benchmark results for reproducible evaluation. | 637 |
| 4 | Interspeech 2026 starts today
https://www.isca-archive.org/interspeech_2026/
Let us be brave to read all this. Surprisingly, very few agentic papers. | 893 |
| 5 | Highly efficient audio inference engine
https://github.com/pegainfer-project/pega-omni
custom kernel for Mimi codec, etc, 128 Personaplex streams in realtime on GB300, should be like 30 streams on 4090 | 738 |
| 6 | Yodas3 was shared yesterday. 55TB, 1.1M hours of speech data
https://huggingface.co/datasets/espnet/yodas3
from previous experience not very useful to be honest | 691 |
| 7 | Some recent advanced TTS evaluation
Blog by @altsoph from Inworld
https://altsoph.substack.com/p/wtf-is-voice-steering
J-HARD-TTS-Eval, a benchmark designed to evaluate the robustness of autoregressive Japanese Text-To-Speech (TTS) models (V2 is coming)
https://github.com/Parakeet-Inc/J-HARD-TTS-Eval
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
https://arxiv.org/abs/2505.23009 | 647 |
| 8 | https://github.com/OPPO-Mente-Lab/CuteTTS | 602 |
| 9 | Facebook's Muse Transcribe is 1st on AA leaderboard and 25th on Huggingface leaderboard
https://x.com/EricBezzam/status/2102368894929514711 | 734 |
| 10 | https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark
Benchmark for the speech recognition with background speech
Data. 265 recordings in real offices, four call-center floors and cars. Each work and call-center recording follows a fixed structure: primary alone, both speakers, secondary alone, primary alone. Speakers read scripts to give exact ground truth; environments, devices and background speech are real, with no synthetic mixing. Segments are hand-labeled primary / mix / secondary / noise.
Evaluation. 11 STT configurations across 9 engines, streaming and batch, run on raw audio and after voice isolation. WER is computed with jiwer 4.0.0 at corpus level (errors pooled across files, not averaged). Reference and hypothesis go through the same normalization: regex splitting of alphanumeric tokens, then NVIDIA NeMo WFST text normalization, then lowercasing, contraction expansion, punctuation and filler removal. Perceptual quality is scored with DNSMOS-C on primary-speaker segments only.
Results.
Corpus WER: 23.29% → 6.26%
All 11 configurations improve
Engine spread narrows from 17–37% to 4–8%
Clean phone audio regresses slightly: 3.48% → 3.91%
Limitation. On clean narrowband audio there's little to remove, and isolation removes some of the primary signal. We report it rather than filtering it out. Segment-level labels also let you check for deletions specifically: a model that correctly outputs silence and one that drops primary-speaker words can post similar corpus WER.
Dataset: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset
Model outputs: https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchmark | 682 |
| 11 | Parakeet v3 ternary quantization with good accuracy (178mb) 113xRT on CPU
https://huggingface.co/moondream/parakeet-redux | 946 |
| 12 | https://github.com/SamsungLabs/samsone
Samsone is a family of open Small Audio Language Models (SALMs) for efficient audio understanding. The release includes Samsone-99M, Samsone-134M, and Samsone-356M checkpoints for both on-device and server-side inference.
https://arxiv.org/abs/2609.21666
Samsone: A Family of Open Small Audio Language Models for On-Device Inference
Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski
The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone. | 954 |
| 13 | Serious issues with Qwen3-TTS stability, long texts and audio prompts break things
https://arxiv.org/abs/2609.16989
https://x.com/RmdW_W/status/2102032894600401231
Taming Long-form Text-to-Speech
Rongxiang Wang, Berkin Durmus, Aysegul Orhon, Eduardo Pacheco, Atila Orhon
Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and speaker similarity (SIM) on short-form prompts but significantly deteriorate when used with long-form prompts. We propose Localized Attention-Constrained Inference (LACI), an inference-only method to detect TTS errors in near real-time, roll back to the error onset and regenerate with temporary guardrails, adding negligible computational overhead. Using LACI, we improve worst-of-N WER across 10 RNG seeds for Qwen3-TTS-0.6B from 35.2% to 3.4% on prompts longer than 1500 words, even surpassing its short-form reliability of 5.4\% on prompts with fewer than 500 words. To demonstrate the efficacy of LACI on voice cloning reliability, we propose a sliding-window version of the SIM metric that we call wSIM. wSIM exposes several novel failure patterns that are not captured by SIM. LACI improves worst-of-N wSIM from 0.01 to 0.47 on 120 seconds of reference audio while reducing the rate of catastrophic generations with WER above 30% from 26% to below 1% | 769 |
| 14 | Voice agents related benchmarks somehow passed our attention, here are few recent ones:
https://github.com/sierra-research/tau2-bench
https://github.com/ServiceNow/eva
https://arxiv.org/abs/2603.13686 | 995 |
| 15 | An advanced dictation model from Whispr, an interesting part is GPRO over user corrections
https://wisprflow.ai/canto | 972 |
| 16 | Another small TTS
https://github.com/lab-emi/GrainSpeech
https://lab-emi.github.io/GrainSpeech/
https://arxiv.org/abs/2609.18856 | 934 |
| 17 | We see a huge decline of the interest in speech technology in China but a rise in Europe and India. Recently I discussed it with one of my Chinese friends - he confirms that nobody is interested anymore in plain speech. It has to be multimodal - video, music generation, etc.
China is ahead of time here. | 930 |
| 18 | https://narilabs.com/blog/making-pyannote-ultrafast/ | 1 054 |
| 19 | 2.1k parameters VAD
https://github.com/AydinAdnan/PulseVAD
reimplementation of KiloVAD
https://arxiv.org/abs/2607.25870 | 1 148 |
| 20 | 10000xRT for Zipformer on A100, 18000xRT on H200 with specialized CUDA tricks
https://github.com/SoundsGoodAI/fast-gpu-asr | 930 |
