uz
Feedback
Speech Technology

Speech Technology

Kanalga Telegram’da oā€˜tish
1 726
Obunachilar
+124 soatlar
+117 kun
+3630 kun
Obunachilarni jalb qilish
Sentabr '26
Sentabr '26
+36
1 kanalda
Avgust '26
+69
1 kanalda
Get PRO
Iyul '26
+55
0 kanalda
Get PRO
Iyun '26
+29
1 kanalda
Get PRO
May '26
+27
0 kanalda
Get PRO
Aprel '26
+36
0 kanalda
Get PRO
Mart '26
+20
1 kanalda
Get PRO
Fevral '26
+38
0 kanalda
Get PRO
Yanvar '26
+23
0 kanalda
Get PRO
Dekabr '25
+16
0 kanalda
Get PRO
Noyabr '25
+25
0 kanalda
Get PRO
Oktabr '25
+26
1 kanalda
Get PRO
Sentabr '25
+45
0 kanalda
Get PRO
Avgust '25
+18
0 kanalda
Get PRO
Iyul '25
+18
0 kanalda
Get PRO
Iyun '25
+28
1 kanalda
Get PRO
May '25
+34
0 kanalda
Get PRO
Aprel '25
+34
1 kanalda
Get PRO
Mart '25
+40
1 kanalda
Get PRO
Fevral '25
+36
0 kanalda
Get PRO
Yanvar '25
+20
1 kanalda
Get PRO
Dekabr '24
+56
2 kanalda
Get PRO
Noyabr '24
+37
1 kanalda
Get PRO
Oktabr '24
+137
3 kanalda
Get PRO
Sentabr '24
+47
2 kanalda
Get PRO
Avgust '24
+51
1 kanalda
Get PRO
Iyul '24
+62
2 kanalda
Get PRO
Iyun '24
+34
1 kanalda
Get PRO
May '24
+41
1 kanalda
Get PRO
Aprel '24
+33
1 kanalda
Get PRO
Mart '24
+55
2 kanalda
Get PRO
Fevral '24
+56
2 kanalda
Get PRO
Yanvar '24
+77
1 kanalda
Get PRO
Dekabr '23
+96
3 kanalda
Get PRO
Noyabr '23
+47
2 kanalda
Get PRO
Oktabr '23
+37
1 kanalda
Get PRO
Sentabr '23
+23
0 kanalda
Get PRO
Avgust '23
+24
0 kanalda
Get PRO
Iyul '23
+34
0 kanalda
Get PRO
Iyun '23
+20
0 kanalda
Get PRO
May '23
+79
0 kanalda
Get PRO
Aprel '23
+51
0 kanalda
Get PRO
Mart '23
+15
0 kanalda
Get PRO
Fevral '23
+6
0 kanalda
Get PRO
Yanvar '23
+11
0 kanalda
Get PRO
Dekabr '22
+12
0 kanalda
Get PRO
Noyabr '22
+22
0 kanalda
Get PRO
Oktabr '22
+14
0 kanalda
Get PRO
Sentabr '22
+25
0 kanalda
Get PRO
Avgust '22
+53
0 kanalda
Get PRO
Iyul '22
+9
0 kanalda
Get PRO
Iyun '22
+16
0 kanalda
Get PRO
May '22
+14
0 kanalda
Get PRO
Aprel '22
+26
0 kanalda
Get PRO
Mart '22
+9
0 kanalda
Get PRO
Fevral '22
+8
0 kanalda
Get PRO
Yanvar '22
+10
0 kanalda
Get PRO
Dekabr '21
+16
0 kanalda
Get PRO
Noyabr '21
+11
0 kanalda
Get PRO
Oktabr '21
+16
0 kanalda
Get PRO
Sentabr '21
+3
0 kanalda
Get PRO
Avgust '21
+24
0 kanalda
Get PRO
Iyul '21
+13
0 kanalda
Get PRO
Iyun '21
+16
0 kanalda
Get PRO
May '21
+220
0 kanalda
Sana
Obunachilarni jalb qilish
Esdaliklar
Kanallar
14 Sentabr+2
13 Sentabr+1
12 Sentabr+3
11 Sentabr+6
10 Sentabr0
09 Sentabr0
08 Sentabr+6
07 Sentabr+2
06 Sentabr+4
05 Sentabr+1
04 Sentabr+5
03 Sentabr+3
02 Sentabr+2
01 Sentabr+1
Kanal postlari
2.1k parameters VAD https://github.com/AydinAdnan/PulseVAD reimplementation of KiloVAD https://arxiv.org/abs/2607.25870

2
10000xRT for Zipformer on A100, 18000xRT on H200 with specialized CUDA tricks https://github.com/SoundsGoodAI/fast-gpu-asr
350
3
Just another reminder RL is important https://x.com/kadirnardev/status/2098425551203475665
334
4
Kokoro distill 82M -> 7M params https://huggingface.co/spaces/oddadmix/Kokoro-7M-Distill-Demo
384
5
https://x.com/unilightwf/status/2098261200480174123 https://arxiv.org/abs/2603.14328 Cross-lingual cloning TTS is still a big
https://x.com/unilightwf/status/2098261200480174123 https://arxiv.org/abs/2603.14328 Cross-lingual cloning TTS is still a big problem, some accent metrics demo interesting results Also worth checking https://iwslt.org/2026/voice-cloning with some useful data https://huggingface.co/datasets/ymoslem/acl-6060
433
6
My friend Miro recommended me Orukeet model https://github.com/Oruk-AI/orukeet https://arxiv.org/abs/2609.10054 it is indeed a good parakeet improvement, about 10% better. Interesting that people left scaling and return back to in-depth architecture analysis. Oruk.AI does some other nice things, for example a visualization of emotion representation in different layers of speech models https://x.com/OrukLabs/status/2073457781018087473 https://oruk.ai/research/how-models-represent-speech
477
7
Interesting TTS, claims 14ms TTFA https://github.com/X-Square-Robot/X2Streaming-TTS
507
8
Suplime results on Russian telephony data, good results actually second after Diarizen Large A bit slow though
Suplime results on Russian telephony data, good results actually second after Diarizen Large A bit slow though
794
9
Pyannote improved https://github.com/rewayai/suplime/
715
10
https://huggingface.co/tencent/AuK AuKĀ is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface.
611
11
ParsVoice, the largest open-source Persian speech dataset, along with a TTS model and an open-source processing pipeline is released. The paper has also been accepted as a main conference paper at EMNLP 2026. Paper: https://arxiv.org/abs/2510.10774 Dataset: https://huggingface.co/datasets/MohammadJRanjbar/ParsVoice TTS model: https://huggingface.co/MohammadJRanjbar/ParsVoice-XTTS Code & pipeline: https://github.com/MohammadJRanjbar/ParsVoice
783
12
Bodhan AI together with AI4Bharat recently released a great update on Indic ASR https://bodhan.ai/research/blogs/indic-transcribe
708
13
Scicom from Malaysia tries Ascend 910B3 https://github.com/Scicom-AI-Enterprise-Organization/TTS-API-Neucodec/blob/main/ASCEND_910B3_PRECISION_REPORT.md
690
14
https://arxiv.org/abs/2609.01246v1 Ready to Speak: Aligning LLMs for TTS-Friendly Text Generation Thibaut Thonet,Ā Jos Rozen,Ā Laurent Besacier Current Large Language Models (LLMs) are primarily optimized for written text, often producing outputs that are grammatically correct and helpful yet poorly suited for spoken delivery via Text-to-Speech (TTS). In this work, we study how to make LLMs natively generate TTS-friendly text, which we frame as a preference alignment problem: instead of relying on downstream rewriting modules, we directly align LLMs to generate text optimized for spoken delivery. We introduce two preference datasets spanning different target domains, CORA and Recipe, which contain paired TTS-friendly and TTS-unfriendly responses. We further propose an evaluation suite combining a pattern-based heuristic metric, a TTS→ASR evaluation pipeline, and a MUSHRA listening study with human judges. Our experiments compare the recently proposed Feature-aware Sampling and Tuning (FaST) framework -- leveraging interpretable features instead of a black-box reward model -- against an array of alignment baselines on the TTS-friendly generation task. Notably, we found that FaST achieves the best overall tradeoff between TTS-friendliness and helpfulness across various settings. We also identified a strong correlation between our different metrics, highlighting the ability to reliably assess TTS-friendliness via an efficient heuristic.
1 170
15
Quite an obvious but so ignored by industry before. There is certainly no need to clone from 3 seconds https://www.linkedin.com/posts/soniox_soniox-texttospeech-voiceai-activity-7501594413660401664-p9Cq
660
16
https://www.nature.com/articles/srep12881 Human starts to plan answer 2 seconds before question end
952
17
We know human scores are useless but anyway https://x.com/datapointai/status/2094829412625654141 today, we're releasing the largest open-source human audio preferences dataset, focused on the customer support use-case - 300K+ annotations by real people - 15 SOTA TTS models ranked (Sonic 3.6, Grok TTS, Simba 3.2, Eleven Labs v3) - 8 categories (IVR menus, empathy, escalations, refunds etc) dataset + benchmark + frontier plot below:
896
18
https://turnbench.sesame.com/
923
19
You can download really huge datasets these days https://x.com/ahochlehnert/status/2092648676829413778 LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training. - 1.3B video URLs from CommonCrawl - 80M downloaded videos - 10M video hours - 55M captioned clips - 300M frame-caption pairs This repository containsĀ 1.7 million audio clipsĀ taken from BVD-V-55M and sampled for uniqueness of the source video, so that the subset maximises source diversity rather than clip count. Each clip comes with a caption, its language, and the timestamps locating it in the source video. https://huggingface.co/datasets/laion/BVD-A-1.7M
1 305
20
https://huggingface.co/BreezeBlue/Breeze-TTS-2 https://breezeblue.ai/breeze-tts-2 English/Chinese only but really good quality
1 032