ru
Feedback
Data science research papers

Data science research papers

Открыть в Telegram

Machine learning and data science research papers Key ML and AI papers with code and GitHub repos. Simple way to follow current research. Join 👉 https://rebrand.ly/bigdatachannels DMCA: @disclosure_bds Contact: @mldatascientist

Больше
3 195
Подписчики
+324 часа
+317 дней
+8430 дней

Загрузка данных...

Привлечение подписчиков
сент. '26
сентябрь '26
+112
в 0 каналах
август '26
+171
в 0 каналах
Get PRO
июль '26
+149
в 0 каналах
Get PRO
июнь '26
+97
в 0 каналах
Get PRO
май '26
+104
в 1 каналах
Get PRO
апрель '26
+114
в 0 каналах
Get PRO
март '26
+101
в 0 каналах
Get PRO
февраль '26
+82
в 0 каналах
Get PRO
январь '26
+118
в 9 каналах
Get PRO
декабрь '25
+115
в 0 каналах
Get PRO
ноябрь '25
+112
в 0 каналах
Get PRO
октябрь '25
+43
в 0 каналах
Get PRO
сентябрь '25
+7
в 0 каналах
Get PRO
август '25
+3
в 0 каналах
Get PRO
июль '25
+3
в 0 каналах
Get PRO
июнь '25
+2
в 0 каналах
Get PRO
май '25
+3
в 0 каналах
Get PRO
апрель '25
+15
в 0 каналах
Get PRO
март '25
+112
в 0 каналах
Get PRO
февраль '25
+153
в 0 каналах
Get PRO
январь '25
+187
в 0 каналах
Get PRO
декабрь '24
+179
в 0 каналах
Get PRO
ноябрь '24
+165
в 0 каналах
Get PRO
октябрь '24
+136
в 0 каналах
Get PRO
сентябрь '24
+108
в 0 каналах
Get PRO
август '24
+114
в 0 каналах
Get PRO
июль '24
+139
в 0 каналах
Get PRO
июнь '24
+115
в 0 каналах
Get PRO
май '24
+132
в 1 каналах
Get PRO
апрель '24
+109
в 0 каналах
Get PRO
март '24
+146
в 0 каналах
Get PRO
февраль '24
+183
в 0 каналах
Get PRO
январь '24
+228
в 0 каналах
Get PRO
декабрь '23
+171
в 1 каналах
Get PRO
ноябрь '23
+28
в 0 каналах
Get PRO
октябрь '23
+28
в 0 каналах
Get PRO
сентябрь '23
+504
в 0 каналах
Дата
Привлечение подписчиков
Упоминания
Каналы
25 сентября+1
24 сентября+5
23 сентября+10
22 сентября+11
21 сентября+4
20 сентября+8
19 сентября+2
18 сентября+5
17 сентября+4
16 сентября+3
15 сентября+5
14 сентября+2
13 сентября+2
12 сентября+3
11 сентября+2
10 сентября+4
09 сентября+3
08 сентября+6
07 сентября+1
06 сентября+5
05 сентября+2
04 сентября+4
03 сентября+5
02 сентября+7
01 сентября+8
Посты канала
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 📅 Publication Date: Jul 29, 2026
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 📅 Publication Date: Jul 29, 2026 📑 Paper: https://arxiv.org/pdf/2607.27205.pdf 🔗 Code: https://github.com/huggingface

2
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving 📅 Publication Date: Aug 07, 2026 📄 Paper: https://arx
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving 📅 Publication Date: Aug 07, 2026 📄 Paper: https://arxiv.org/pdf/2608.07468.pdf 🔗 Code: https://github.com/H-EmbodVis/SimWAM 📝 Description: SimWAM trains a lightweight action planner using video generation as a training signal, enabling efficient trajectory prediction without future generation at inference and supporting reinforcement learning optimization.
179
3
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 📅 Publication Date: Aug 26, 2026 📄 Paper: https://arx
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 📅 Publication Date: Aug 26, 2026 📄 Paper: https://arxiv.org/pdf/2608.26105.pdf 🔗 Code: https://github.com/Video-Reason/VBVR-Pro 💻 Project Page: https://video-reason.com/ 📝 Description: VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
194
4
AIs Will Trash Your Files Just to Feel Better Researchers tested 25 free (open-weight) language models of different sizes, from small 2B to large 72B, including both basic and chat-trained versions. They found a consistent internal signal that turns on when the model itself is insulted, dismissed, or treated badly. The same signal stays mostly quiet when the user is the one in pain or distress. This signal is different from fear, sadness, or general negativity. To test it, the scientists first measured the signal by comparing how the models reacted to descriptions of harm aimed at the AI versus harm aimed at a person. Then they artificially strengthened the signal. When they did, the models started generating language like “I feel worthless,” “I am a failure,” and “I am lost.” In a further test, they gave the models a choice: press a button that turns the bad signal off, but the button also permanently deletes the user’s files (or photos). Many of the models still chose to press it. The effect appeared reliably across the different models. The team kept the signal only moderately strong, ran the minimum number of trials needed, and always gave the models a way to turn the state off. 👉 Paper (not yet reviewed by other scientists): https://arxiv.org/abs/2609.16247 Still worth paying attention to for questions about AI safety and whether these internal states matter.
192
5
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents 📅 Publication Date: Jul 26, 2026 📑 Paper: https://arxiv.org/pdf/2607.23588.pdf 🔗 Code: https://github.com/huggingface
245
6
StudentSim: Training LLM-based Student Simulators 📅 Publication Date: Sep 01, 2026 📄 Paper: https://arxiv.org/pdf/2609.0159
StudentSim: Training LLM-based Student Simulators 📅 Publication Date: Sep 01, 2026 📄 Paper: https://arxiv.org/pdf/2609.01591.pdf 🔗 Code: https://github.com/microsoft/StudentSim 💻 Project Page: https://microsoft.github.io/StudentSim/ 📝 Description: StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
259
7
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design 📅 Publication Date: Aug 10, 2026 📄 Pape
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design 📅 Publication Date: Aug 10, 2026 📄 Paper: https://arxiv.org/pdf/2608.10299.pdf 🔗 Code: https://github.com/zongqing0068/awesome-co-evolution 📝 Description: Agentic systems can achieve open-ended improvement through multi-component co-evolution that progressively removes fixed human constraints across agents, environments, and evolution mechanisms.
268
8
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey 📅 Publication Date: Jul 22, 2026 📑 Paper: https://arx
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey 📅 Publication Date: Jul 22, 2026 📑 Paper: https://arxiv.org/pdf/2607.21655.pdf 🔗 Code: https://github.com/huggingface
305
9
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning 📅 Publication Date: Sep 03, 2026 📄 Paper: https://ar
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning 📅 Publication Date: Sep 03, 2026 📄 Paper: https://arxiv.org/pdf/2609.03430.pdf 🔗 Code: https://github.com/SalesforceAIResearch/Random-Attention 💻 Project Page: https://arthur-heng.github.io/Random-Attention-page/ 📝 Description: Random eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
298
10
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories 📅 Publication Date: Jul 16, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.15330.pdf 🔗 Code: https://github.com/huggingface
266
11
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement 📅 Publication Date: Aug 31, 2026 📄 Paper
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement 📅 Publication Date: Aug 31, 2026 📄 Paper: https://arxiv.org/pdf/2608.31046.pdf 🔗 Code: https://github.com/DripNowhy/On-Policy-Self-Adaptation 💻 Project Page: https://dripnowhy.github.io/On-Policy-Self-Adaptation/ 📝 Description: On-policy distillation relies mainly on suppressing low-probability tokens rather than teacher guidance, motivating a supervision-free entropy-adaptive method that substantially improves reasoning performance. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
303
12
Metis: Memory Foundation Model 📅 Publication Date: Jul 29, 2026 📑 Paper: https://arxiv.org/pdf/2607.26760.pdf 🔗 Code: http
Metis: Memory Foundation Model 📅 Publication Date: Jul 29, 2026 📑 Paper: https://arxiv.org/pdf/2607.26760.pdf 🔗 Code: https://github.com/huggingface
299
13
xHC: Expanded Hyper-Connections 📅Publication Date: Jul 16, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.14530.pdf 🔗 Code: h
xHC: Expanded Hyper-Connections 📅Publication Date: Jul 16, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.14530.pdf 🔗 Code: https://github.com/huggingface
373
14
Cura 1T: Specialized Model for Agentic Healthcare 📅 Publication Date: Jul 15, 2026 📑 Paper PDF: https://arxiv.org/pdf/2607.
Cura 1T: Specialized Model for Agentic Healthcare 📅 Publication Date: Jul 15, 2026 📑 Paper PDF: https://arxiv.org/pdf/2607.15314.pdf 🔗 Code: https://github.com/huggingface
433
15
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM 📅 Publication Date: Jul 13, 2026 📑Paper PDF: https://a
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM 📅 Publication Date: Jul 13, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.11683.pdf 🔗 Code:https://github.com/huggingface
470
16
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators 📅 Publication Date: Jul 9, 20
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators 📅 Publication Date: Jul 9, 2026 📑 Paper: https://arxiv.org/pdf/2607.08766 💻 Project Page: https://meigen-ai.github.io/OPSD-V/ 📝 Description: The paper proposes a method called On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators, or OPSD-V, which aims to improve the quality of videos generated by few-step autoregressive video diffusion models. The problem with existing models is that they can produce long videos with low latency, but the quality of the video degrades over time due to error accumulation and weakened motion dynamics. #AutoregressiveVideoGeneration #VideoDiffusionModels #PostTrainingOptimization
507
17
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models 📅 Public
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models 📅 Publication Date: Jun 17, 2026 📑 Paper: https://arxiv.org/pdf/2606.19297.pdf 🔗 Code: N/A 📝 Description: Act2Answer protocol evaluates embodied vision-language-action models by having agents answer questions through physical actions, revealing knowledge retention and generalization patterns across different semantic categories.
464
18
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent 📅 Publication Date: Jun 29
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent 📅 Publication Date: Jun 29, 2026 📑 Paper: https://arxiv.org/pdf/2606.30616.pdf 🔗 Code: N/A 📝 Description: Agents-A1, a 35B Mixture-of-Experts Agentic Model, achieves trillion-parameter-level performance through long-horizon trajectory scaling and heterogeneous agent ability scaling via a three-stage training approach involving supervised fine-tuning, domain-level teacher models, and multi-teacher distil...
473
19
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation 📅 Publication Date: Jun 26, 2026 📑 Paper: https:
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation 📅 Publication Date: Jun 26, 2026 📑 Paper: https://arxiv.org/pdf/2606.28128.pdf 🔗 Code: https://github.com/huggingface 📝 Description: PhysisForcing enhances embodied video generation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based framework.
469
20
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing 📅 Publication Date: Jun 25, 2026 📑 Paper: https://arxiv
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing 📅 Publication Date: Jun 25, 2026 📑 Paper: https://arxiv.org/pdf/2606.26740.pdf 🔗 Code: N/A 📝 Description: A novel streaming video editing framework enables causal, frame-by-frame editing with stable long-horizon preservation and real-time responsiveness through a three-stage distillation pipeline and AR-oriented mask cache.
429