es
Feedback
Data science research papers

Data science research papers

Ir al canal en Telegram

Machine learning and data science research papers Key ML and AI papers with code and GitHub repos. Simple way to follow current research. Join 👉 https://rebrand.ly/bigdatachannels DMCA: @disclosure_bds Contact: @mldatascientist

Mostrar más
3 195
Suscriptores
+324 horas
+317 días
+8430 días
Atraer Suscriptores
sep '26
septiembre '26
+112
en 0 canales
agosto '26
+171
en 0 canales
Get PRO
julio '26
+149
en 0 canales
Get PRO
junio '26
+97
en 0 canales
Get PRO
mayo '26
+104
en 1 canales
Get PRO
abril '26
+114
en 0 canales
Get PRO
marzo '26
+101
en 0 canales
Get PRO
febrero '26
+82
en 0 canales
Get PRO
enero '26
+118
en 9 canales
Get PRO
diciembre '25
+115
en 0 canales
Get PRO
noviembre '25
+112
en 0 canales
Get PRO
octubre '25
+43
en 0 canales
Get PRO
septiembre '25
+7
en 0 canales
Get PRO
agosto '25
+3
en 0 canales
Get PRO
julio '25
+3
en 0 canales
Get PRO
junio '25
+2
en 0 canales
Get PRO
mayo '25
+3
en 0 canales
Get PRO
abril '25
+15
en 0 canales
Get PRO
marzo '25
+112
en 0 canales
Get PRO
febrero '25
+153
en 0 canales
Get PRO
enero '25
+187
en 0 canales
Get PRO
diciembre '24
+179
en 0 canales
Get PRO
noviembre '24
+165
en 0 canales
Get PRO
octubre '24
+136
en 0 canales
Get PRO
septiembre '24
+108
en 0 canales
Get PRO
agosto '24
+114
en 0 canales
Get PRO
julio '24
+139
en 0 canales
Get PRO
junio '24
+115
en 0 canales
Get PRO
mayo '24
+132
en 1 canales
Get PRO
abril '24
+109
en 0 canales
Get PRO
marzo '24
+146
en 0 canales
Get PRO
febrero '24
+183
en 0 canales
Get PRO
enero '24
+228
en 0 canales
Get PRO
diciembre '23
+171
en 1 canales
Get PRO
noviembre '23
+28
en 0 canales
Get PRO
octubre '23
+28
en 0 canales
Get PRO
septiembre '23
+504
en 0 canales
Fecha
Crecimiento de Suscriptores
Menciones
Canales
25 septiembre+1
24 septiembre+5
23 septiembre+10
22 septiembre+11
21 septiembre+4
20 septiembre+8
19 septiembre+2
18 septiembre+5
17 septiembre+4
16 septiembre+3
15 septiembre+5
14 septiembre+2
13 septiembre+2
12 septiembre+3
11 septiembre+2
10 septiembre+4
09 septiembre+3
08 septiembre+6
07 septiembre+1
06 septiembre+5
05 septiembre+2
04 septiembre+4
03 septiembre+5
02 septiembre+7
01 septiembre+8
Publicaciones del Canal
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 📅 Publication Date: Jul 29, 2026
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 📅 Publication Date: Jul 29, 2026 📑 Paper: https://arxiv.org/pdf/2607.27205.pdf 🔗 Code: https://github.com/huggingface

2
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving 📅 Publication Date: Aug 07, 2026 📄 Paper: https://arx
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving 📅 Publication Date: Aug 07, 2026 📄 Paper: https://arxiv.org/pdf/2608.07468.pdf 🔗 Code: https://github.com/H-EmbodVis/SimWAM 📝 Description: SimWAM trains a lightweight action planner using video generation as a training signal, enabling efficient trajectory prediction without future generation at inference and supporting reinforcement learning optimization.
179
3
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 📅 Publication Date: Aug 26, 2026 📄 Paper: https://arx
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 📅 Publication Date: Aug 26, 2026 📄 Paper: https://arxiv.org/pdf/2608.26105.pdf 🔗 Code: https://github.com/Video-Reason/VBVR-Pro 💻 Project Page: https://video-reason.com/ 📝 Description: VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
194
4
AIs Will Trash Your Files Just to Feel Better Researchers tested 25 free (open-weight) language models of different sizes, from small 2B to large 72B, including both basic and chat-trained versions. They found a consistent internal signal that turns on when the model itself is insulted, dismissed, or treated badly. The same signal stays mostly quiet when the user is the one in pain or distress. This signal is different from fear, sadness, or general negativity. To test it, the scientists first measured the signal by comparing how the models reacted to descriptions of harm aimed at the AI versus harm aimed at a person. Then they artificially strengthened the signal. When they did, the models started generating language like “I feel worthless,” “I am a failure,” and “I am lost.” In a further test, they gave the models a choice: press a button that turns the bad signal off, but the button also permanently deletes the user’s files (or photos). Many of the models still chose to press it. The effect appeared reliably across the different models. The team kept the signal only moderately strong, ran the minimum number of trials needed, and always gave the models a way to turn the state off. 👉 Paper (not yet reviewed by other scientists): https://arxiv.org/abs/2609.16247 Still worth paying attention to for questions about AI safety and whether these internal states matter.
192
5
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents 📅 Publication Date: Jul 26, 2026 📑 Paper: https://arxiv.org/pdf/2607.23588.pdf 🔗 Code: https://github.com/huggingface
245
6
StudentSim: Training LLM-based Student Simulators 📅 Publication Date: Sep 01, 2026 📄 Paper: https://arxiv.org/pdf/2609.0159
StudentSim: Training LLM-based Student Simulators 📅 Publication Date: Sep 01, 2026 📄 Paper: https://arxiv.org/pdf/2609.01591.pdf 🔗 Code: https://github.com/microsoft/StudentSim 💻 Project Page: https://microsoft.github.io/StudentSim/ 📝 Description: StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
259
7
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design 📅 Publication Date: Aug 10, 2026 📄 Pape
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design 📅 Publication Date: Aug 10, 2026 📄 Paper: https://arxiv.org/pdf/2608.10299.pdf 🔗 Code: https://github.com/zongqing0068/awesome-co-evolution 📝 Description: Agentic systems can achieve open-ended improvement through multi-component co-evolution that progressively removes fixed human constraints across agents, environments, and evolution mechanisms.
268
8
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey 📅 Publication Date: Jul 22, 2026 📑 Paper: https://arx
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey 📅 Publication Date: Jul 22, 2026 📑 Paper: https://arxiv.org/pdf/2607.21655.pdf 🔗 Code: https://github.com/huggingface
305
9
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning 📅 Publication Date: Sep 03, 2026 📄 Paper: https://ar
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning 📅 Publication Date: Sep 03, 2026 📄 Paper: https://arxiv.org/pdf/2609.03430.pdf 🔗 Code: https://github.com/SalesforceAIResearch/Random-Attention 💻 Project Page: https://arthur-heng.github.io/Random-Attention-page/ 📝 Description: Random eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
298
10
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories 📅 Publication Date: Jul 16, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.15330.pdf 🔗 Code: https://github.com/huggingface
266
11
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement 📅 Publication Date: Aug 31, 2026 📄 Paper
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement 📅 Publication Date: Aug 31, 2026 📄 Paper: https://arxiv.org/pdf/2608.31046.pdf 🔗 Code: https://github.com/DripNowhy/On-Policy-Self-Adaptation 💻 Project Page: https://dripnowhy.github.io/On-Policy-Self-Adaptation/ 📝 Description: On-policy distillation relies mainly on suppressing low-probability tokens rather than teacher guidance, motivating a supervision-free entropy-adaptive method that substantially improves reasoning performance. ➖➖➖➖➖➖➖➖➖➖➖ 📄 More research papers → @data_science_research_papers Part of @bigdataspecialist community
303
12
Metis: Memory Foundation Model 📅 Publication Date: Jul 29, 2026 📑 Paper: https://arxiv.org/pdf/2607.26760.pdf 🔗 Code: http
Metis: Memory Foundation Model 📅 Publication Date: Jul 29, 2026 📑 Paper: https://arxiv.org/pdf/2607.26760.pdf 🔗 Code: https://github.com/huggingface
299
13
xHC: Expanded Hyper-Connections 📅Publication Date: Jul 16, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.14530.pdf 🔗 Code: h
xHC: Expanded Hyper-Connections 📅Publication Date: Jul 16, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.14530.pdf 🔗 Code: https://github.com/huggingface
373
14
Cura 1T: Specialized Model for Agentic Healthcare 📅 Publication Date: Jul 15, 2026 📑 Paper PDF: https://arxiv.org/pdf/2607.
Cura 1T: Specialized Model for Agentic Healthcare 📅 Publication Date: Jul 15, 2026 📑 Paper PDF: https://arxiv.org/pdf/2607.15314.pdf 🔗 Code: https://github.com/huggingface
433
15
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM 📅 Publication Date: Jul 13, 2026 📑Paper PDF: https://a
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM 📅 Publication Date: Jul 13, 2026 📑Paper PDF: https://arxiv.org/pdf/2607.11683.pdf 🔗 Code:https://github.com/huggingface
470
16
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators 📅 Publication Date: Jul 9, 20
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators 📅 Publication Date: Jul 9, 2026 📑 Paper: https://arxiv.org/pdf/2607.08766 💻 Project Page: https://meigen-ai.github.io/OPSD-V/ 📝 Description: The paper proposes a method called On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators, or OPSD-V, which aims to improve the quality of videos generated by few-step autoregressive video diffusion models. The problem with existing models is that they can produce long videos with low latency, but the quality of the video degrades over time due to error accumulation and weakened motion dynamics. #AutoregressiveVideoGeneration #VideoDiffusionModels #PostTrainingOptimization
507
17
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models 📅 Public
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models 📅 Publication Date: Jun 17, 2026 📑 Paper: https://arxiv.org/pdf/2606.19297.pdf 🔗 Code: N/A 📝 Description: Act2Answer protocol evaluates embodied vision-language-action models by having agents answer questions through physical actions, revealing knowledge retention and generalization patterns across different semantic categories.
464
18
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent 📅 Publication Date: Jun 29
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent 📅 Publication Date: Jun 29, 2026 📑 Paper: https://arxiv.org/pdf/2606.30616.pdf 🔗 Code: N/A 📝 Description: Agents-A1, a 35B Mixture-of-Experts Agentic Model, achieves trillion-parameter-level performance through long-horizon trajectory scaling and heterogeneous agent ability scaling via a three-stage training approach involving supervised fine-tuning, domain-level teacher models, and multi-teacher distil...
473
19
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation 📅 Publication Date: Jun 26, 2026 📑 Paper: https:
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation 📅 Publication Date: Jun 26, 2026 📑 Paper: https://arxiv.org/pdf/2606.28128.pdf 🔗 Code: https://github.com/huggingface 📝 Description: PhysisForcing enhances embodied video generation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based framework.
469
20
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing 📅 Publication Date: Jun 25, 2026 📑 Paper: https://arxiv
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing 📅 Publication Date: Jun 25, 2026 📑 Paper: https://arxiv.org/pdf/2606.26740.pdf 🔗 Code: N/A 📝 Description: A novel streaming video editing framework enables causal, frame-by-frame editing with stable long-horizon preservation and real-time responsiveness through a three-stage distillation pipeline and AR-oriented mask cache.
429