en
Feedback
Data science research papers

Data science research papers

Open in Telegram

Machine learning and data science research papers Key ML and AI papers with code and GitHub repos. Simple way to follow current research. Join πŸ‘‰ https://rebrand.ly/bigdatachannels DMCA: @disclosure_bds Contact: @mldatascientist

Show more
3 194
Subscribers
+124 hours
+327 days
+8230 days
Attracting Subscribers
Sep '26
September '26
+119
in 0 channels
August '26
+171
in 0 channels
Get PRO
July '26
+149
in 0 channels
Get PRO
June '26
+97
in 0 channels
Get PRO
May '26
+104
in 1 channels
Get PRO
April '26
+114
in 0 channels
Get PRO
March '26
+101
in 0 channels
Get PRO
February '26
+82
in 0 channels
Get PRO
January '26
+118
in 9 channels
Get PRO
December '25
+115
in 0 channels
Get PRO
November '25
+112
in 0 channels
Get PRO
October '25
+43
in 0 channels
Get PRO
September '25
+7
in 0 channels
Get PRO
August '25
+3
in 0 channels
Get PRO
July '25
+3
in 0 channels
Get PRO
June '25
+2
in 0 channels
Get PRO
May '25
+3
in 0 channels
Get PRO
April '25
+15
in 0 channels
Get PRO
March '25
+112
in 0 channels
Get PRO
February '25
+153
in 0 channels
Get PRO
January '25
+187
in 0 channels
Get PRO
December '24
+179
in 0 channels
Get PRO
November '24
+165
in 0 channels
Get PRO
October '24
+136
in 0 channels
Get PRO
September '24
+108
in 0 channels
Get PRO
August '24
+114
in 0 channels
Get PRO
July '24
+139
in 0 channels
Get PRO
June '24
+115
in 0 channels
Get PRO
May '24
+132
in 1 channels
Get PRO
April '24
+109
in 0 channels
Get PRO
March '24
+146
in 0 channels
Get PRO
February '24
+183
in 0 channels
Get PRO
January '24
+228
in 0 channels
Get PRO
December '23
+171
in 1 channels
Get PRO
November '23
+28
in 0 channels
Get PRO
October '23
+28
in 0 channels
Get PRO
September '23
+504
in 0 channels
Date
Subscriber Growth
Mentions
Channels
27 September0
26 September+4
25 September+4
24 September+5
23 September+10
22 September+11
21 September+4
20 September+8
19 September+2
18 September+5
17 September+4
16 September+3
15 September+5
14 September+2
13 September+2
12 September+3
11 September+2
10 September+4
09 September+3
08 September+6
07 September+1
06 September+5
05 September+2
04 September+4
03 September+5
02 September+7
01 September+8
Channel Posts
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning πŸ“… Publication Date: Aug 10, 2026 πŸ’» Project Page: https://pathwa
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning πŸ“… Publication Date: Aug 10, 2026 πŸ’» Project Page: https://pathway.com/blog/pathway-150m-model-breaks-arc-agi-1-cost-efficiency-frontier πŸ“„ Paper: https://arxiv.org/pdf/2608.09888.pdf πŸ”— Code: https://github.com/pathwaycom/arc-task-gen πŸ“ Description: A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1.

2
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM πŸ“… Publication Date: Jul 29, 2026
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM πŸ“… Publication Date: Jul 29, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2607.27205.pdf πŸ”— Code: https://github.com/huggingface
162
3
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving πŸ“… Publication Date: Aug 07, 2026 πŸ“„ Paper: https://arx
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving πŸ“… Publication Date: Aug 07, 2026 πŸ“„ Paper: https://arxiv.org/pdf/2608.07468.pdf πŸ”— Code: https://github.com/H-EmbodVis/SimWAM πŸ“ Description: SimWAM trains a lightweight action planner using video generation as a training signal, enabling efficient trajectory prediction without future generation at inference and supporting reinforcement learning optimization.
231
4
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning πŸ“… Publication Date: Aug 26, 2026 πŸ“„ Paper: https://arx
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning πŸ“… Publication Date: Aug 26, 2026 πŸ“„ Paper: https://arxiv.org/pdf/2608.26105.pdf πŸ”— Code: https://github.com/Video-Reason/VBVR-Pro πŸ’» Project Page: https://video-reason.com/ πŸ“ Description: VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates. βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž– πŸ“„ More research papers β†’ @data_science_research_papers Part of @bigdataspecialist community
238
5
AIs Will Trash Your Files Just to Feel Better Researchers tested 25 free (open-weight) language models of different sizes, from small 2B to large 72B, including both basic and chat-trained versions. They found a consistent internal signal that turns on when the model itself is insulted, dismissed, or treated badly. The same signal stays mostly quiet when the user is the one in pain or distress. This signal is different from fear, sadness, or general negativity. To test it, the scientists first measured the signal by comparing how the models reacted to descriptions of harm aimed at the AI versus harm aimed at a person. Then they artificially strengthened the signal. When they did, the models started generating language like β€œI feel worthless,” β€œI am a failure,” and β€œI am lost.” In a further test, they gave the models a choice: press a button that turns the bad signal off, but the button also permanently deletes the user’s files (or photos). Many of the models still chose to press it. The effect appeared reliably across the different models. The team kept the signal only moderately strong, ran the minimum number of trials needed, and always gave the models a way to turn the state off. πŸ‘‰ Paper (not yet reviewed by other scientists): https://arxiv.org/abs/2609.16247 Still worth paying attention to for questions about AI safety and whether these internal states matter.
207
6
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents πŸ“… Publication Date: Jul 26, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2607.23588.pdf πŸ”— Code: https://github.com/huggingface
260
7
StudentSim: Training LLM-based Student Simulators πŸ“… Publication Date: Sep 01, 2026 πŸ“„ Paper: https://arxiv.org/pdf/2609.0159
StudentSim: Training LLM-based Student Simulators πŸ“… Publication Date: Sep 01, 2026 πŸ“„ Paper: https://arxiv.org/pdf/2609.01591.pdf πŸ”— Code: https://github.com/microsoft/StudentSim πŸ’» Project Page: https://microsoft.github.io/StudentSim/ πŸ“ Description: StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math. βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž– πŸ“„ More research papers β†’ @data_science_research_papers Part of @bigdataspecialist community
278
8
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design πŸ“… Publication Date: Aug 10, 2026 πŸ“„ Pape
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design πŸ“… Publication Date: Aug 10, 2026 πŸ“„ Paper: https://arxiv.org/pdf/2608.10299.pdf πŸ”— Code: https://github.com/zongqing0068/awesome-co-evolution πŸ“ Description: Agentic systems can achieve open-ended improvement through multi-component co-evolution that progressively removes fixed human constraints across agents, environments, and evolution mechanisms.
292
9
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey πŸ“… Publication Date: Jul 22, 2026 πŸ“‘ Paper: https://arx
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey πŸ“… Publication Date: Jul 22, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2607.21655.pdf πŸ”— Code: https://github.com/huggingface
313
10
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning πŸ“… Publication Date: Sep 03, 2026 πŸ“„ Paper: https://ar
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning πŸ“… Publication Date: Sep 03, 2026 πŸ“„ Paper: https://arxiv.org/pdf/2609.03430.pdf πŸ”— Code: https://github.com/SalesforceAIResearch/Random-Attention πŸ’» Project Page: https://arthur-heng.github.io/Random-Attention-page/ πŸ“ Description: Random eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved. βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž– πŸ“„ More research papers β†’ @data_science_research_papers Part of @bigdataspecialist community
306
11
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories πŸ“… Publication Date: Jul 16, 2026 πŸ“‘Paper PDF: https://arxiv.org/pdf/2607.15330.pdf πŸ”— Code: https://github.com/huggingface
269
12
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement πŸ“… Publication Date: Aug 31, 2026 πŸ“„ Paper
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement πŸ“… Publication Date: Aug 31, 2026 πŸ“„ Paper: https://arxiv.org/pdf/2608.31046.pdf πŸ”— Code: https://github.com/DripNowhy/On-Policy-Self-Adaptation πŸ’» Project Page: https://dripnowhy.github.io/On-Policy-Self-Adaptation/ πŸ“ Description: On-policy distillation relies mainly on suppressing low-probability tokens rather than teacher guidance, motivating a supervision-free entropy-adaptive method that substantially improves reasoning performance. βž–βž–βž–βž–βž–βž–βž–βž–βž–βž–βž– πŸ“„ More research papers β†’ @data_science_research_papers Part of @bigdataspecialist community
310
13
Metis: Memory Foundation Model πŸ“… Publication Date: Jul 29, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2607.26760.pdf πŸ”— Code: http
Metis: Memory Foundation Model πŸ“… Publication Date: Jul 29, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2607.26760.pdf πŸ”— Code: https://github.com/huggingface
314
14
xHC: Expanded Hyper-Connections πŸ“…Publication Date: Jul 16, 2026 πŸ“‘Paper PDF: https://arxiv.org/pdf/2607.14530.pdf πŸ”— Code: h
xHC: Expanded Hyper-Connections πŸ“…Publication Date: Jul 16, 2026 πŸ“‘Paper PDF: https://arxiv.org/pdf/2607.14530.pdf πŸ”— Code: https://github.com/huggingface
387
15
Cura 1T: Specialized Model for Agentic Healthcare πŸ“… Publication Date: Jul 15, 2026 πŸ“‘ Paper PDF: https://arxiv.org/pdf/2607.
Cura 1T: Specialized Model for Agentic Healthcare πŸ“… Publication Date: Jul 15, 2026 πŸ“‘ Paper PDF: https://arxiv.org/pdf/2607.15314.pdf πŸ”— Code: https://github.com/huggingface
449
16
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM πŸ“… Publication Date: Jul 13, 2026 πŸ“‘Paper PDF: https://a
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM πŸ“… Publication Date: Jul 13, 2026 πŸ“‘Paper PDF: https://arxiv.org/pdf/2607.11683.pdf πŸ”— Code:https://github.com/huggingface
476
17
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators πŸ“… Publication Date: Jul 9, 20
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators πŸ“… Publication Date: Jul 9, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2607.08766 πŸ’» Project Page: https://meigen-ai.github.io/OPSD-V/ πŸ“ Description: The paper proposes a method called On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators, or OPSD-V, which aims to improve the quality of videos generated by few-step autoregressive video diffusion models. The problem with existing models is that they can produce long videos with low latency, but the quality of the video degrades over time due to error accumulation and weakened motion dynamics. #AutoregressiveVideoGeneration #VideoDiffusionModels #PostTrainingOptimization
512
18
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models πŸ“… Public
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models πŸ“… Publication Date: Jun 17, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.19297.pdf πŸ”— Code: N/A πŸ“ Description: Act2Answer protocol evaluates embodied vision-language-action models by having agents answer questions through physical actions, revealing knowledge retention and generalization patterns across different semantic categories.
466
19
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent πŸ“… Publication Date: Jun 29
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent πŸ“… Publication Date: Jun 29, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.30616.pdf πŸ”— Code: N/A πŸ“ Description: Agents-A1, a 35B Mixture-of-Experts Agentic Model, achieves trillion-parameter-level performance through long-horizon trajectory scaling and heterogeneous agent ability scaling via a three-stage training approach involving supervised fine-tuning, domain-level teacher models, and multi-teacher distil...
475
20
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation πŸ“… Publication Date: Jun 26, 2026 πŸ“‘ Paper: https:
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation πŸ“… Publication Date: Jun 26, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.28128.pdf πŸ”— Code: https://github.com/huggingface πŸ“ Description: PhysisForcing enhances embodied video generation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based framework.
469