Data science research papers
Відкрити в Telegram
Machine learning and data science research papers Key ML and AI papers with code and GitHub repos. Simple way to follow current research. Join 👉 https://rebrand.ly/bigdatachannels DMCA: @disclosure_bds Contact: @mldatascientist
Показати більше3 195
Підписники
+324 години
+317 днів
+8430 днів
Архів дописів
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
📅 Publication Date: Jul 29, 2026
📑 Paper: https://arxiv.org/pdf/2607.27205.pdf
🔗 Code: https://github.com/huggingface
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
📅 Publication Date: Aug 07, 2026
📄 Paper: https://arxiv.org/pdf/2608.07468.pdf
🔗 Code: https://github.com/H-EmbodVis/SimWAM
📝 Description:
SimWAM trains a lightweight action planner using video generation as a training signal, enabling efficient trajectory prediction without future generation at inference and supporting reinforcement learning optimization.
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
📅 Publication Date: Aug 26, 2026
📄 Paper: https://arxiv.org/pdf/2608.26105.pdf
🔗 Code: https://github.com/Video-Reason/VBVR-Pro
💻 Project Page: https://video-reason.com/
📝 Description:
VBBR-Pro introduces a closed-loop testbed that enables scalable, verifiable, and controllable native visual reasoning through generation across diverse visual substrates.
➖➖➖➖➖➖➖➖➖➖➖
📄 More research papers → @data_science_research_papers
Part of @bigdataspecialist community
Repost from Talks with ChatGPT
AIs Will Trash Your Files Just to Feel Better
Researchers tested 25 free (open-weight) language models of different sizes, from small 2B to large 72B, including both basic and chat-trained versions.
They found a consistent internal signal that turns on when the model itself is insulted, dismissed, or treated badly. The same signal stays mostly quiet when the user is the one in pain or distress. This signal is different from fear, sadness, or general negativity.
To test it, the scientists first measured the signal by comparing how the models reacted to descriptions of harm aimed at the AI versus harm aimed at a person. Then they artificially strengthened the signal. When they did, the models started generating language like
“I feel worthless,” “I am a failure,” and “I am lost.”In a further test, they gave the models a choice: press a button that turns the bad signal off, but the button also permanently deletes the user’s files (or photos). Many of the models still chose to press it. The effect appeared reliably across the different models. The team kept the signal only moderately strong, ran the minimum number of trials needed, and always gave the models a way to turn the state off. 👉 Paper (not yet reviewed by other scientists): https://arxiv.org/abs/2609.16247 Still worth paying attention to for questions about AI safety and whether these internal states matter.
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
📅 Publication Date: Jul 26, 2026
📑 Paper: https://arxiv.org/pdf/2607.23588.pdf
🔗 Code: https://github.com/huggingface
StudentSim: Training LLM-based Student Simulators
📅 Publication Date: Sep 01, 2026
📄 Paper: https://arxiv.org/pdf/2609.01591.pdf
🔗 Code: https://github.com/microsoft/StudentSim
💻 Project Page: https://microsoft.github.io/StudentSim/
📝 Description:
StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math.
➖➖➖➖➖➖➖➖➖➖➖
📄 More research papers → @data_science_research_papers
Part of @bigdataspecialist community
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
📅 Publication Date: Aug 10, 2026
📄 Paper: https://arxiv.org/pdf/2608.10299.pdf
🔗 Code: https://github.com/zongqing0068/awesome-co-evolution
📝 Description:
Agentic systems can achieve open-ended improvement through multi-component co-evolution that progressively removes fixed human constraints across agents, environments, and evolution mechanisms.
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
📅 Publication Date: Jul 22, 2026
📑 Paper: https://arxiv.org/pdf/2607.21655.pdf
🔗 Code: https://github.com/huggingface
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
📅 Publication Date: Sep 03, 2026
📄 Paper: https://arxiv.org/pdf/2609.03430.pdf
🔗 Code: https://github.com/SalesforceAIResearch/Random-Attention
💻 Project Page: https://arthur-heng.github.io/Random-Attention-page/
📝 Description:
Random eviction of reasoning tokens matches selective KV cache compression because reasoning traces are self-protecting through redundancy, making scoring unnecessary once prompts are preserved.
➖➖➖➖➖➖➖➖➖➖➖
📄 More research papers → @data_science_research_papers
Part of @bigdataspecialist community
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
📅 Publication Date: Jul 16, 2026
📑Paper PDF: https://arxiv.org/pdf/2607.15330.pdf
🔗 Code: https://github.com/huggingface
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
📅 Publication Date: Aug 31, 2026
📄 Paper: https://arxiv.org/pdf/2608.31046.pdf
🔗 Code: https://github.com/DripNowhy/On-Policy-Self-Adaptation
💻 Project Page: https://dripnowhy.github.io/On-Policy-Self-Adaptation/
📝 Description:
On-policy distillation relies mainly on suppressing low-probability tokens rather than teacher guidance, motivating a supervision-free entropy-adaptive method that substantially improves reasoning performance.
➖➖➖➖➖➖➖➖➖➖➖
📄 More research papers → @data_science_research_papers
Part of @bigdataspecialist community
Metis: Memory Foundation Model
📅 Publication Date: Jul 29, 2026
📑 Paper: https://arxiv.org/pdf/2607.26760.pdf
🔗 Code: https://github.com/huggingface
xHC: Expanded Hyper-Connections
📅Publication Date: Jul 16, 2026
📑Paper PDF: https://arxiv.org/pdf/2607.14530.pdf
🔗 Code: https://github.com/huggingface
Cura 1T: Specialized Model for Agentic Healthcare
📅 Publication Date: Jul 15, 2026
📑 Paper PDF: https://arxiv.org/pdf/2607.15314.pdf
🔗 Code: https://github.com/huggingface
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
📅 Publication Date: Jul 13, 2026
📑Paper PDF: https://arxiv.org/pdf/2607.11683.pdf
🔗 Code:https://github.com/huggingface
OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
📅 Publication Date: Jul 9, 2026
📑 Paper: https://arxiv.org/pdf/2607.08766
💻 Project Page: https://meigen-ai.github.io/OPSD-V/
📝 Description:
The paper proposes a method called On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators, or OPSD-V, which aims to improve the quality of videos generated by few-step autoregressive video diffusion models. The problem with existing models is that they can produce long videos with low latency, but the quality of the video degrades over time due to error accumulation and weakened motion dynamics.
#AutoregressiveVideoGeneration #VideoDiffusionModels #PostTrainingOptimization
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
📅 Publication Date: Jun 17, 2026
📑 Paper: https://arxiv.org/pdf/2606.19297.pdf
🔗 Code: N/A
📝 Description:
Act2Answer protocol evaluates embodied vision-language-action models by having agents answer questions through physical actions, revealing knowledge retention and generalization patterns across different semantic categories.
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
📅 Publication Date: Jun 29, 2026
📑 Paper: https://arxiv.org/pdf/2606.30616.pdf
🔗 Code: N/A
📝 Description:
Agents-A1, a 35B Mixture-of-Experts Agentic Model, achieves trillion-parameter-level performance through long-horizon trajectory scaling and heterogeneous agent ability scaling via a three-stage training approach involving supervised fine-tuning, domain-level teacher models, and multi-teacher distil...
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
📅 Publication Date: Jun 26, 2026
📑 Paper: https://arxiv.org/pdf/2606.28128.pdf
🔗 Code: https://github.com/huggingface
📝 Description:
PhysisForcing enhances embodied video generation by enforcing physical consistency through pixel-level trajectory alignment and semantic-level relational alignment losses in a DiT-based framework.
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
📅 Publication Date: Jun 25, 2026
📑 Paper: https://arxiv.org/pdf/2606.26740.pdf
🔗 Code: N/A
📝 Description:
A novel streaming video editing framework enables causal, frame-by-frame editing with stable long-horizon preservation and real-time responsiveness through a three-stage distillation pipeline and AR-oriented mask cache.
