en
Feedback
Data science research papers

Data science research papers

Open in Telegram

Machine learning and data science research papers Key ML and AI papers with code and GitHub repos. Simple way to follow current research. Join πŸ‘‰ https://rebrand.ly/bigdatachannels DMCA: @disclosure_bds Contact: @mldatascientist

Show more
3 012
Subscribers
+324 hours
+177 days
+7930 days
Posts Archive
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions πŸ“… Publication Date: Jun 22, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.23654.pdf πŸ”— Code: https://github.com/huggingface πŸ“ Description: EnterpriseClawBench presents a benchmark for enterprise agents based on real-world sessions with 852 reproducible tasks, emphasizing comprehensive evaluation metrics beyond single performance scores.

Heterogeneous Scientific Foundation Model Collaboration πŸ“… Publication Date: Apr 30, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/260
Heterogeneous Scientific Foundation Model Collaboration πŸ“… Publication Date: Apr 30, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2604.27351.pdf πŸ”— Code: https://github.com/Violet24K/Eywa πŸ“ Description: Eywa is a heterogeneous agentic framework that extends language-centric systems to scientific foundation models by integrating domain-specific models with language-based reasoning interfaces for improved performance across diverse scientific domains.

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents πŸ“… Publication Date: May 6, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.05185.pdf πŸ”— Code: https://github.com/shawn0728/OpenSearch-VL πŸ“ Description: OpenSearch-VL presents an open-source framework for training advanced multimodal search agents using reinforcement learning, featuring specialized data curation, diverse tool environments, and a novel training algorithm that improves performance across multiple benchmarks.

WorldOlympiad: Can Your World Model Survive a Triathlon? πŸ“… Publication Date: Jun 9, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/260
WorldOlympiad: Can Your World Model Survive a Triathlon? πŸ“… Publication Date: Jun 9, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.11129 πŸ’» Project Page: https://alibaba-damo-academy.github.io/WorldOlympiad/ πŸ“ Description: The paper introduces WorldOlympiad, a comprehensive benchmark for evaluating video-based world models. The problem with current generative models is that they often focus on visual quality, but lack physical faithfulness, geometric consistency, and interaction fidelity. To address this gap, WorldOlympiad decomposes world-model evaluation into three dimensions: physical faithfulness, geometric consistency, and interaction fidelity. WorldOlympiad covers three major downstream scenarios, including gaming, robotics, and general real-world videos, capturing diverse challenges from interactive control and embodied manipulation to open-domain motion and camera dynamics. #WorldModelEvaluation #VideoBasedWorldModels #PhysicalFaithfulness #GeometricConsistency

Qwen-AgentWorld: Language World Models for General Agents πŸ“… Publication Date: Jun 23, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.24597.pdf πŸ”— Code: https://github.com/huggingface πŸ“ Description: Language-based world models enable agentic environment simulation across multiple domains and enhance general agent performance through scalable simulation and improved downstream task performance.

HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation πŸ“… Publication Date: Apr 30, 2026 πŸ“‘
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation πŸ“… Publication Date: Apr 30, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2604.28196.pdf πŸ”— Code: https://github.com/H-EmbodVis/HERMESV2 πŸ“ Description: HERMES++ combines 3D scene understanding and future geometry prediction through BEV representation, LLM-enhanced queries, temporal linking, and joint geometric optimization for autonomous driving applications.

ABot-Earth 0.5: Generative 3D Earth Model πŸ“… Publication Date: Jun 8, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.09967 πŸ’» Proj
ABot-Earth 0.5: Generative 3D Earth Model πŸ“… Publication Date: Jun 8, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.09967 πŸ’» Project Page: https://abot-earth.amap.com/ πŸ“ Description: The paper presents ABot-Earth 0.5, a generative framework that creates realistic 3D environments from satellite imagery. The problem addressed is the need for large-scale 3D reconstruction, which is currently expensive and technically challenging. The authors propose a novel generative model based on 3D Gaussian Splatting representation, which is trained on a diverse set of real-world urban reconstructions. This model learns to generate realistic geometry and textures, and can synthesize novel 3D scenes conditioned solely on satellite imagery in under 10 minutes per square kilometer. #Generative3DModeling #3DGaussianSplatting #SatelliteImageryReconstruction #GeospatialModeling

OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories πŸ“… Publication Date: May 5, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.04036.pdf πŸ”— Code: https://github.com/PolarSeeker/OpenSeeker πŸ“ Description: A simple supervised fine-tuning approach achieves state-of-the-art performance in deep search capabilities using minimal data, outperforming complex industrial pipelines and demonstrating the effectiveness of academic-led development in large language model agents.

World Model for Robot Learning: A Comprehensive Survey πŸ“… Publication Date: Apr 30, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605
World Model for Robot Learning: A Comprehensive Survey πŸ“… Publication Date: Apr 30, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.00080 πŸ’» Project Page: https://ntumars.github.io/wm-robot-survey/ πŸ”— Code: https://github.com/NTUMARS/Awesome-World-Model-for-Robotics-Policy πŸ“ Description: The paper provides a comprehensive survey of world models for robot learning, which are predictive representations of environmental dynamics that support policy learning, planning, and simulation. The authors note that the literature on world models is fragmented across different architectures, functional roles, and application domains, making it difficult to understand the current state of the field. To address this gap, the authors present a systematic review of world models from a robot learning perspective, examining how they are coupled with robot policies, used as learned simulators for reinforcement learning and evaluation, and have progressed in terms of robotic video world models. #RobotLearning #RobotPolicies

Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond πŸ“… Publication Date: Apr 24, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2604.22748.pdf πŸ”— Code: https://github.com/matrix-agent/awesome-agentic-world-modeling πŸ“ Description: World models are categorized into three capability levels and four law regimes to better understand and develop predictive environment models for AI agents across diverse domains.

Asymmetric Flow Models πŸ“… Publication Date: May 13, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.12964 πŸ’» Project Page: https://
Asymmetric Flow Models πŸ“… Publication Date: May 13, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.12964 πŸ’» Project Page: https://hanshengchen.com/asymflow/ πŸ”— Code: https://github.com/Lakonik/LakonLab ⭐️ 324 πŸ“Š Models citing this paper: β€’ https://huggingface.co/Lakonik/AsymFLUX.2-klein-9B β€’ https://huggingface.co/Lakonik/AsymFlow-ImageNet β€’ https://huggingface.co/OJ-1/AsymFLUX.2-klein-9B πŸ“ Description: The paper introduces Asymmetric Flow Modeling, a method for efficient high-dimensional flow-based generation. The problem with existing flow-based generation methods is that they require modeling high-dimensional noise, which is difficult even when the data has a strong low-rank structure. To address this, the authors propose a rank-asymmetric velocity parameterization that restricts noise prediction to a low-rank subspace while keeping data prediction full-dimensional. #AsymmetricFlowModels #FlowBasedGeneration #RankAsymmetricVelocity

From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company πŸ“… Publication Date: Apr 24, 2026 πŸ“‘ Paper: ht
From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company πŸ“… Publication Date: Apr 24, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2604.22446.pdf πŸ”— Code: https://github.com/1mancompany/OneManCompany πŸ“ Description: OneManCompany (OMC) introduces an organizational framework for multi-agent systems that enables dynamic team assembly, governance, and improvement through portable agent identities and hierarchical decision-making processes.

πŸ“’ Advertising in this channel You can place an ad via Telegaβ€€io. It takes just a few minutes. Formats and current rates: Vie
πŸ“’ Advertising in this channel You can place an ad via Telegaβ€€io. It takes just a few minutes. Formats and current rates: View details

πŸ”₯ Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning πŸ“… Publication Date: Jun 9, 2026 πŸ“‘ Paper: https://
πŸ”₯ Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning πŸ“… Publication Date: Jun 9, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2606.11087 πŸ’»Project Page: https://q-guided-flow.github.io/ πŸ“ Description: The paper proposes a reinforcement learning algorithm called QGF that improves policies at test time by using a value gradient to guide a pre-trained flow policy. The problem addressed is that incorporating flow models into reinforcement learning pipelines for policy improvement can be difficult due to stability and scalability issues. The method involves pre-training a reference flow policy and a value function critic, then using the value gradient to guide the reference policy to generate higher-value actions at test time, without any additional policy learning. #ReinforcementLearningAlgorithms #TestTimePolicyImprovement #QGFAlgorithm

LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling πŸ“… Publication Date: May 8, 2026 πŸ“‘ Paper: https://arxiv.org/pdf
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling πŸ“… Publication Date: May 8, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.08083 πŸ’» Project Page: https://zhengkid.github.io/AutoTTS-web/ πŸ”— Code: https://github.com/zhengkid/AutoTTS πŸ“ Description: The paper proposes a novel approach to improve the performance of large language models through test-time scaling, which involves allocating additional computation during inference. Existing test-time scaling strategies are typically hand-crafted, relying on manual design and tuning of reasoning patterns and heuristics. This approach leaves much of the computation-allocation space unexplored, resulting in potential inefficiencies. To address this limitation, the authors introduce AutoTTS, an environment-driven framework that automates the discovery of test-time scaling strategies. Instead of designing individual strategies, researchers can create environments where optimal strategies can be discovered automatically. #LargeLanguageModels

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation πŸ“… Publication Date: Apr 27, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2604.24764.pdf πŸ”— Code: https://github.com/microsoft/World-R1 πŸ“ Description: World-R1 framework improves video generation by incorporating 3D constraints through reinforcement learning and specialized text datasets while maintaining visual quality and scalability.

Code as Agent Harness πŸ“… Publication Date: May 18, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.18747 πŸ”— Code: N/A πŸ“ Descriptio
Code as Agent Harness πŸ“… Publication Date: May 18, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.18747 πŸ”— Code: N/A πŸ“ Description: The paper discusses the concept of code as agent harness, where large language models are used as operational substrates for agent reasoning and execution in agentic systems. The authors argue that code is no longer just a target output, but serves as a unified infrastructure layer across multiple domains and applications. They introduce a unified view that centers code as the basis for agent infrastructure, and organize their survey around three connected layers: the harness interface, harness mechanisms, and scaling the harness. #AgenticSystems #LargeLanguageModels #AgentReasoning

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration πŸ“… Publication Date: May 19, 2026 πŸ“‘ Paper
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration πŸ“… Publication Date: May 19, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2605.20025 πŸ’» Project Page: https://github.com/aiming-lab/AutoResearchClaw πŸ”— Code: https://github.com/huggingface πŸ—ƒDatasets citing this paper: β€’ https://huggingface.co/datasets/AIMING-Lab-UNC/ARC-Bench πŸ“Description: AutoResearchClaw is a new autonomous research system that improves scientific discovery by incorporating human collaboration and iterative learning. The problem with existing autonomous research systems is that they often model the research process as a linear pipeline, relying on single agent reasoning and stopping when execution fails, without carrying experience across runs. #AutonomousResearchSystems #MultiAgentLearning #SelfReinforcingSystems

AgentSearchBench: A Benchmark for AI Agent Search in the Wild πŸ“… Publication Date: Apr 24, 2026 πŸ“‘ Paper: https://arxiv.org/p
AgentSearchBench: A Benchmark for AI Agent Search in the Wild πŸ“… Publication Date: Apr 24, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2604.22436 πŸ”— Code: N/A πŸ“ Description: AgentSearchBench is a new benchmark for finding suitable AI agents using execution-grounded performance signals from nearly 10,000 real-world agents. It shows that description-based similarity is insufficient, and lightweight behavioral signals significantly improve agent ranking. #AI #AIAgents #Benchmarking #AgentSearch #MachineLearning

Omnilingual MT: Machine Translation for 1,600 Languages πŸ“… Publication Date: Mar 17, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/260
Omnilingual MT: Machine Translation for 1,600 Languages πŸ“… Publication Date: Mar 17, 2026 πŸ“‘ Paper: https://arxiv.org/pdf/2603.16309 πŸ—ƒ Datasets citing this paper: https://huggingface.co/datasets/facebook/bouquet πŸ”— Code: N/A πŸ“ Description: Omnilingual MT OMT is the first system to support over 1,600 languages. It uses specialized smaller LLMs 1B-8B to outperform 70B baselines, achieving high-quality translation and coherent generation in low-compute settings. #AI #DataScience #MachineLearning #HuggingFace #Research