Data Scientology
Kanalga Telegramāda oātish
Hot data science related posts every hour. Chat: https://telegram.me/r_channels Contacts: @lgyanf
Ko'proq ko'rsatish1 147
Obunachilar
+124 soatlar
-37 kun
-630 kun
Ma'lumot yuklanmoqda...
O'xshash kanallar
Taglar buluti
Kirish va chiqish esdaliklari
---
---
---
---
---
---
Obunachilarni jalb qilish
Sentabr '26
Sentabr '26
+7
0 kanalda
Avgust '26
+11
0 kanalda
Get PRO
Iyul '26
+13
0 kanalda
Get PRO
Iyun '26
+6
0 kanalda
Get PRO
May '26
+4
0 kanalda
Get PRO
Aprel '26
+30
0 kanalda
Get PRO
Mart '26
+40
0 kanalda
Get PRO
Fevral '26
+36
0 kanalda
Get PRO
Yanvar '26
+13
0 kanalda
Get PRO
Dekabr '25
+10
0 kanalda
Get PRO
Noyabr '25
+7
0 kanalda
Get PRO
Oktabr '25
+4
0 kanalda
Get PRO
Sentabr '25
+3
0 kanalda
Get PRO
Avgust '25
+8
0 kanalda
Get PRO
Iyul '25
+10
0 kanalda
Get PRO
Iyun '25
+10
0 kanalda
Get PRO
May '25
+9
0 kanalda
Get PRO
Aprel '25
+3
0 kanalda
Get PRO
Mart '25
+7
0 kanalda
Get PRO
Fevral '25
+3
0 kanalda
Get PRO
Yanvar '25
+4
0 kanalda
Get PRO
Dekabr '24
+6
0 kanalda
Get PRO
Noyabr '24
+2
0 kanalda
Get PRO
Oktabr '24
+8
0 kanalda
Get PRO
Sentabr '24
+8
0 kanalda
Get PRO
Avgust '24
+11
0 kanalda
Get PRO
Iyul '24
+13
0 kanalda
Get PRO
Iyun '24
+11
0 kanalda
Get PRO
May '24
+10
0 kanalda
Get PRO
Aprel '24
+12
0 kanalda
Get PRO
Mart '24
+20
0 kanalda
Get PRO
Fevral '24
+19
0 kanalda
Get PRO
Yanvar '24
+19
0 kanalda
Get PRO
Dekabr '23
+17
6 kanalda
Get PRO
Noyabr '23
+23
3 kanalda
Get PRO
Oktabr '23
+27
3 kanalda
Get PRO
Sentabr '23
+21
0 kanalda
Get PRO
Avgust '23
+9
0 kanalda
Get PRO
Iyul '23
+18
0 kanalda
Get PRO
Iyun '23
+20
0 kanalda
Get PRO
May '23
+26
0 kanalda
Get PRO
Aprel '23
+53
0 kanalda
Get PRO
Mart '23
+15
0 kanalda
Get PRO
Fevral '23
+12
0 kanalda
Get PRO
Yanvar '23
+18
0 kanalda
Get PRO
Dekabr '22
+31
0 kanalda
Get PRO
Noyabr '22
+14
0 kanalda
Get PRO
Oktabr '22
+47
0 kanalda
Get PRO
Sentabr '22
+45
0 kanalda
Get PRO
Avgust '22
+18
0 kanalda
Get PRO
Iyul '22
+18
0 kanalda
Get PRO
Iyun '22
+18
0 kanalda
Get PRO
May '22
+27
0 kanalda
Get PRO
Aprel '22
+43
0 kanalda
Get PRO
Mart '22
+47
0 kanalda
Get PRO
Fevral '22
+89
0 kanalda
Get PRO
Yanvar '22
+57
0 kanalda
Get PRO
Dekabr '21
+36
0 kanalda
Get PRO
Noyabr '21
+27
0 kanalda
Get PRO
Oktabr '21
+43
0 kanalda
Get PRO
Sentabr '21
+64
0 kanalda
Get PRO
Avgust '21
+55
0 kanalda
Get PRO
Iyul '21
+28
0 kanalda
Get PRO
Iyun '21
+26
0 kanalda
Get PRO
May '21
+29
0 kanalda
Get PRO
Aprel '21
+26
0 kanalda
Get PRO
Mart '21
+49
0 kanalda
Get PRO
Fevral '21
+21
0 kanalda
Get PRO
Yanvar '21
+45
0 kanalda
Get PRO
Dekabr '20
+882
0 kanalda
| Sana | Obunachilarni jalb qilish | Esdaliklar | Kanallar | |
| 13 Sentabr | 0 | |||
| 12 Sentabr | +1 | |||
| 11 Sentabr | 0 | |||
| 10 Sentabr | 0 | |||
| 09 Sentabr | +2 | |||
| 08 Sentabr | 0 | |||
| 07 Sentabr | 0 | |||
| 06 Sentabr | 0 | |||
| 05 Sentabr | 0 | |||
| 04 Sentabr | +2 | |||
| 03 Sentabr | 0 | |||
| 02 Sentabr | 0 | |||
| 01 Sentabr | +2 |
Kanal postlari
I made a computer vision tool for running analysis!
https://redd.it/1wcx27z
@datascientology
| 2 | OpenAl Says It Has Cracked One of Math's āMillennium Problemsā (Navier-Stokes) N
As reported by the New York Times:
https://www.nytimes.com/2026/09/08/science/openai-proof-millennium-problem.html?smid=nytcore-ios-share
OpenAIās announcement:
https://openai.com/index/navier-stokes-solution/
https://redd.it/1wavdi7
@datascientology | 167 |
| 3 | Camera recommendation for real-time object tracking on a conveyor belt (Budget: ~$150 - $280)
https://redd.it/1w9nud3
@datascientology | 227 |
| 4 | World Labs' new Atlas model: Space-time simulation, "bullet time" from 3 cell phones, and scalable Real-to-Sim
https://redd.it/1w4m5yr
@datascientology | 333 |
| 5 | Where to submit stat/prob ML D
I'm a researcher in statistical and probabilistic ML, I have a steady record of top ML publications and really used to enjoy going to conferences.
Over the last few years LLM based works have completely taken over the top conferences. At this year's ICLR, walking among the rows of posters you were lucky to find one paper per row of 10 that wasn't about how their favourite LLM could or couldn't solve their niche benchmark. The workshops tell the same story, most are some kind of agentic flavour. Looking at this year's NeurIPS workshops it's the same thing, basically all are about agents.
I'm wondering where do the stat/prob ML communities go from here? I look up to people like Arnaud Doucet, Aapo HyvƤrinen, Christian Naesseth, Stefano Ermon, they seem to still publish at the top 3? On my end, I m thinking AISTATS/UAI might be the way to go.
All in all, the top 3 might never really have been intended as the home for prob/statML works, it just happened to be the 'prestigious' venue.
https://redd.it/1w0kipf
@datascientology | 431 |
| 6 | Qwen 3.6 VLM playing āWhereās Waldo?ā
https://redd.it/1vylz7a
@datascientology | 334 |
| 7 | I honestly did not think I will be able with on-device models
https://redd.it/1vwdapl
@datascientology | 330 |
| 8 | EMNLP 2026 Notifications
EMNLP 2026 notifications are expected in approximately 14 hours, so Iām creating this thread for everyone waiting for the results.
Good luck, everyone! Hopefully the next 14 hours pass quickly. š¤
https://redd.it/1vtxi3u
@datascientology | 257 |
| 9 | AI fatigue is killing motivation
I am about to start my MSc. I wish to specialize in computer vision, then pursue a PhD. I eventually want to work in industry. I was initially excited about this path. However, AI fatigue is killing my motivation.
Honestly, I don't have any hope for the future. It has been around four years since GPT-3.5 was introduced. AI is now proving major conjectures. It recently came close to proving Riemann's hypothesis, and dominated(not only defeated) the best competitive programmers in the world at AtCoder World Finals. I can't see a place for myself in the future because of AI.
I keep going because I feel like I don't have any other choice. I was genuinely excited about computer vision, robotics, and autonomous driving. But I have convinced myself that all my effort is in vain.
I wish to ask people in a similar situation, what makes you keep going? What are your plans for the future?
https://redd.it/1vqnl7w
@datascientology | 304 |
| 10 | Decoupled Descent: Enforcing Exact Train-Test Error Tracking Via AMP Onsager Corrections R
Link: https://arxiv.org/pdf/2604.27883
Hi,
Most of use are familiar with the headache of training a neural network using gradient descent where the training error may go to zero but the test error may stay the same as initialization or even increases.
My paper treats this phenomena as a consequence of data reuse bias and can be isolated by studying full batch gradient descent on a set of stylize Gaussian mixture models. I turns out that this fundamental issue can be avoided using some clever tricks from high-dimensional statistical theory, specifically approximate message passing (which is beyond the scope of this post but I would be happy to explain more).
By doing so I created a training method called Decoupled Descent (DD) which generates a certificate that the training error of the network will asymptotically equal the testing error at each parameter iterate. I think this method gives a cool way to approach how to train networks and I was hoping to get y'alls input on it. It opens up some nice ideas for optimal stopping or hyperparameter tuning and future directions of pushing to something like SGD or more general models.
I have attached the train-test curves on a simple model fitting problem to compare the performance of GD with with DD (my algorithm) to give a high-level idea of what the method can guarantee. I stress this is a theory paper so there is a long way to go to get to very large models but I think it is a good first step.
100 simulations of a simple high dimensional XOR model for a bespoke two layer network. Left is training with GD, right its training with my method. The colored bands are 25% to 75% quantile.
Happy to answer whatever questions people have, I plan on writing a PyTorch compatible package for this training method one day so any feature suggestions would be welcome as well.
https://redd.it/1vlu1se
@datascientology | 358 |
| 11 | Run SAM3 and RTMPose over 1950s-era factory footage. No fine-tuning. It just works
https://redd.it/1vhp0h6
@datascientology | 443 |
| 12 | I Compressed Bad Apple into a 3MB Neural Network [P]
https://redd.it/1vfrco1
@datascientology | 468 |
| 13 | Is it too late regain some coherence in the ML research space in our life time? D
Was just looking at the list of preprints on Arxiv cs.LG https://arxiv.org/list/cs.LG/recent?skip=0&show=500
Everyday 100 - 400 new machine learning papers gets uploaded on this server.
Looking at this unending list of preprints is as if you stepped into a crowded room, like the stock trading floor on wall st. in the 1980s. Everyone is shouting over each other. Nobody is talking to each other. Everyone's trying to prove something, to someone, to themselves, to build some credentials in the ML/AI space to meet those job requirements, or dying to get their truth out. Every title contains some new terminology invented by the authors that feels not worth the effort in keeping it in your working memory. Burn-out by endless novelty.
Frontier research are now corporate trade secrets that politicians and military are watching closely. Research papers are ir/unreproducible he-said-she-saids. Marketing material are research paper and vice versa. Extremely major breakthroughs are announced via tweets, whereas extremely minor results are unannounced via journals. Everything feels simultaneously mostly true and possibly false (because nobody is seriously checking). Nobody knows what's going on, and people who knows what's going on has a non-disclosure clause in their job contract. Is the theory of generalization that we learned in school true or false? It feels false, why hasn't there been any retractions? Many questions like these.
Is it too late to regain some coherence in this field??
https://redd.it/1ve7chh
@datascientology | 410 |
| 14 | Beginner here: My pothole detection model mistakes the roadside for potholes.
https://redd.it/1v90113
@datascientology | 334 |
| 15 | 30+ officially free AI/ML books, all in one curated repo
https://redd.it/1v7cvqr
@datascientology | 369 |
| 16 | Are there some textbooks that take a primarily engineering approach to machine learning (as opposed to a "scientific" approach)? D
As someone who studied stats undergrad and industrial engineering operations research grad, and who thinks about the practical business of ML components in software....
I get lost and a bit hopeless when I think about how to make useful software out of ML models in a reasonable amount of time, and in the current business environment.
And when I look at the businesses where I have worked that have mountains of middle management running tiny bits of the ML model lifecycle (think feature extraction, data ingestion and integration, training infra, hosting infra, more hosting infra, applied science)... that only makes my head hurt even more.
How do you go about making practical software out of ML components?
Edit: I should mention that I mean from scratch ML components, not just a call to a third party hosted tool.
https://redd.it/1v16l6a
@datascientology | 494 |
| 17 | SenseNova-Vision is open-sourced: handle every CV task as unified multimodal generation
https://redd.it/1uyorje
@datascientology | 283 |
| 18 | I reviewed that boy.
https://redd.it/1uwsmbo
@datascientology | 265 |
| 19 | Prompt-engineering paper accepted to ICML R
"Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity"
This paper was accepted to ICML this year. Its main idea is a very simple prompt-engineering trick: "changing the prompt this way led to more diverse sampling". Naturally, it is difficult to provide a rigorous theoretical analysis for something like this.
Even if it works, Iām not sure this kind of prompt engineering belongs at a top-tier machine learning conference. Some people seems to call this kind of work āmodern machine learningā, but I think it should be categorized as less technical venues.
How do you think? Am I being too rigid?
https://redd.it/1uv1xb3
@datascientology | 203 |
| 20 | Hyperparameter tuning approach question R
I am doing some work with cell type classification, where I have 4.3 million cells and 512 features (condensed embeddings from the encoder of a transformer).
The broader goal is to implement a contextual bandit for augmenting the training set of the dataset, as it is currently imbalanced, and rare cell type classification is poor when I tried a baseline logistic regression classifier.
Dataset:
Feature matrix shape: (4290471, 512)
Labels shape: (4290471,)
Class distribution:
T cell 1966941
DC 858451
NK cell 561904
Monocyte 411170
B cell 375882
Platelet 54576
Progenitor cell 24689
ILC 24254
Erythrocyte 12604
I didn't do any hyperparameter tuning for the LR classifier, but I want to try other ML models (LightGBM, XGBoost, SVM)
However, I face a bottleneck with hyperparameter tuning. I want to do 80/10/10 train/validate/test split, but the training set is so large and takes a long time even on H100.
What are some solutions to this? I tried optuna but still very long for each hyperparameter trial. I then tried optuna but instead of using the full 80% for training each time, only 15% of the 80% is used (subsampling from the training set). I'm not sure if this is robust or not. I also couldn't really find anything in the literature.
Anyone been in a similar situation?
https://redd.it/1usa46w
@datascientology | 156 |
