ar
Feedback
Artificial Intelligence||DL

Artificial Intelligence||DL

الذهاب إلى القناة على Telegram

Channel for who have a passion for - * Artificial Intelligence * Machine Learning * Deep Learning * Data Science * Computer vision * IT news Admin: @devuz77

إظهار المزيد
567
المشتركون
لا توجد بيانات24 ساعات
-27 أيام
-330 أيام
أرشيف المشاركات
🐈 TTT Long Video Generation🐈 👉A novel architecture for video generation adapting the CogVideoX 5B model by incorporating Test-Time Training layers. Adding TTT layers into a pre-trained Transformer -> one-minute clip from text storyboards. Videos, code & annotations released💙 👉Review https://t.ly/mhlTN 👉Paper arxiv.org/pdf/2504.05298 👉Project test-time-training.github.io/video-dit/ 👉Repo github.com/test-time-training/ttt-video-dit

🐟Segment Any Motion in Video🐟 👉From CVPR2025 a novel approach for moving object segmentation that combines DINO-based semantic features and SAM2. Code under MIT license💙 👉Review https://t.ly/4aYjJ 👉Paper arxiv.org/pdf/2503.22268 👉Project motion-seg.github.io/ 👉Repo github.com/nnanhuang/SegAnyMo

🔥 Dereflection Any Image 🔥 👉SJTU & #Huawei unveils DAI, novel diffusion-based framework able to recover from a wide range of reflection types. One-step diffusion with deterministic outputs & fast inference. Inference, pretrained models & training released💙 👉Review https://t.ly/PDA9K 👉Paper https://arxiv.org/pdf/2503.17347 👉Project abuuu122.github.io/DAI.github.io/ 👉Repo github.com/Abuuu122/Dereflection-Any-Image

🥎LLM Spatial Understanding🥎 👉SpatialLM by Manycore: novel LLM designed to process 3D point cloud data and generate structured 3D scene understanding outputs. Code, model & data 💙 👉Review https://t.ly/ejr1s 👉Project manycore-research.github.io/SpatialLM/ 👉Code github.com/manycore-research/SpatialLM 🤗Models https://huggingface.co/manycore-research

🧸 Occluded 3D Reconstruction 🧸 👉Oxford unveils a novel 3D generative model to reconstruct 3D objects from partial observations. Code (TBR), demo, model on HF💙 👉Review https://t.ly/Lr5D7 👉Paper arxiv.org/pdf/2503.13439 👉Project sm0kywu.github.io/Amodal3R/ 🤗huggingface.co/spaces/Sm0kyWu/Amodal3R

🔥Distill-Any-Depth: SOTA MDE🔥 👉Distill-Any-Depth is the new SOTA monocular depth estimation model trained with a novel knowledge distillation. Authors: ZJUT, WestLake University, LZU & NTU. Source Code, pre-trained models & HF-demo released💙 👉Review https://t.ly/GBJgi 👉Paper arxiv.org/pdf/2502.19204 👉Repo https://lnkd.in/dPtxNrQh 🤗Demo https://lnkd.in/d2TMPf4b

🧠 Distractor-Aware SAM2 🧠 👉A novel distractor-aware memory for SAM2 and an introspection-based update strategy for VOT. Code & Dataset released💙 👉Review https://t.ly/RBRpQ 👉Paper arxiv.org/pdf/2411.17576 👉Project jovanavidenovic.github.io/dam-4-sam 👉Repo github.com/jovanavidenovic/DAM4SAM/

🧪 SUPIR: SOTA restoration 🧪 👉SUPIR is the new SOTA in image restoration; suitable for restoration of blurry objects, defining the material texture of objects, and adjusting restoration based on high-level semantics 👉Review https://t.ly/wgObH 👉Project https://supir.xpixel.group/ 👉Paper https://lnkd.in/dZPYcUuq 👉Demo coming 🩷 but no code announced :(

🕷️ Gen-NeRF2NeRF Translation 🕷️ 👉GenN2N: unified NeRF-to-NeRF translation for editing tasks such as text-driven NeRF editing, colorization, super-resolution, inpainting, etc. 👉Review https://t.ly/VMWAH 👉Paper arxiv.org/pdf/2404.02788.pdf 👉Project xiangyueliu.github.io/GenN2N/ 👉Code github.com/Lxiangyue/GenN2N

👩‍🦰 SOTA Gaussian Haircut 👩‍🦰 👉ETH et. al unveils Gaussian Haircut, the new SOTA in hair reconstruction via dual representation (classic + 3D Gaussian). Code and Model announced💙 👉Review https://t.ly/aiOjq 👉Paper arxiv.org/pdf/2409.14778 👉Project https://lnkd.in/dFRm2ycb 👉Repo https://lnkd.in/d5NWNkb5

📫MeshPose: DensePose+HMR📫 👉MeshPose: novel approach to jointly tackle DensePose and Human Mesh Reconstruction in a while. A natural fit for #AR applications requiring real-time mobile inference. 👉Review https://t.ly/a-5uN 👉Paper arxiv.org/pdf/2406.10180 👉Project https://meshpose.github.io/

🎹 PianoMotion10M for gen-hands 🎹 👉PianoMotion10M: 116 hours of piano playing videos from a bird’s-eye view with 10M+ annotated hand poses. A big contribution in hand motion generation. Code & Dataset released💙 👉Review https://t.ly/_pKKz 👉Paper arxiv.org/pdf/2406.09326 👉Code https://lnkd.in/dcBP6nvm 👉Project https://lnkd.in/d_YqZk8x 👉Dataset https://lnkd.in/dUPyfNDA

👽Neural-Free Sparse Voxels Rasterization👽 👉#Nvidia unveils a novel efficient radiance field rendering algorithm that incorporates a rasterization process on adaptive sparse voxels without neural networks or 3D Gaussians. Code released (custom license)💙 👉Review https://t.ly/Nh_ic 👉Paper https://lnkd.in/g8k8Zs6R 👉Project https://lnkd.in/gR-bD4Wx 👉Repo https://lnkd.in/gNHX-w4t

🌾 New SOTA Edge Detection 🌾 👉CUP (+ ESPOCH) unveils the new SOTA for Edge Detection (NBED); superior performance consistently across multiple benchmarks, even compared with huge computational cost and complex training models. Source Code released💙 👉Review https://t.ly/zUMcS 👉Paper arxiv.org/pdf/2409.14976 👉Code github.com/Li-yachuan/NBED

🔥 YOLOv12 is out (new SOTA) 🔥 👉YOLOv12 is a novel attention-centric YOLO framework that matches the speed of previous CNN-based ones while harnessing the performance benefits of attention mechanisms. Source Code & Demo released💙 👉Review https://t.ly/jj1oR 👉Paper arxiv.org/pdf/2502.12524 👉Repo github.com/sunsmarterjie/yolov12 🤗Demo https://t.ly/w5rno

🔥 Animate Anyone 2 🔥 👉 The evolution of the first version that enables character animation w/ environment affordance. Amazing results but no code announced 🥲 👉Review https://t.ly/iNNLB 👉Paper https://arxiv.org/pdf/2502.06145 👉Project https://humanaigc.github.io/animate-anyone-2

🤖 META Human-Robot 🤖 👉#META PARTNR: novel benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration. The largest benchmark of its kind: 100,000+ natural language tasks, spanning 60 houses and 5,819 unique objects. Code & Data (🤗) under MIT💙 👉Review https://t.ly/zcN0K 👉Paper arxiv.org/pdf/2411.00081 👉Repo github.com/facebookresearch/partnr-planner 🤗Data huggingface.co/datasets/ai-habitat/partnr_episodes

🛸Real-Time Differentiable Tracing🛸 👉 Radiant Foam is a novel scene representation by leveraging the decades-old efficient volumetric mesh ray tracing algorithm (largely overlooked in recent research). Performing like Gaussian Splatting, without the constraints of rasterization. Code announced💙 👉Review https://shorturl.at/26U06 👉Paper https://arxiv.org/pdf/2502.01157 👉Project https://radfoam.github.io/ 👉Repo https://github.com/theialab/radfoam

🐙MambaGlue: SOTA feats. matching🐙 👉MambaGlue is a hybrid neural network combining the Mamba and the Transformer architectures to match local features. Source Code announced, to be released💙 👉Review https://shorturl.at/LxDG1 👉Paper arxiv.org/pdf/2502.00462 👉Repo https://lnkd.in/dAujfGZQ

🈯SOTA 0-Shot Multi-View Diffusion🈯 👉MVGD by #TOYOTA is the SOTA method that generates images and scale-consistent depth maps from novel viewpoints given an arbitrary number of posed input views. A novel diffusion-based architecture capable of direct pixel-level generation. Code announced 💙 👉Review https://t.ly/_ecKl 👉Paper arxiv.org/pdf/2501.18804 👉Project mvgd.github.io/ 👉Repo TBA