fa
Feedback
Artificial Intelligence||DL

Artificial Intelligence||DL

رفتن به کانال در Telegram

Channel for who have a passion for - * Artificial Intelligence * Machine Learning * Deep Learning * Data Science * Computer vision * IT news Admin: @devuz77

نمایش بیشتر
کشور مشخص نشده استفناوری و برنامه‌ها42 322
567
مشترکین
اطلاعاتی وجود ندارد24 ساعت
-27 روز
-330 روز
آرشیو پست ها
☀️ Relightable Full-Body Avatars ☀️ 👉#Meta unveils the first approach ever to jointly model the relightable appearance of the body, face, and hands of drivable avatars. 👉Review https://t.ly/kx9gf 👉Paper arxiv.org/pdf/2501.14726 👉Project neuralbodies.github.io/RFGCA

🌅 Generative Human Mesh Recovery 🌅 👉GenHMR is a novel generative framework that reformulates monocular HMR as an image-conditioned generative task, explicitly modeling and mitigating uncertainties in 2D-to-3D mapping process. Impressive results but no code announced 🥺 👉Review https://t.ly/Rrzpj 👉Paper https://arxiv.org/pdf/2412.14444 👉Project m-usamasaleem.github.io/publication/GenHMR/GenHMR.html

🔥 [SOTA] Long-Video Depth Anything 🔥 👉ByteDance unveils Video Depth Anything: HQ, consistent depth estimation in SUPER-long videos (over several minutes) without sacrificing efficiency. Based on Depth Anything V2 with a novel efficient spatial-temporal head. Repo available under Apache 2.0💙 👉Review https://t.ly/Q4ZZd 👉Paper arxiv.org/pdf/2501.12375 👉Project https://lnkd.in/dKNwJzbM 👉Repo https://lnkd.in/ddfwwpCj

🌈 #Nvidia Foundation ZS-Stereo 🌈 👉Nvidia unveils FoundationStereo, a foundation model for stereo depth estimation with strong zero-shot generalization. In addition, a large-scale (1M stereo pairs) synthetic training dataset featuring large diversity and high photorealism. Code, model & dataset to be released💙 👉Review https://t.ly/rfBr5 👉Paper arxiv.org/pdf/2501.09898 👉Project nvlabs.github.io/FoundationStereo/ 👉Repo github.com/NVlabs/FoundationStereo/tree/master

🧽 Diffusion Video Inpainting 🧽 👉#Alibaba unveils a technical report about DiffuEraser, a video inpainting model based on stable diffusion, designed to fill masked regions with greater details and more coherent structures. Code & weights released under Apache💙 👉Review https://t.ly/7rEll 👉Paper arxiv.org/pdf/2501.10018 👉Project lixiaowen-xw.github.io/DiffuEraser-page/ 👉Repo github.com/lixiaowen-xw/DiffuEraser

🏄‍♀️ GSTAR: Gaussian Surface Tracking 🏄‍♀️ 👉ETH Zurich unveils GSTAR, a novel framework for photo-realistic rendering, surface reconstruction, and 3D tracking for dynamic scenes while handling topology changes. Code announced💙 👉Review https://t.ly/udpMq 👉Paper arxiv.org/pdf/2501.10283 👉Project chengwei-zheng.github.io/GSTAR/ 👉Repo TBA

🎁Free Book: LLM Foundations🎁 👉A fully free book just released on arXiv to outline the basic concepts of #LLMs and related techniques with a focus on the foundational aspects. ✅Chapter 1: basics of pre-training ✅Chapter 2: gen-models & LLMs ✅Chapter 3: prompting methods ✅Chapter 4: alignment methods 👉If you have any background in ML, along with a certain understanding of stuff like Transformers, this book will be "smooth". However, even without this prior knowledge, it is still perfectly fine because the contents of each chapter are self-contained. 👉Review https://t.ly/9LGCa 👉Book https://lnkd.in/d3VkswZf

🧞‍♂️Omni-RGPT: SOTA MLLM Understanding🧞‍♂️ 👉 #NVIDIA presents Omni-RGPT, MLLM for region-level comprehension for both images & videos. New SOTA on image/video-based commonsense reasoning. 👉Review https://t.ly/KHnQ7 👉Paper arxiv.org/pdf/2501.08326 👉Project miranheo.github.io/omni-rgpt/ 👉Repo TBA soon

🏆Universal Detector-Free Match🏆 👉MatchAnything: novel detector-free universal matcher across unseen real-world single/cross-modality domains. Same weights for everything. Code announced, to be released 💙 👉Review https://t.ly/sx92L 👉Paper https://lnkd.in/dWwRwGyY 👉Project https://lnkd.in/dCwb2Yte 👉Repo https://lnkd.in/dnUXYzQ5

🔥 Depth Any Camera (SOTA) 🔥 👉DAC is a novel and powerful zero-shot metric depth estimation framework that extends a perspective-trained model to effectively handle cams with varying FoVs (including large fisheye & 360◦). Code announced (not available yet)💙 👉Review https://t.ly/1qz4F 👉Paper arxiv.org/pdf/2501.02464 👉Project yuliangguo.github.io/depth-any-camera/ 👉Repo github.com/yuliangguo/depth_any_camera

⚽ FIFA 3D Human Pose ⚽ 👉#FIFA WorldPose is a novel dataset for multi-person global pose estimation in the wild, featuring footage from the 2022 World Cup. 2.5M+ annotation, released 💙 👉Review https://t.ly/kvGVQ 👉Paper arxiv.org/pdf/2501.02771 👉Project https://lnkd.in/d5hFWpY2 👉Dataset https://lnkd.in/dAphJ9WA

🧤World-Space Ego 3D Hands🧤 👉The Imperial College unveils HaWoR, a novel world-space 3D hand motion estimation for egocentric videos. The new SOTA on both cam pose estimation & hand motion reconstruction. Code under Attribution-NC-ND 4.0 Int.💙 👉Review https://t.ly/ozJn7 👉Paper arxiv.org/pdf/2501.02973 👉Project hawor-project.github.io/ 👉Code github.com/ThunderVVV/HaWoR

🥮 SOTA probabilistic tracking🥮 👉ProTracker is a novel framework for robust and accurate long-term dense tracking of arbitrary points in videos. Code released under CC Attribution-NonCommercial💙 👉Review https://t.ly/YY_PH 👉Paper https://arxiv.org/pdf/2501.03220 👉Project michaelszj.github.io/protracker/ 👉Code github.com/Michaelszj/pro-tracker

🌳 HD Video Object Insertion 🌳 👉VideoAnydoor is a novel zero-shot video object insertion #AI with high-fidelity detail preservation and precise motion control. All-in-one: video VTON, face swapping, logo insertion, multi-region editing, etc. 👉Review https://t.ly/hyvRq 👉Paper arxiv.org/pdf/2501.01427 👉Project videoanydoor.github.io/ 👉Repo TBA

🌲MagicEdit: Magic Video Edit🌲 👉MagicEdit: explicit disentangling content, structure & motion for Hi-Fi and temporally coherent video editing 😎Report https://t.ly/tREX4 😎Paper arxiv.org/pdf/2308.14749.pdf 😎Project magic-edit.github.io 😎Code github.com/magic-research/magic-edit

⭐TOP 10 Papers you loved - 2024⭐ 👉Here the list of my posts you liked the most in 2024, thank you all 💙 𝐏𝐚𝐩𝐞𝐫𝐬: ⭐"Look Ma, no markers" ⭐T-Rex 2 Detector ⭐Models at Any Resolution 👉The full list with links: https://t.ly/GvQVy

📞FacET: VideoCall Change Your Expression📞 👉Columbia University unveils FacET: discovering behavioral differences between conversing face-to-face (F2F) and on video-calls (VCs). 👉Review https://t.ly/qsQmt 👉Paper arxiv.org/pdf/2406.00955 👉Project facet.cs.columbia.edu/ 👉Repo (empty) github.com/stellargo/facet

🧊 Universal 6D Pose/Tracking 🧊 👉Omni6DPose is a novel dataset for 6D Object Pose with 1.5M+ annotations. Extra: GenPose++, the novel SOTA in category-level 6D estimation/tracking thanks to two pivotal improvements. 👉Review https://t.ly/Ywgl1 👉Paper arxiv.org/pdf/2406.04316 👉Project https://lnkd.in/dHBvenhX 👉Lib https://lnkd.in/d8Yc-KFh

🔄️ Orient Anything in 3D 🔄️ ️ 👉Orient Anything is a novel robust image-based object orientation estimation model. By training on 2M rendered labeled images, it achieves strong zero-shot generalization in the wild. Code released💙 👉Review https://t.ly/ro5ep 👉Paper arxiv.org/pdf/2412.18605 👉Project orient-anything.github.io/ 👉Code https://lnkd.in/d_3k6Nxz

🧬Event-driven SuperResolution🧬 👉USTC unveils EvTexture, the first VSR method that utilizes event signals for texture enhancement. It leverages high-freq details of events to better recover texture in VSR. Code available💙 👉Review https://t.ly/zlb4c 👉Paper arxiv.org/pdf/2406.13457 👉Code github.com/DachunKai/EvTexture