ch
Feedback
Artificial Intelligence AI News

Artificial Intelligence AI News

前往频道在 Telegram

We are a community of machine learning enthusiasts/researchers/journalists/writers who share interesting news and articles about the applications of AI. You will never miss any updates on ML/AI/CV/NLP fields because we post them daily. JOIN NOW

显示更多
未指定国家技术与应用22 105
3 443
订阅者
+724 小时
+197 天
+10430 天
吸引订阅者
9月 '26
九月 '26
+146
在0个频道中
八月 '26
+160
在0个频道中
Get PRO
七月 '26
+132
在1个频道中
Get PRO
六月 '26
+134
在2个频道中
Get PRO
五月 '26
+150
在1个频道中
Get PRO
四月 '26
+456
在2个频道中
Get PRO
三月 '26
+846
在2个频道中
Get PRO
二月 '26
+141
在0个频道中
Get PRO
一月 '26
+81
在1个频道中
Get PRO
十二月 '25
+58
在3个频道中
Get PRO
十一月 '25
+52
在0个频道中
Get PRO
十月 '25
+74
在1个频道中
Get PRO
九月 '25
+16
在0个频道中
Get PRO
八月 '25
+16
在0个频道中
Get PRO
七月 '25
+14
在0个频道中
Get PRO
六月 '25
+15
在0个频道中
Get PRO
五月 '25
+20
在0个频道中
Get PRO
四月 '25
+18
在0个频道中
Get PRO
三月 '25
+33
在0个频道中
Get PRO
二月 '25
+43
在0个频道中
Get PRO
一月 '25
+30
在1个频道中
Get PRO
十二月 '24
+32
在2个频道中
Get PRO
十一月 '24
+107
在0个频道中
Get PRO
十月 '24
+164
在0个频道中
Get PRO
九月 '24
+136
在0个频道中
Get PRO
八月 '24
+1 029
在0个频道中
日期
订阅者增长
提及
频道
28 九月+2
27 九月+7
26 九月+5
25 九月+4
24 九月+3
23 九月+5
22 九月+7
21 九月+4
20 九月+6
19 九月+7
18 九月+7
17 九月+7
16 九月+6
15 九月+5
14 九月+7
13 九月+5
12 九月+3
11 九月+5
10 九月+9
09 九月+4
08 九月0
07 九月+3
06 九月+4
05 九月+10
04 九月+4
03 九月+6
02 九月+3
01 九月+8
频道帖子
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time It answers "wh
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time It answers "who spoke when" in a conversation, even when people talk over each other. One open-weight checkpoint handles both offline audio and real-time streaming. ✅ 100M parameters, open weights on Hugging Face ✅ OpenMDW-1.1 license, commercial use allowed ✅ Tracks up to 8 speakers, overlap included ✅ 14.72% DER on Diarization-Bench vs 19.3% runner-up ✅ 4 streaming presets, from 30.4 s down to 0.32 s ✅ 41.0% average relative DER cut vs Streaming Sortformer ✅ 15,113x batched RTFx on RTX PRO 5000, 30.4 s preset ✅ Runs through NVIDIA NeMo on Linux GPUs Full breakdown + interactive explainer: https://www.marktechpost.com/2026/09/23/nvidia-releases-nemotron-3-diarization/ Model: https://huggingface.co/nvidia/Nemotron-3-Diarization NVIDIA blog: https://huggingface.co/blog/nvidia/nemotron-diarization Live demo: https://huggingface.co/spaces/nvidia/nemotron-diarization

2
Nokia AI Research Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model It read
Nokia AI Research Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model It reads typed answers straight from the model's next-token distribution. It removes position and label bias without any training. ▪️ Apache-2.0, pip install "anyjev[hf]" ▪️ Choice, yes/no and score questions, nothing generated ▪️ Hugging Face transformers and vLLM backends ▪️ L0 needs zero labels, L1 needs 100 to 500 ▪️ Order-flip rate 0.230 to 0.073 at L0 ▪️ Calibration error 0.240 to 0.095 at L1 ▪️ Auto-decidable at 5% error: 7.7% to 52.0% at L1 ▪️ Tested on Qwen3-8B, BANKING77 20-way, 300 test items Accuracy moves 6 points, but the traffic you can safely automate grows 6.8x. Full analysis: https://www.marktechpost.com/2026/09/23/nokia-open-sources-anyjev-a-training-free-layer-that-turns-any-open-llm-into-a-calibrated-decision-model/ Repo: https://github.com/nokia-applied-research/AnyJev
454
3
TypeSafe AI released Jev, a model from an RLHF (reinforcement learning from human feedback) co-inventor that does not generat
TypeSafe AI released Jev, a model from an RLHF (reinforcement learning from human feedback) co-inventor that does not generate text. You send it state and typed questions. It returns decisions with probabilities your code can branch on, so there is nothing to parse. - Founder Diogo Almeida co-invented RLHF at OpenAI - $40M seed round led by DCVC - 3 primitives: Choice, Score, Noul - Every question in a request scored in parallel - $0.042 per 1M input tokens, output tokens free - 70ms to 500ms end to end, vendor reported - 0.114s vs 8.566s in TypeSafe's own recorded demo - Vercel: up to 18x faster than GPT Luna, p95 - Listed on Vercel AI Gateway as typesafe-ai/jev Confidence is the product. Your code sets the threshold where automation stops and a human takes over. The 193.6x faster and 444.6x cheaper claims come from TypeSafe's own evals, not independent tests. Read the full analysis: https://www.marktechpost.com/2026/09/19/typesafe-ai-releases-jev/ Technical details: https://typesafe.ai/blog/introducing-system-one-models-and-jev Join the waitlist: https://typesafe.ai/
644
4
Linkup just open-sourced SPARSEUP, a 149M sparse embedding model scoring 56.4 on BEIR-13. It is a SPLADE-style retriever that
Linkup just open-sourced SPARSEUP, a 149M sparse embedding model scoring 56.4 on BEIR-13. It is a SPLADE-style retriever that outputs readable token weights, not an opaque dense vector. It fills the sparse slot next to LightOn's DenseOn and LateOn, trained on the same open data. ✅ Apache 2.0, weights on Hugging Face ✅ ModernBERT backbone, English only ✅ 56.4 nDCG@10 on BEIR-13, excluding MS MARCO ✅ Beats splade-v3 (51.7) and opensearch doc-v3-gte (54.6) ✅ ~380µs per query at >97% recall, Seismic, MS MARCO ✅ 47 non-zeros per query, 190 per document ✅ Logit shift of 15 keeps vectors sparse from the start ✅ Top-12 expansion cap per token, not per vector ✅ Case folding cuts output from ~50k to ~34k dims Every dimension is a token, so you can read exactly why a document matched. With identical backbone and data, it still trails DenseOn by 1.52 points and LateOn by 2.5. Linkup also documents weak expansions, like number queries spilling into unrelated terms. Full breakdown: https://www.marktechpost.com/2026/09/19/linkup-research-releases-sparseup/ Model: https://huggingface.co/Linkup-Platform/linkup-sparseup-embed-v1 Technical blog: https://www.linkup.so/blog/introducing-sparseup-by-linkup
602
5
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data P
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data Paper2Agent turns a research paper and its codebase into a tested AI agent. You don't have to clone repos, install dependencies or debug environments before using a paper's method. It builds an MCP server from the paper's code, then generates and runs tests against the paper's own reference outputs. Tools that keep failing are excluded, so every tool in the final server has passed validation. You can connect the server to Claude Code or any MCP-compatible agent and ask it to apply the paper's method to your own data in plain language. On the AlphaGenome paper, Paper2Agent built 22 validated tools in about 45 minutes for US $14. The resulting agent scored 100% on 15 novel queries, compared with 78.7% for Claude Code with direct repo access and 56.0% for Biomni. Across 100 computational biology papers from bioRxiv, 74 were agentified and 593 of 599 proposed tools passed validation. On 300 benchmark questions, it reached 91.2% accuracy at US $0.20 per query. Learn more with the following resources: 📰 Full breakdown: https://www.marktechpost.com/2026/09/16/stanford-researchers-release-paper2agent-turning-research-papers-into-ai-agents-that-reproduce-results-and-run-on-new-data/ 📄 Paper: https://www.nature.com/articles/s41586-026-11044-y ⭐ GitHub: https://github.com/jmiao24/Paper2Agent 🤗 AlphaGenome MCP server: https://huggingface.co/spaces/Paper2Agent/alphagenome_mcp 🧪 Try it: https://paper2agent.ai/live
612
6
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, native speech to speech models built for production voice age
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, native speech to speech models built for production voice agents. Both run in the Gemini Live API and replace cascaded ASR plus LLM plus TTS pipelines with a single model. The defining capability: the conversation never pauses while reasoning and tool calls finish in the background. ✅ #1 on Artificial Analysis Speech to Speech Index, 82.6 ✅ Executes tools and API calls in the background mid conversation ✅ Auto switches between 97 supported languages mid conversation ✅ $0.005/min audio input, $0.018/min audio output Extended Thinking still scores just 35.1% on Sierra's τ-Voice-banking benchmark, so hard multi-step voice workflows are far from solved. Weights are not open; access is API only, with enterprise access in private preview. Full analysis: https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/ Technical details: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
630
7
NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing Here's how it works. 👇 1
NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing Here's how it works. 👇 1. One YAML, three computers Robot pipelines span 3 compute tiers: training on GB200/H100, simulation on RTX PRO 6000, hardware-in-the-loop on Jetson AGX Thor. OSMO treats all 3 as backends of one control plane. → Tasks name a platform (gb200, rtx-pro-6000, jetson-agx-thor), never a cluster 2. Dependencies are data The README example chains 3 tasks: Isaac Sim → PyTorch training with 8 GPUs → ROS eval on Jetson, writing results to a named dataset. → Placement from platform, ordering from inputs, persistence from outputs 3. Same file, laptop to cloud Runs on KIND locally and on EKS, AKS, GKE, on-prem, or air-gapped clusters. Release 6.3.0 added a multi-provider deploy script with MinIO, Azure Blob, or S3 storage. → Zero code changes between environments 4. Built for production NVIDIA KAI Scheduler by default, NVLink topology-aware placement, per-group timeouts, RBAC sidecar, OAuth2 login, TLS at the gateway, cloud workload identity. → Powers Project GR00T, Isaac Lab, Isaac Sim, and Isaac ROS internally Full analysis: https://www.marktechpost.com/2026/09/14/nvidia-open-sources-osmo-one-yaml-orchestrates-physical-ai-training-simulation-and-robot-testing/ Repo: https://github.com/NVIDIA/OSMO
641
8
A Princeton Researcher Proposes Recurrent Looped Transformer (RLT): a decoder that carries its full state across every prompt
A Princeton Researcher Proposes Recurrent Looped Transformer (RLT): a decoder that carries its full state across every prompt and response token, with 96 logical blocks per token and no reset at the boundary. Here's how it works. 👇 1. One state, never reset A causal encoder builds KV memory in parallel. A recurrent decoder merges each token's encoder feature with the previous final hidden state, then runs sliding-window attention, cross-attention to encoder memory, and an FFN. The state H_t = (s_t, C^D_t) crosses the serving split untouched. → Proposition 3.1: moving the prompt-response split does not change the conditional distribution 2. Fixed work, growing depth Reference config ties 48 encoder and 48 decoder layers with shared attention and FFN weights. → 96 logical blocks per token, fixed → 48t decoder blocks on the state path after t tokens 3. RL replay under current parameters The sampler records behavior log-probs. The trainer rebuilds encoder memory, recurrent output, and every SWA cache from H_0 before scoring each action. → Old rollout states are never reused; weight updates invalidate all caches The key takeaway: a design spec that makes pretraining, SFT, sampling, and RL replay share one transition. Full analysis: https://www.marktechpost.com/2026/09/13/a-princeton-researcher-proposes-recurrent-looped-transformer-rlt/ Paper: https://github.com/yifanzhang-pro/recurrent-looped-tranformer/blob/master/Recurrent_Looped_Transformer.pdf Git Repo: https://github.com/yifanzhang-pro/recurrent-looped-tranformer Amazing explanation in author's research blog: https://yifanzhang-pro.github.io/recurrent-looped-tranformer/
619
9
Sparse attention already cut long-context compute. The KV cache sitting in HBM and on SSD is now the bottleneck DeepSeek AI w
Sparse attention already cut long-context compute. The KV cache sitting in HBM and on SSD is now the bottleneck DeepSeek AI went after, and they cut theirs to 890 bytes per token. They released DeepSeek-V4.1-Flash, a 552B MoE model with 1M-token context that activates only 8B parameters per token during prefill and 16B during decode. The global KV cache footprint is roughly 1/4 of DeepSeek-V4-Flash and about 437x smaller than DeepSeek-V1, and the weights are open under an MIT license. Here's what's actually interesting: → Causal Encoder-Decoder: prompt tokens stop at layer 20, decoder global KV is projected from the encoder's final hidden state, prefill compute nearly halves → Compressed Sparse Attention 2: three statically assigned layer modes, Reindex reuses main KV and indexer K, Reuse also reuses Top-K indices → FP4 main KV cache via quantization-aware training, SWA Bounded Replay replays only 128 tokens instead of layers times window → Terminal-Bench 2.1: 90.6 vs 89.1 for Opus-5 → DeepSWE v1.1: 74.2 vs 74.0 for Opus-5 and 73.0 for GPT-5.6 Sol → Terminal-Bench 4.0: 31.2 vs 51.8 for Opus-5, so the gap on expert-level science tasks is still real Full analysis: https://www.marktechpost.com/2026/09/10/deepseek-ai-released-deepseek-v4-1-flash-with-1m-context-fp4-kv-cache-and-cross-layer-attention-reuse/ Technical Details: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
723
10
NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels NVIDIA has announced
NVIDIA Announces CUDA Rust with cuda-oxide (SIMT) and cutile-rs (Tile) for Compile-Time-Safe GPU Kernels NVIDIA has announced CUDA Rust, a push to make Rust a first-class language for writing GPU kernels. Rust code could already launch CUDA kernels, but the kernel body usually had to be written elsewhere. CUDA Rust closes that gap with two NVlabs open-source projects: cuda-oxide for the SIMT model and cutile-rs for the newer Tile model. Both compile Rust kernels natively and use Rust’s ownership rules to reject aliasing bugs at compile time. Here's how it works. 👇 1. cuda-oxide (SIMT track) A custom rustc codegen backend. #[kernel] functions go from Rust MIR through the Pliron IR framework and LLVM down to PTX. Host and device code live in one file. → Needs a pinned nightly (2026-04-03), CUDA 12.x+, compute capability 8.0+ → Status: early alpha 2. cutile-rs (Tile track) Each tile block runs the kernel body once as a single logical thread. The compiler owns thread mapping and memory layout. The kernel AST is embedded in the host binary and JIT-compiled through CUDA Tile IR on first launch. → Stable Rust 1.89+, CUDA 13.3, no nightly, no custom LLVM → Published on crates.io, already used in Hugging Face's Grout... 3. What the compiler catches Pass a kernel's output buffer as one of its own inputs and it does not build. → SIMT: error[E0502]: cannot borrow c_dev as mutable because it is also borrowed as immutable → Tile: error[E0382]: use of moved value: z cutile-rs carries ownership across the launch boundary, which NVIDIA calls the stronger guarantee.... Full analysis: https://www.marktechpost.com/2026/09/08/nvidia-announces-cuda-rust-with-cuda-oxide-simt-and-cutile-rs-tile-for-compile-time-safe-gpu-kernels/ Technical details: https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels
686
11
Google DeepMind Releases AlphaGenome Atlas: Precomputed Molecular Effect Predictions and AVI Scores for All 9 Billion Single-
Google DeepMind Releases AlphaGenome Atlas: Precomputed Molecular Effect Predictions and AVI Scores for All 9 Billion Single-Letter DNA Changes in the Human Genome. No lab assay. No per-variant model run. No coding required to query it. Here is how it works: 1. Precompute instead of predict on demand AlphaGenome was run across every possible single-nucleotide variant in the human genome and the outputs were stored. Researchers now look up a variant instead of running a model on it. → 9 billion variants scored → 1 petabyte dataset, over 30x the size of the AlphaFold Database 2. One score for coding and non-coding DNA The AlphaGenome Variant Impact (AVI) score folds AlphaGenome's regulatory predictions together with AlphaMissense's protein-impact predictions into a single rankable number. → Works across the 2% coding genome and the 98% non-coding genome → DeepMind reports best-in-class results on variant pathogenicity and rare disease benchmarks 3. The score is decomposed, not opaque Each AVI score splits into additive feature attributions across categories like chromatin accessibility, splicing, and conservation, so you see which process a variant is predicted to disrupt. → Paired with a compendium of 2,500+ recurrent DNA motifs and their locations 4. Rare disease result At the Broad Institute, the AVI score reprioritized variants that earlier analyses had missed and surfaced one in DNM1, a gene linked to epileptic encephalopathy. The prediction: an incorrect splice site extending the protein. → Experimental screens validated it and found nearby variants with similar effects.... Full analysis: https://marktechpost.com/2026/09/08/google-deepmind-releases-alphagenome-atlas-with-precomputed-molecular-effect-predictions-and-avi-scores-for-9-billion-human-dna-variants/ Paper: https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/alphagenome-atlas.pdf Technical details: https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/
557
12
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal D
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder No SigLIP2 tower. No causal decoder. No VLM to repurpose. Here's how it works. 👇 (1) Raw patches, not vision-tower features Images are split into non-overlapping 32×32 RGB patches and projected by a 2-layer MLP trained from scratch. Text enters through a 256-dimensional factorized embedding. Both then share every Transformer layer. → 16,384-token context, enough for two 3840×2160 4K UHD images (2) Masked diffusion, not masked language modeling Pretraining is a discrete masked-diffusion text denoiser. Text-only segments draw a corruption rate from U(0,1). Multimodal segments draw from U(0.30,1), which kills the "guess it from the surrounding words" shortcut. → +38.4 points masked-token accuracy from visible page patches at 90% masking (260M) (3) Trained from scratch on a small budget About 524B packed input tokens, roughly 290B of them text-only. ModernBERT saw around 2T text tokens. They used the NorMuon optimizer to squeeze more out of the smaller budget. → 16 H100s for the 260M run, 32 for the 800M Full analysis: https://www.marktechpost.com/2026/09/06/h-company-releases-neomme-a-family-of-260m-and-800m-single-tower-multimodal-encoders-that-drop-the-vision-tower-and-causal-decoder/ Paper: https://arxiv.org/pdf/2609.01657 Technical details: https://huggingface.co/blog/Hcompany/neomme? HF: https://huggingface.co/collections/Hcompany/neomme
583
13
NVIDIA Releases Personal AI Router (PAIR): An Open Inference Router That Turns The RTX, DGX Spark And Mac Boxes You Already O
NVIDIA Releases Personal AI Router (PAIR): An Open Inference Router That Turns The RTX, DGX Spark And Mac Boxes You Already Own Into One Local AI Cluster. No new cluster API. No agent harness changes. No prompts leaving your network. Here's how it works. 👇 1. It routes, it doesn't execute PAIR is not a new inference engine. Ollama or LM Studio still runs the model on whichever machine PAIR picks. It takes over the default port each engine uses, so the agent keeps talking to the endpoint it already knows. → Proxies Ollama-compatible, LM Studio-compatible and OpenAI-compatible endpoints → The agent decides what to request, PAIR decides where it runs 2. Discovery and trust mDNS finds nearby machines automatically, or you add a node by IP. Trust is bootstrapped by a six-digit PIN shown on one machine and entered on the other. → All node-to-node traffic is blocked until pairing completes → Paired nodes then communicate over mTLS with generated certificates 3. The eligibility filter A node only becomes a candidate once it can actually serve the request. The scheduler weighs five signals: is the node online and ready, is a supported engine enabled, is the exact model present, what is the current job load, what is GPU utilization. → Models don't need to be identical across nodes — PAIR routes by model location → Loading the same tag on more nodes just widens the eligible pool Full analysis: https://www.marktechpost.com/2026/09/04/nvidia-releases-personal-ai-router-pair-an-open-source-virtual-inference-router-that-distributes-local-ai-requests-across-rtx-dgx-spark-and-mac-nodes/ Repo: https://github.com/NVIDIA/Personal-AI-Router Technical details: https://www.nvidia.com/en-us/ai-on-rtx/personal-ai-router/
757
14
Anthropic Released Claude Commerce Agents: An Apache-2.0 Blueprint for Shopping and Merchant Agents Across Retail, Travel, Te
Anthropic Released Claude Commerce Agents: An Apache-2.0 Blueprint for Shopping and Merchant Agents Across Retail, Travel, Telecom, and Entertainment. No intent router. No subagent per domain. No custom markup for the UI. Here's how it works. 👇 1. One agent loop, skills for the long tail A commerce session is one tightly coupled conversation across many intents. Every handoff to a subagent is state-lossy — the orchestrator holds the cart, the preferences, the history. → Anthropic reports a single agent with skills beat both the one-big-prompt and the subagent design on quality, often at lower cost and latency 2. Prompt or skill, decided by frequency Loading a skill costs a model turn, so anything the agent needs on most turns goes in the system prompt. Safety rules, brand constraints, and key user facts always sit there. → Roughly a third or more of traffic → system prompt; everything else → skill → 5 skills on the shopping agent, 5 on the merchant agent 3. UI components are tools, not tags Most commerce responses are carousels, itineraries, and seat maps. The model calls present_products or present_itinerary with typed arguments; the server validates and the client renders. Old conversations reload without a custom parser. → The layout lives in the messages array, so "the third one down" resolves 4. Prompt caching carries the cost Caching is prefix-based, so the request is ordered global → session → volatile. A timestamp at the top of the system prompt breaks the cache on every request. → 90–99% hit rate is the range Anthropic says to design for → Cached reads cost a tenth of fresh tokens; cache writes carry a ~1.25x premium Full analysis: https://www.marktechpost.com/2026/09/03/anthropic-released-claude-commerce-agents-an-apache-2-0-blueprint-for-shopping-and-merchant-agents-across-retail-travel-telecom-and-entertainment/ Repo: https://github.com/anthropics/commerce-agents Technical details: https://claude.com/blog/the-anatomy-of-effective-commerce-agents
787
15
Meta AI Released Muse Spark 1.3: An Agentic Coding Model Doing the Same Work With ~20% Fewer Tool Calls and ~25% Fewer Tokens
Meta AI Released Muse Spark 1.3: An Agentic Coding Model Doing the Same Work With ~20% Fewer Tool Calls and ~25% Fewer Tokens Than Muse Spark 1.2. No price increase. No new harness. No open weights either. Here's how it works. 👇 1. Fewer round trips, not just better answers Meta trained 1.3 to take fewer turns where they aren't needed, with less verbosity and a cleaner coding style. → ~20% fewer tool calls and ~25% fewer tokens in Meta's internal engineer comparisons 2. It asks instead of guessing On ambiguous prompts it asks a clarifying question. When it stalls it invokes you. Before consequential actions it confirms. → Better calibration on what counts as irreversible 3. One thread, several workflows Given an open-ended objective, it generates its own context from messy and conflicting sources and patches gaps in its own plan. → Maps an incoming prompt to the right task inside a cluttered thread, whether you're steering or interrupting 4. The numbers (Meta's launch scorecard) → 75.4 on DeepSWE v1.1, ahead of Claude Opus 5 at 74.0 and GPT-5.6 Sol at 72.7 → 88.8 on Terminal-Bench 2.1, tied with GPT-5.6 Sol → 59.4 on SWE-Atlas Codebase QnA → 98.5 and 98.1 on MRCR v2 long-context retrieval, inside a 1,048,576-token window Full analysis: https://marktechpost.com/2026/09/03/meta-ai-released-muse-spark-1-3-an-agentic-coding-model-that-uses-20-fewer-tool-calls-and-25-fewer-tokens-than-muse-spark-1-2/ Technical details: https://research.meta.ai/blog/introducing-muse-spark-1-3
595
16
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon Most "runs locally on your
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon Most "runs locally on your Mac" stacks are a general-purpose runtime pointed at whatever model you downloaded. Perplexity just argued that the generality itself is the bottleneck. They open sourced Lily — the local inference engine behind Hybrid Compute in Perplexity Computer. A Rust runtime with hand-written Metal kernels, built for exactly one model (Qwen3.6-35B-A3B) on exactly one chip family (Apple silicon). Neither PyTorch nor MLX is anywhere in the execution path. Here's what's actually interesting: → 4,156 vs 3,388 prefill tokens/s and 170.0 vs 126.4 decode tokens/s against MLX-LM — mean across ten lengths from 256 to 128K tokens, batch 1, one 40-core / 128 GB M5 Max → Fusing 4-bit dequantization into the grouped GEMM, so the expanded weight array never touches unified memory: +77.4% prefill at a 512-token prompt → Keeping the whole routing sequence — histogram, prefix scan, scatter, block map — inside one GPU command buffer: +89% prefill at 512 tokens → GQA packing, so four query heads share one KV row load: +23.8% decode at 32K context → Fixed-block attention layout above 32K: +40.2% decode at 128K Full analysis: https://www.marktechpost.com/2026/09/02/perplexity-open-sources-lily-a-rust-metal-inference-engine-for-qwen3-6-35b-a3b-on-apple-silicon/ GitHub: https://github.com/perplexityai/pplx-garden/tree/main/lily Technical details: https://www.perplexity.ai/hub/blog/optimizing-on-device-inference-for-apple-silicon
667
17
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointi
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing Most production voice stacks are three systems stitched together. One model transcribes, a second separates speakers, and a detector decides when the user stopped talking. Each hand-off adds latency and a new failure mode. Muse Voice Transcribe, announced by Meta Superintelligence Labs this week, collapses those three jobs into a single autoregressive model. Meta calls it its first real-time audio perception model. It performs streaming ASR, speaker diarization for 20+ speakers, and endpointing in one pass, with no required post-processing.... Full analysis: https://www.marktechpost.com/2026/09/01/meta-superintelligence-labs-releases-muse-voice-transcribe-one-real-time-model-for-streaming-asr-diarization-and-endpointing/ Technical details: https://x.com/AIatMeta/status/2094839236016976028
653
18
Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, Gated On Device Agentic assistants
Perplexity Releases Hybrid Compute on Mac: Cloud Agents Orchestrate Down to a Local Model, Gated On Device Agentic assistants have a structural problem: the context that makes them useful — deal documents, privileged files, client records — is exactly the context users cannot send to a cloud endpoint. This week, Perplexity shipped its answer for Mac. Hybrid compute splits a single Perplexity Computer task between frontier models in the cloud and a compact model on the user’s Mac, with an on-device privacy gate deciding what may cross the boundary. Perplexity also open-sourced the classifier behind that gate.... Full read: https://www.marktechpost.com/2026/09/01/perplexity-releases-hybrid-compute-on-mac-cloud-agents-orchestrate-down-to-a-local-model-gated-on-device/ Technical details: https://www.perplexity.ai/hub/blog/introducing-hybrid-compute-on-mac
646
19
Google AI Releases TimesFM-3: A 330M Parameter Zero-Shot Foundation Model For Multivariate Time Series Forecasting TimesFM-3
Google AI Releases TimesFM-3: A 330M Parameter Zero-Shot Foundation Model For Multivariate Time Series Forecasting TimesFM-3 is a 330 million parameter time series foundation model that forecasts multiple related series in a single forward pass. Every TimesFM checkpoint through 2.5 was univariate: one series, its own history, nothing else. TimesFM-3 is pretrained natively for multivariate forecasting on more than 1 trillion time points, and accepts multiple targets, past covariates, and past-future covariates with no task-specific fine-tuning. It takes the top average rank among pretrained foundation models on GIFT-Eval, fev-bench, and the TIME leaderboard, on both point and probabilistic metrics...... Full analysis: https://www.marktechpost.com/2026/08/31/google-ai-releases-timesfm-3-a-330m-parameter-zero-shot-foundation-model-for-multivariate-time-series-forecasting/ HF Card: https://huggingface.co/google/timesfm-3.0-pytorch Technical details: https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/
682
20
I've read a lot of environment-scaling papers this year. This is the first one that doesn't generate anything. Google AI Intr
I've read a lot of environment-scaling papers this year. This is the first one that doesn't generate anything. Google AI Introduces EnvHarness: A Programmable Layer That Turns Static Agent Environments Into Adaptive Training Worlds It wraps an existing environment through the standard reset/step interface, so the original tasks and human-built verifiers stay in place. An LLM designer writes the wrappers against flaws it finds in the agent's own rollouts. - Apache-2.0, code and reproduction drivers on GitHub - Three components: Stage, Contract, Chain - Stage replays actions to move the episode start state - Contract hooks actions, transitions and observations per step - Chain joins two environments into one episode - EnvRigger loop: observe, diagnose, write, validate - Five benchmarks, four domains, one interface - +9.0 points on held-out ALFWorld tasks - 49.6 vs 55.0 average steps on SWE-bench Verified Full analysis: https://www.marktechpost.com/2026/08/30/google-ai-introduces-envharness-a-programmable-layer-that-turns-static-agent-environments-into-adaptive-training-worlds/ Paper: https://arxiv.org/pdf/2608.19880 GitHub Repo: https://github.com/google-research/envharness
733