uk
Feedback
🚨 AI News | TestingCatalog

🚨 AI News | TestingCatalog

Відкрити в Telegram

Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors 🗞

Показати більше
Країна не вказанаТехнології та додатки14 474
7 462
Підписники
+924 години
+667 днів
+28430 днів
Архів дописів
GOOGLE 🔥: Gemini 3.8 Flash is rolling out on Gemini, Google AI Studio and APIs. Gemini 3.8 Flash scores 71% on DeepSWE 1.1,
GOOGLE 🔥: Gemini 3.8 Flash is rolling out on Gemini, Google AI Studio and APIs. Gemini 3.8 Flash scores 71% on DeepSWE 1.1, compared to 74% for Claude Opus 5, at a much lower price. > Input price $0.75 through December 31, 2026. $1.50 starting January 1, 2027. > Output price (including thinking tokens) $3.75 through December 31, 2026. $7.50 starting January 1, 2027. This is big 👀

Wonderful raises $550M for its enterprise AI agent platform Wonderful raised $550M in Series C funding at a $5B valuation to expand AI agents for enterprise work. Backed by Insight Partners and Salesforce, it pairs software with embedded engineering teams to move deployments into production fast. 🗞 #sponsored @testingcatalog

GOOGLE 🔥: Gemini 3.8 Flash is already available in Agent Studio on GCP. Best for - Complex multimodal data processing - Codi
GOOGLE 🔥: Gemini 3.8 Flash is already available in Agent Studio on GCP. Best for - Complex multimodal data processing - Coding use cases - Supporting software engineering–related agentic tasks Use case - Processing data with images and text - Coding problems - Web research and application testing

Catch launches an AI assistant with phone calling support Catch launched a $99-a-month AI admin assistant for executives that manages inboxes, calendars, and outbound calls, plus a seed round. It works across email, chat, and phone, with travel booking in limited beta. 🗞 #sponsored @testingcatalog

GOOGLE 🔥: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor. Earlier today,
+1
GOOGLE 🔥: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor. Earlier today, users also spotted that Gemini 3.8 Flash has been powering some of there conversations on Gemini already. Very soon 👀

Muse superapp from Meta and Ava model with computer use Meta is preparing its agent super app, formerly Project Hatch, for launch as Muse. iOS now shows a waitlist, while desktop updates point to computer and browser control, suggesting a phased rollout for an autonomous AI product. 🗞 #meta @testingcatalog

DAILY AI BRIEF 🗞 — Sept 2 OPENAI 🔥: > Official “Path to Astra” post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is “coming soon” — advanced cyber tools stay limited to testers / Daybreak Blue at first. > M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha. GOOGLE 🔥: > Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway. > WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today. > Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later. ANTHROPIC 🔥: > Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads — about 25% cheaper typically, up to 45% on heavy agent runs. META 🔥: > Muse Voice Transcribe is live — MSL’s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available. XAI 🔥: > Elon: “Grok 4.7 comes out in 10 days.” That’s ~Sept 12. Reply to Tobi on Grok 4.6. ALIBABA 🔥: > Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 — 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max. WORLD LABS 🔥: > Fei-Fei Li’s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet. * Too much is happening, and I have some scoops planned for today, so I don't want to spam the algo. ** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.

Anthropic launches Claude Fable 5.1 and Mythos 5.1 Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds. 🗞 #anthropic @testingcatalog

GOOGLE 🔥: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ. “Jetski” has been mentioned in the article as a Goog
+1
GOOGLE 🔥: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ.
“Jetski” has been mentioned in the article as a Google’s internal coding tool too.
Soon 👀

OPENAI 🔥: Astra will be "available soon," but its cybersecurity capabilities will be limited. > Astra scored 100% on Exploit
+1
OPENAI 🔥: Astra will be "available soon," but its cybersecurity capabilities will be limited. > Astra scored 100% on ExploitBench. > OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently. > Astra achieved "much higher arbitrary code-execution rates than GPT‑5.6 Sol". > During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains. Soon 👀

ANTHROPIC 🔥: Fable 5.1 is now available on Claude and Claude Code. It requires extra usage credits while priced the same as
ANTHROPIC 🔥: Fable 5.1 is now available on Claude and Claude Code. It requires extra usage credits while priced the same as Fable 5, with 75% cheaper API cache reads. > Writes in plain language and sticks to what you asked for. > Creates finished spreadsheets and checks each number as it goes. > Shows its sources and separates what's known from what's estimated.

BREAKING 🔥: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1! It scores 52.6% on Terminal-Bench-Science 0.1, m
+1
BREAKING 🔥: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1!
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.
On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
Rolling out on Claude now 👀

ANTHROPIC 🔥: Claude Fable 5.1 is being prepared for the upcoming release! > The "Thought Preserved: Modifying the way the Me
+1
ANTHROPIC 🔥: Claude Fable 5.1 is being prepared for the upcoming release! > The "Thought Preserved: Modifying the way the Messages API handles thought blocks to protect against distillation" support page has been updated. > Both Fable 5.1 and Mythos 5.1 are expected soon. Soon? 👀

META 🔥: A new Muse Voice Transcribe model from MSL is now available on Meta models API.
SOTA in streaming speech-to-text. Trained on 70+ languages.
As it has been foretold 👀

META 🔥: Project Hatch will be released under the name “Muse” and will arrive with a waitlist! Hatch was an internal codename
+1
META 🔥: Project Hatch will be released under the name “Muse” and will arrive with a waitlist! Hatch was an internal codename of the upcoming superapp from Meta. Read more about Hatch in the post above. Joined 👀👀👀

Visko launches Orbis, a new real-time AI video model Visko launched Orbis, a real-time video model that keeps generating as prompts change mid-stream. It maintains scene memory over long runs, outputs 4K/24 fps with sub-second updates, and leads tests in video quality, motion, and prompt match. 🗞 #sponsored @testingcatalog

PERPLEXITY 🔥: A Hybrid mode for Perplexity Computer on Mac has been officially announced! > The Hybrid mode is powered by PPLX Qwen 3.8 27B, a custom post-trained model from Perplexity. > It also comes with a Privacy Gate feature to detect PII data before it is sent to the cloud. > Perplexity also open-sourced the Privacy Gate classifier on Hugging Face.

Claude app for iOS now has an explicit warning next to the Max effort option that it consumes 1.5 more usage. Max 1.5 ⚠️
+1
Claude app for iOS now has an explicit warning next to the Max effort option that it consumes 1.5 more usage. Max 1.5 ⚠️

Antigravity now supports /boost command in order to power its harness system with more tokens. Deep boost 👀
+1
Antigravity now supports /boost command in order to power its harness system with more tokens. Deep boost 👀

Grok Bot now supports apps from Microsoft ecosystem via plugins for Outlook, Calendar, and OneDrive. Expansion 🤖
Grok Bot now supports apps from Microsoft ecosystem via plugins for Outlook, Calendar, and OneDrive. Expansion 🤖