🚨 AI News | TestingCatalog
Відкрити в Telegram
Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors 🗞
Показати більшеКраїна не вказанаТехнології та додатки14 474
7 462
Підписники
+924 години
+667 днів
+28430 днів
Архів дописів
GOOGLE 🔥: Gemini 3.8 Flash is rolling out on Gemini, Google AI Studio and APIs.
Gemini 3.8 Flash scores 71% on DeepSWE 1.1, compared to 74% for Claude Opus 5, at a much lower price.
> Input price
$0.75 through December 31, 2026.
$1.50 starting January 1, 2027.
> Output price (including thinking tokens)
$3.75 through December 31, 2026.
$7.50 starting January 1, 2027.
This is big 👀
Wonderful raises $550M for its enterprise AI agent platform
Wonderful raised $550M in Series C funding at a $5B valuation to expand AI agents for enterprise work. Backed by Insight Partners and Salesforce, it pairs software with embedded engineering teams to move deployments into production fast.
🗞 #sponsored @testingcatalog
GOOGLE 🔥: Gemini 3.8 Flash is already available in Agent Studio on GCP.
Best for
- Complex multimodal data processing
- Coding use cases
- Supporting software engineering–related agentic tasks
Use case
- Processing data with images and text
- Coding problems
- Web research and application testing
Catch launches an AI assistant with phone calling support
Catch launched a $99-a-month AI admin assistant for executives that manages inboxes, calendars, and outbound calls, plus a seed round. It works across email, chat, and phone, with travel booking in limited beta.
🗞 #sponsored @testingcatalog
+1
GOOGLE 🔥: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor.
Earlier today, users also spotted that Gemini 3.8 Flash has been powering some of there conversations on Gemini already.
Very soon 👀
Muse superapp from Meta and Ava model with computer use
Meta is preparing its agent super app, formerly Project Hatch, for launch as Muse. iOS now shows a waitlist, while desktop updates point to computer and browser control, suggesting a phased rollout for an autonomous AI product.
🗞 #meta @testingcatalog
DAILY AI BRIEF 🗞 — Sept 2
OPENAI 🔥:
> Official “Path to Astra” post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is “coming soon” — advanced cyber tools stay limited to testers / Daybreak Blue at first.
> M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha.
GOOGLE 🔥:
> Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway.
> WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today.
> Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later.
ANTHROPIC 🔥:
> Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads — about 25% cheaper typically, up to 45% on heavy agent runs.
META 🔥:
> Muse Voice Transcribe is live — MSL’s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available.
XAI 🔥:
> Elon: “Grok 4.7 comes out in 10 days.” That’s ~Sept 12. Reply to Tobi on Grok 4.6.
ALIBABA 🔥:
> Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 — 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max.
WORLD LABS 🔥:
> Fei-Fei Li’s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet.
* Too much is happening, and I have some scoops planned for today, so I don't want to spam the algo.
** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
Anthropic launches Claude Fable 5.1 and Mythos 5.1
Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds.
🗞 #anthropic @testingcatalog
+1
GOOGLE 🔥: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ.
“Jetski” has been mentioned in the article as a Google’s internal coding tool too.Soon 👀
+1
OPENAI 🔥: Astra will be "available soon," but its cybersecurity capabilities will be limited.
> Astra scored 100% on ExploitBench.
> OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently.
> Astra achieved "much higher arbitrary code-execution rates than GPT‑5.6 Sol".
> During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains.
Soon 👀
ANTHROPIC 🔥: Fable 5.1 is now available on Claude and Claude Code.
It requires extra usage credits while priced the same as Fable 5, with 75% cheaper API cache reads.
> Writes in plain language and sticks to what you asked for.
> Creates finished spreadsheets and checks each number as it goes.
> Shows its sources and separates what's known from what's estimated.
+1
BREAKING 🔥: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1!
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.
On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.Rolling out on Claude now 👀
+1
ANTHROPIC 🔥: Claude Fable 5.1 is being prepared for the upcoming release!
> The "Thought Preserved: Modifying the way the Messages API handles thought blocks to protect against distillation" support page has been updated.
> Both Fable 5.1 and Mythos 5.1 are expected soon.
Soon? 👀
META 🔥: A new Muse Voice Transcribe model from MSL is now available on Meta models API.
SOTA in streaming speech-to-text. Trained on 70+ languages.As it has been foretold 👀
+1
META 🔥: Project Hatch will be released under the name “Muse” and will arrive with a waitlist!
Hatch was an internal codename of the upcoming superapp from Meta. Read more about Hatch in the post above.
Joined 👀👀👀
Visko launches Orbis, a new real-time AI video model
Visko launched Orbis, a real-time video model that keeps generating as prompts change mid-stream. It maintains scene memory over long runs, outputs 4K/24 fps with sub-second updates, and leads tests in video quality, motion, and prompt match.
🗞 #sponsored @testingcatalog
PERPLEXITY 🔥: A Hybrid mode for Perplexity Computer on Mac has been officially announced!
> The Hybrid mode is powered by PPLX Qwen 3.8 27B, a custom post-trained model from Perplexity.
> It also comes with a Privacy Gate feature to detect PII data before it is sent to the cloud.
> Perplexity also open-sourced the Privacy Gate classifier on Hugging Face.
+1
Claude app for iOS now has an explicit warning next to the Max effort option that it consumes 1.5 more usage.
Max 1.5 ⚠️
+1
Antigravity now supports /boost command in order to power its harness system with more tokens.
Deep boost 👀
Grok Bot now supports apps from Microsoft ecosystem via plugins for Outlook, Calendar, and OneDrive.
Expansion 🤖
