en
Feedback
🚨 AI News | TestingCatalog

🚨 AI News | TestingCatalog

Open in Telegram

Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors 🗞

Show more
The country is not specifiedTechnologies & Applications14 411
7 472
Subscribers
+1224 hours
+717 days
+29130 days
Posts Archive
DAILY AI BRIEF 🗞 — Sept 3 GOOGLE 🔥: > Gemini 3.8 Flash is live — third Flash in six weeks. Workhorse for coding and agents, 1M context. Intro price $0.75 / $3.75 through Dec 31. > Flash Cyber is the defender twin (vuln find + auto-patch). Gated to Fairwind: governments, infra, trusted maintainers. META 🔥: > Muse Spark 1.3 is out in Muse Code and the Meta Model API. Zuck: biggest jump yet on coding and agents, “almost too cheap to meter.” > Beat GPT-5.6 and Opus 5 on DeepSWE 1.1. Watermelon 🍉 and Spark open weights are next. OPENAI 🔥: > gpt-6-astra slug is on the APIs. Employees and testers are posting about a possible drop today. Still not public. > Already tagged Critical for cyber under the Preparedness Framework — extra safeguards, limited tools first. XAI 🔥: > Grok Bot is live on Android. Elon: “Grok @Bot now on Android.” Play Store, same “give it real work” agent as iOS. * Didn't plan to post it initially cuz all the news were covered already but here you go :P

Google releases Gemini 3.8 Flash and Flash Cyber Google launched Gemini 3.8 Flash for coding, agents, and multi-step reasoning, plus 3.8 Flash Cyber for vulnerability discovery and patching. Flash keeps 3.7 pricing, while Cyber access is limited to trusted defense users. 🗞 #google @testingcatalog

ICYMI: OpenAI's Astra crosses Critical cybersecurity threshold OpenAI says Astra has reached its Critical cyber threshold, showing top exploit and zero-day capability on hardened systems. Release is planned soon, with strongest access limited to alpha testers and Daybreak Blue for defense use. 🗞 #openai @testingcatalog

Astralogy 🔮 OpenAI employees are teasing an upcoming release of the Astra model.
+1
Astralogy 🔮
OpenAI employees are teasing an upcoming release of the Astra model.

META 🔥: Muse Spark 1.3 has been officially announced, and it scored above GPT-5.6 and Opus 5 on DeepSWE 1.1! > Muse Spark 1.
+1
META 🔥: Muse Spark 1.3 has been officially announced, and it scored above GPT-5.6 and Opus 5 on DeepSWE 1.1! > Muse Spark 1.3 is rolling out on Meta model APIs. > "Watermelon", the next big model upgrade from Meta, and the Muse Spark open-weight version are coming soon! The competition is getting hotter 👀

OPENAI 🔥: GPT-6-Astra model slug has been spotted on the APIs. If we will actually get it tomorrow, it would be a huge week.
+1
OPENAI 🔥: GPT-6-Astra model slug has been spotted on the APIs. If we will actually get it tomorrow, it would be a huge week. Routing first 👀

SPACEXAI 🔥: Grok Bot is now available on Android platform! Bot testing time 👀

GOOGLE 🔥: Gemini 3.8 Flash is rolling out on Gemini, Google AI Studio and APIs. Gemini 3.8 Flash scores 71% on DeepSWE 1.1,
GOOGLE 🔥: Gemini 3.8 Flash is rolling out on Gemini, Google AI Studio and APIs. Gemini 3.8 Flash scores 71% on DeepSWE 1.1, compared to 74% for Claude Opus 5, at a much lower price. > Input price $0.75 through December 31, 2026. $1.50 starting January 1, 2027. > Output price (including thinking tokens) $3.75 through December 31, 2026. $7.50 starting January 1, 2027. This is big 👀

Wonderful raises $550M for its enterprise AI agent platform Wonderful raised $550M in Series C funding at a $5B valuation to expand AI agents for enterprise work. Backed by Insight Partners and Salesforce, it pairs software with embedded engineering teams to move deployments into production fast. 🗞 #sponsored @testingcatalog

GOOGLE 🔥: Gemini 3.8 Flash is already available in Agent Studio on GCP. Best for - Complex multimodal data processing - Codi
GOOGLE 🔥: Gemini 3.8 Flash is already available in Agent Studio on GCP. Best for - Complex multimodal data processing - Coding use cases - Supporting software engineering–related agentic tasks Use case - Processing data with images and text - Coding problems - Web research and application testing

Catch launches an AI assistant with phone calling support Catch launched a $99-a-month AI admin assistant for executives that manages inboxes, calendars, and outbound calls, plus a seed round. It works across email, chat, and phone, with travel booking in limited beta. 🗞 #sponsored @testingcatalog

GOOGLE 🔥: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor. Earlier today,
+1
GOOGLE 🔥: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor. Earlier today, users also spotted that Gemini 3.8 Flash has been powering some of there conversations on Gemini already. Very soon 👀

Muse superapp from Meta and Ava model with computer use Meta is preparing its agent super app, formerly Project Hatch, for launch as Muse. iOS now shows a waitlist, while desktop updates point to computer and browser control, suggesting a phased rollout for an autonomous AI product. 🗞 #meta @testingcatalog

DAILY AI BRIEF 🗞 — Sept 2 OPENAI 🔥: > Official “Path to Astra” post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is “coming soon” — advanced cyber tools stay limited to testers / Daybreak Blue at first. > M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha. GOOGLE 🔥: > Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway. > WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today. > Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later. ANTHROPIC 🔥: > Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads — about 25% cheaper typically, up to 45% on heavy agent runs. META 🔥: > Muse Voice Transcribe is live — MSL’s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available. XAI 🔥: > Elon: “Grok 4.7 comes out in 10 days.” That’s ~Sept 12. Reply to Tobi on Grok 4.6. ALIBABA 🔥: > Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 — 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max. WORLD LABS 🔥: > Fei-Fei Li’s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet. * Too much is happening, and I have some scoops planned for today, so I don't want to spam the algo. ** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.

Anthropic launches Claude Fable 5.1 and Mythos 5.1 Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds. 🗞 #anthropic @testingcatalog

GOOGLE 🔥: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ. “Jetski” has been mentioned in the article as a Goog
+1
GOOGLE 🔥: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ.
“Jetski” has been mentioned in the article as a Google’s internal coding tool too.
Soon 👀

OPENAI 🔥: Astra will be "available soon," but its cybersecurity capabilities will be limited. > Astra scored 100% on Exploit
+1
OPENAI 🔥: Astra will be "available soon," but its cybersecurity capabilities will be limited. > Astra scored 100% on ExploitBench. > OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently. > Astra achieved "much higher arbitrary code-execution rates than GPT‑5.6 Sol". > During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains. Soon 👀

ANTHROPIC 🔥: Fable 5.1 is now available on Claude and Claude Code. It requires extra usage credits while priced the same as
ANTHROPIC 🔥: Fable 5.1 is now available on Claude and Claude Code. It requires extra usage credits while priced the same as Fable 5, with 75% cheaper API cache reads. > Writes in plain language and sticks to what you asked for. > Creates finished spreadsheets and checks each number as it goes. > Shows its sources and separates what's known from what's estimated.

BREAKING 🔥: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1! It scores 52.6% on Terminal-Bench-Science 0.1, m
+1
BREAKING 🔥: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1!
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.
On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
Rolling out on Claude now 👀

ANTHROPIC 🔥: Claude Fable 5.1 is being prepared for the upcoming release! > The "Thought Preserved: Modifying the way the Me
+1
ANTHROPIC 🔥: Claude Fable 5.1 is being prepared for the upcoming release! > The "Thought Preserved: Modifying the way the Messages API handles thought blocks to protect against distillation" support page has been updated. > Both Fable 5.1 and Mythos 5.1 are expected soon. Soon? 👀