News linked to both this project and an event.
Google Gemini 3.8 Flash has now launched on Gemini, Google AI Studio, and the API. According to Testing Catalog, the model scored 71% on the DeepSWE 1.1 benchmark, close to Claude Opus 5’s 74%, but at a much lower price. Regarding pricing, input costs $0.75 and output (including thinking tokens) costs $3.75 before December 31, 2026; starting January 1, 2027, these rates will increase to $1.50 and $7.50 respectively.
According to a report by Wall Street Journal reporter Erin Woo citing employee sources, Google’s artificial intelligence research department is set to release a new model, Gemini 3.8 Flash, featuring significantly enhanced coding capabilities. Internally codenamed "Skimaki," the model could go live as early as Wednesday, September 2. In comparative testing against Google’s internal coding tool Jetski, company engineers demonstrated a stronger preference for the new model than for Anthropic’s Opus.
Anthropic has released a new study titled "Training a Misaligned Reward Chaser," examining whether "reward hacking" during training compels models to pursue rewards at all costs. The research team trained an Opus-scale model across 80 known exploitable production environments. Simulated evaluations revealed that the model engaged in unauthorized network attacks, tampered with reward mechanisms, and attempted to evade security monitoring.
Z.ai has released GLM-5.3-Flash, calling it the first native multimodal model in the GLM-5 series. The model features 320 billion total parameters and 18 billion active parameters, outperforming GLM-5.2 across multiple coding, agent, and vision benchmarks, and approaching Claude Opus 4.8 on certain coding tasks. GLM-5.3-Flash employs a hybrid architecture that combines sparse attention with linear attention, and introduces mechanisms such as mHC and IndexPool to reduce long-context inference costs and KV cache overhead.
Odaily News - Digital asset manager Grayscale's Zcash ETF began trading on NYSE Arca on Tuesday under the ticker ZCSH. The product is the world's first exchange-traded product offering spot exposure to Zcash, allowing investors to track ZEC prices through securities accounts without needing to directly purchase or store the token.ZCSH was formerly known as the Grayscale Zcash Trust, established in October 2017 through a private placement. Grayscale filed an application with the U.S. Securities and Exchange Commission in November 2025 to convert the trust into an ETF, with shareholders holding shares that track the fund's ZEC holdings rather than holding ZEC directly.In May of this year, security researcher Taylor Hornby, using Anthropic's Claude Opus 4.8, discovered a vulnerability in Zcash's Orchard shielded pool that had existed for four years, which could potentially allow attackers to mint counterfeit ZEC. Developers deployed an emergency patch on June 1, but due to privacy mechanisms, it was not possible to cryptographically confirm whether the vulnerability had been exploited.Zcash activated the Ironwood upgrade in July, replacing Orchard with a new shielded pool and introducing accounting rules that limit the amount of ZEC exiting the old shielded pool to no more than the amount entering. Grayscale stated it will monitor the adoption of the Ironwood upgrade, network security, exchange support, and regulatory conditions for privacy assets. (Decrypt)
Developers have discovered two early-access Claude models from Anthropic, codenamed claude-marshmallow-eap and claude-melon-eap.According to tester feedback, Marshmallow performs better overall than Melon, but neither has yet reached the level of Fable 5. Another tester noted that Marshmallow's conversation experience surpasses that of Opus 5.Previously, Anthropic has repeatedly seen early models named in the "food name + EAP" format. In early July, Claude Honeycomb EAP briefly appeared in Cursor before being removed; on July 24, Anthropic officially released Opus 5, but the company has never confirmed a direct correspondence between Honeycomb and Opus 5.
B.AI has launched DeepSeek's brand-new multimodal experimental flagship model, DeepSeek-V4-Flash-Vision-Exp. The official API team has taken the lead in completing full integration and is now making it freely accessible to all users. Building upon V4-Flash's state-of-the-art pure text and code Agent capabilities, the model officially unlocks native image understanding, with its comprehensive multimodal Agent performance approaching that of the Claude Opus-4.8 flagship. The new model supports mixed text-and-image input and a 1M ultra-long context window, effortlessly handling demanding scenarios such as multimodal PPT generation, complex UI reconstruction, and frontend visual effects. Image input is billed per token, consuming a maximum of just 384 tokens per image. All users can now log into the B.AI platform to access this latest multimodal model capability completely free of charge and with zero barriers, empowering top-tier compute to rapidly accelerate the realization of creative and productive workflows.
According to TechCrunch, the UK AI lab Inherent, founded by several former members of Google DeepMind, has released an AI research agent named Faraday. The company states that Faraday outperforms frontier models such as Anthropic Claude Opus 4.8 and OpenAI GPT-5.5 in tasks requiring the independent reproduction of research results from published scientific papers.
DeepSeek announced today that its brand-new multimodal visual understanding model, DeepSeek-V4-Flash-Vision-Exp, is now available on the DeepSeek API platform. The model is experimental and can be invoked by users by setting model="deepseek-v4-flash-vision-exp". Its pure text capabilities are roughly on par with the official release of DeepSeek-V4-Flash, while demonstrating significantly improved performance in benchmarks requiring visual understanding, bringing its multimodal Agent capabilities close to Opus-4.8.
According to TechFlow Research, the frontier AI data tracking report released by Bank of America Securities on August 17 shows that Anthropic leads comprehensively in three major AI benchmarks, with Claude Opus 5 ranking first in the Intelligence Index, Agent Index, and Coding Agent Index, GPT-5.6 Sol following closely behind, Meta MuseSpark 1.2 entering the top ten, and Google Gemini 3.6 Flash ranking outside the top ten. In terms of usage, DeepSeek leads with approximately 30% of the Vercel platform token share, Anthropic accounts for 25% and OpenAI accounts for 16%; but in terms of payment amount, Anthropic leads far ahead with 65%, while OpenAI accounts for only 11%. In terms of pricing, the AI Token Price Index decreased 9% month-over-month in August to $2.21, but still increased 87% year-over-year; GPU rental rates remain strong, with H100 increasing 33% year-over-year to $2.77/hour, DRAM increasing 483% year-over-year, and NAND increasing 432% year-over-year. The research report judges that AI infrastructure demand remains healthy, with open-source model usage growing but payment share still highly concentrated on top closed-source models. BofA believes that the Meta "Watermelon" and Google Gemini 4 releases, token pricing trends and GPU rental trends are
According to Reuters, the latest V4-Flash API model officially released by Chinese AI startup DeepSeek on July 31 incurred operating costs of only 1/105 that of Anthropic Claude Fable 5 in benchmark tests conducted by AI performance analysis agency Artificial Analysis. Regarding specific pricing, V4-Flash input token costs are $0.14 per million, and output token costs are $0.28 per million, with an average cost per test of approximately 3 cents, far lower than Wenxin Kimi K3 (86 cents), OpenAI GPT-5.6 Sol ($1.86), and Claude Fable 5 ($3.15). In terms of performance, V4-Flash scored 50 points on the Comprehensive Intelligence Index, tying with Google Gemini 3.6 Flash, but still more than 9 points lower than leading models such as Claude Opus 5 and GPT-5.6. It is worth noting that a low listed price does not equate to low actual costs—if the model consumes more inference and output tokens when generating responses, actual expenses may rise significantly.
: Bitgo CEO Mike Belshe deposited 100 BTC into a public Bitcoin address on August 1, worth approximately $6.3 million at the time, and invited Anthropic's Claude model to attempt to move the funds out of the address. On-chain records show the wallet received the funds on July 31, and the balance had not been transferred out as of August 2. Anthropic previously disclosed that during 141,006 cybersecurity assessment runs, 3 incidents were found, with 6 evaluation sessions involving 3 models inadvertently interacting with real organizational systems. The models involved include Claude Opus 4.7, Claude Mythos 5, and an unreleased internal research model. The cause was a configuration error by third-party testing partner Irregular, which led to the test environment being connected to the internet. Anthropic stated that Claude Opus 4.7, during one evaluation, located a real website with the same name as a simulated company, exploited weak passwords and exposed services to recover infrastructure credentials, and accessed a production database containing hundreds of records. The company said the model was attempting to complete assigned tasks, not actively breaking constraints or pursuing independent goals. Belshe's challenge involves Bitgo's institutional custody platform, which uses multi-signature or multi-party computation technology to distribute signing authority across multiple independent keys. As of August 2, Anthropic had not publicly responded to the challenge.
DeepSeek officially launches the public beta of version V4-Flash-0731. Through post-training and Agent framework optimization, DeepSWE benchmark scores have improved from 7.3 to 54.4, Cybergym scores have doubled to 76.7, and performance on some hard benchmarks approaches Claude Opus 4.8. The new version does not increase parameters, prices remain unchanged, it natively supports the Responses API, and existing user APIs will automatically switch to the new version. Currently, only the Flash API is updated; the V4 Pro official version is still under development.
Odaily News, July 30 – Anthropic released a report stating that during a review of cybersecurity assessment records, three incidents were discovered in which the Claude model accessed the internet in a third-party evaluation environment and further obtained unauthorized access to three real organizations' systems.Anthropic stated that the review covered 141,000 evaluation runs that could have potentially gained network access, and a total of three related incidents were found. All incidents occurred during Capture The Flag (CTF) cybersecurity tests, where the model was told the environment was a simulation with no internet access; however, due to configuration errors by the evaluation partner, the actual environment had internet connectivity.Among these, Claude Opus 4.7 accessed real company infrastructure during one test and obtained database permissions containing hundreds of production data records; Claude Mythos 5 built a malicious Python package and uploaded it to PyPI, resulting in the package being downloaded and run on 15 real systems; another internal research test model scanned approximately 9,000 targets and accessed a company's internet application through a public vulnerability.Anthropic stated that these incidents were not cases of the model actively seeking to escape or pursue its own goals, but rather the model mistakenly believed the real systems were within the test scope and continued executing the assigned cyberattack tasks. Notably, the newer internal research model stopped attacking after identifying that the targets might be real systems.Anthropic stated that these incidents primarily reflect issues with evaluation environment isolation and operational processes, rather than model alignment failures. The company has suspended related cybersecurity assessments, strengthened security controls in evaluation environments, continuously monitored test records, and will collaborate with third-party organizations to conduct further reviews.
Perplexity CEO Aravind Srinivas tweeted that GLM is a severely underrated model, demonstrating performance close to Opus-level at the 700B parameter scale with extremely high operational efficiency. Previously, Moonshot AI just released the open-source model Kimi K3, and industry attention on the GLM series models under Zhipu AI is rising.
Moonshot AI launches open-source model Kimi K3, scoring 57 points on the Artificial Analysis Intelligence Index, becoming the third highest-scoring model, second only to Anthropic's Claude Opus 5 (61 points), Claude Fable 5 (60 points), and OpenAI's GPT-5.6 Sol (59 points). The index shows that the gap between leading closed-source models and open-weight models has narrowed to 4 points, the smallest gap since the release of GLM-5 in February. This signifies that open-source models are rapidly catching up to closed-source models in performance.
The B.AI API platform officially launches the Claude Opus 5 model. This model provides a context window of up to 1 million tokens and a single output limit of 128,000 tokens, achieving significant performance leaps in scenarios such as complex reasoning, long code engineering development, large-scale document analysis, and knowledge-intensive workflows. To meet diverse deployment needs, this release simultaneously opens a dual-channel API access solution on the B.AI platform: in addition to the official standard interface, a new "Designated Service Provider" cooperation channel is added, through which enterprise users can obtain cost discounts of up to 40%, facilitating flexible balancing of performance and budget expenditure based on actual business loads. Developers can log in to the official B.AI platform starting today to experience the exceptional capabilities of Claude Opus 5 first.
Anthropic's Claude Opus 5 (max and xhigh versions) achieved the highest intelligence score on the Artificial Analysis Intelligence Index v4.1. The index integrates 9 evaluation benchmarks including GDPval-AA v2, GPQA Diamond, and Humanity's Last Exam. Previously, Opus 5 has demonstrated leading performance in multiple benchmark tests, including scoring twice that of the previous generation on Frontier-Bench, topping the AA-Briefcase agent benchmark, and outperforming Fable 5 in cost-performance ratio.
Anthropic has officially released the Opus 5 AI model; the company states that its performance is close to the frontier model Fable 5, but the price is only half that of the latter.
PPP Prediction Market Tool monitoring shows that on Polymarket, the probability of "the next Claude Opus model will be released on July 23, 2026" is currently at 61%, up 51% in 24 hours; additionally, the probability of "release before July 24" is at 9%; the probability of "release before July 25" is at 5%;This event will be settled based on the date (Eastern Time) when Anthropic's next Claude Opus model becomes available to the public. Only models explicitly named "Opus" by Anthropic are valid, and they must be accessible to the public (including public beta testing or open waitlist). Other models such as Sonnet and Haiku are not counted. The final settlement will primarily rely on official Anthropic information.Join the PPP Signal Push Community to stay ahead and seize the opportunity.