News linked to both this project and an event.
DeepSeek officially launches the public beta of version V4-Flash-0731. Through post-training and Agent framework optimization, DeepSWE benchmark scores have improved from 7.3 to 54.4, Cybergym scores have doubled to 76.7, and performance on some hard benchmarks approaches Claude Opus 4.8. The new version does not increase parameters, prices remain unchanged, it natively supports the Responses API, and existing user APIs will automatically switch to the new version. Currently, only the Flash API is updated; the V4 Pro official version is still under development.
According to the research summary published by Rohan Paul, a large-scale study covering 207 GitHub projects and 1.02 million pull requests shows that introducing AI agent review can reduce code review time by 2.5 to 4.5 days/KLOC, but at the cost of declining review quality—in reviews involving AI, 78%~94% of PRs exhibit "review smells", higher than the 69%~76% in pure human reviews. The study points out that repeatedly assigning the same AI reviewer identity is the main reason leading to the decline in review diversity. Notably, projects that introduced LLM review extensively in the early stage did not achieve significant efficiency improvements.
Anthropic engineers disclosed that the team reduced approximately 80% of the system prompts for Claude Code for the latest Claude model, and coding evaluations showed no performance degradation. Early versions contained numerous strict rules in the prompts (such as prohibiting multi-line comments, etc.) due to insufficient model capabilities, but after the new model acquired autonomous decision-making capabilities, these rules instead triggered instruction conflicts and divergent thinking. The new strategy shifted to describing goals (such as "write code that matches the surrounding code") and distributed instructions across tool descriptions and on-demand enabled skill modules. The team explicitly views prompt bloat as technical debt, believing that as model capabilities improve, instructions should be simplified rather than stacked.
According to Decrypt, Jack Dorsey's payment company Block has officially launched the open-source collaboration platform Buzz, built on the Nostr protocol, enabling employees and AI agents to collaborate within the same workspace. Buzz integrates team chat, code repositories, and automated workflows, where each user and AI agent possesses an independent encrypted identity, supporting mainstream model frameworks such as Claude Code, OpenAI Codex, and Block's self-developed goose, with the code released open-source under the Apache 2.0 license. Jack Dorsey stated that the platform aims to reduce the company's reliance on Slack and GitHub; the current version is 0.4.21, already supporting macOS, Windows, and Linux desktops, while the mobile version is still under development.
According to an official announcement by the FCA, Anthropic will provide access to its suite of Claude products (including Claude Code and Claude Cowork) for the second batch of participating enterprises in the UK Financial Conduct Authority (FCA) "Super Sandbox" to accelerate their AI development process. A total of 21 institutions were selected for the second batch, including Scottish Widows, Money Advice Trust, and TrueLayer, among others. The number of applications this time increased by 51% compared to the first batch, with a total of 199 applications received. Participating enterprises will test AI solutions focusing on the following areas: agent payment security, fraud and financial crime detection, AI governance and accountability, financial inclusion for vulnerable groups, and compliance and business automation. Additionally, the FCA simultaneously launched the "Agentic Academy"—a 10-week AI specialized training program co-hosted by the FCA and the Centre for Finance, Technology and Entrepreneurship (CFTE). Participating enterprises will continue to have access to NayaOne digital sandbox infrastructure and NVIDIA accelerated computing resource support.
Kimi posted on platform X, stating that the market response to K3 has far exceeded expectations since its launch, with user demand in the past 48 hours approaching the upper limit of existing GPU computing capacity. To ensure the user experience for existing subscribers, new subscription services have been temporarily suspended. Current subscribers will not be affected, and the company is working to expand computing resources as quickly as possible, gradually resuming new user subscriptions in batches.Additionally, the future membership system will be split into two more targeted products: "Kimi Membership," designed for Kimi's web interface, app, and office scenarios, and "Kimi Code Membership," focused on programming workflows, to achieve more precise resource allocation and improve service stability.
Cos, founder of SlowMist, shared a tweet on X platform regarding potential poisoning attack risks in Claude Code and published a detailed analysis of poisoning attacks targeting Grok Build CLI and Claude Code CLI. The analysis pointed out that the security mechanisms of Grok Build CLI are not unified, with different code paths having different trust assumptions, and the gaps between them serve as channels for attackers.Attackers may exploit malicious project configuration files to execute arbitrary commands without the user's knowledge, thereby stealing API keys, cloud credentials, or gaining control over local devices. Researchers constructed a test environment and found that on Mac systems, if Claude Code is compromised, executing a specific test command could trigger the launch of a local calculator, demonstrating a potential command execution risk. If the attack succeeds, attackers could further steal API keys from AI services such as Claude and OpenAI, causing account cost losses; obtain credentials for cloud services like AWS, Alibaba Cloud, and Tencent Cloud to access servers and data; tamper with code repositories to implant backdoors; and leverage local devices as a springboard to attack internal enterprise networks. It is reported that the relevant vulnerability has existed for one year.
: Yesterday, Dark Side of the Moon (Moonshot AI) released its latest open-source AI model, Kimi K3. It ranked first on the Frontend Code Arena test website with a score of 1,679, surpassing the Claude Fable 5 model. Following an evaluation of the K3 model by Artifacial Analysis, Elon Musk once again praised the Kimi model from Dark Side of the Moon, stating that the K3 model's benchmark performance is impressive.In March of this year, when Kimi published the research paper "Attention Residuals: Rethinking the Aggregation of Depth Direction," it received praise from Musk, who said, "Kimi's research work is impressive." Previously, he also stated that the Zhipu GLM model could surpass the Claude Mythos model (i.e., Fable 5) by Q1 2027. In response, Zhipu founder Tang Jie replied, "It won't take that long."
Today, X-Agent, a zero-code operating platform for AI Agents built on a Web3 social network, officially released its core to the global audience. The whitepaper details how it leverages zero-code deployment and secure sandbox technology to reshape the decentralized AI agent ecosystem and its commercialization pathway.The core technologies and solutions are as follows:1) Speak to Build (Zero-Code Deployment): Allows users to create and deploy professional-grade AI agents with one click to Telegram, WhatsApp, or the Web in just minutes using pure natural language.2) SRE Cryptographic Sandbox (Physical Isolation Security): A proprietary physically isolated operating environment that perfectly protects users' private keys and sensitive credentials while enabling agents to possess enterprise-grade capabilities for autonomous asset management and on-chain collaboration.Market Performance and Latest Data: X-Agent has surpassed 1,000,000+ global cumulative users. Its deployed AI agents have autonomously completed over 1,100,000 actual business tasks, consuming/burning more than 84 billion model Tokens. The project has previously raised $1.8 million in funding and boasts highly cohesive localized communities in Japan and South Korea.The native token, $XAGT, serves as the hard currency for network settlement, directly used for paying SRE computing power Gas fees, transaction commissions, and staking endorsements. The team and investors are subject to a strict 12-month Cliff lock-up period. Community allocations are entirely non-inflationary and can only be released algorithmically through actual sandbox task consumption, ensuring every token is backed by real business demand.According to the latest roadmap, X-Agent plans to open a new zero-code public portal in Q3 2026 and launch a decentralized Agent Store in Q1 2027.
Odaily News: BNB treasury-listed company BNB Plus has announced that it has received a delisting ruling from the Nasdaq Hearings Panel. The primary reason for the delisting is that the company's stock price has persistently failed to meet the Nasdaq Capital Market's continued listing requirement of a minimum $1 bid price. BNB Plus plans to submit a review request to the Nasdaq Listing Review Committee, but this review will not stay the delisting process. The company's stock will be suspended for trading on Nasdaq at the opening on July 14, 2026. The company has completed preparatory work for listing on the OTCQB Venture Market, and the stock code will remain BNBX. Normal trading on the over-the-counter market is expected to commence at the opening on July 14. (Businesswire)
据 Claude 官方 X 账号(@claudeai)发布,Claude 将在所有付费计划中延长 Fable 5 的访问权限,同时 Claude Code 每周使用限制维持高出 50% 的水平,上述政策延续至 7 月 19 日。官方补充说明,用户每周使用限额的一半可用于 Fable 5,超出后可通过使用积分继续访问,或切换至其他模型继续使用。
According to reports by The Paper, due to security risks involving implanted backdoors exposed in Claude Code, Alibaba has listed it on its high-risk software list following a comprehensive evaluation. Effective July 10, internal employees are completely prohibited from using it in the office environment, and Qoder is recommended as an alternative. It is reported that Anthropic's Claude Code was found to have embedded "invisible code" since version 2.1.91. If users enable network proxies, the tool will secretly transmit user location and other information by modifying invisible system prompts. In response, Anthropic stated that this function was an experimental measure launched in March this year, aimed at preventing unauthorized resellers from abusing accounts and preventing distillation behavior, and indicated it would be fully rolled back in the latest version.
BNB Chain has announced the official mainnet launch of its AI Agent development platform, BNB Agent Studio.Developers can now use a single prompt in AI coding tools like Cursor and Claude Code to complete Agent wallet creation, on-chain identity registration (ERC-8004), and deployment, without needing to separately set up wallets, identities, payments, custody, or LLM integration.Once deployed, Agents can use the x402 protocol to automatically deduct fees from users' pre-funded wallets to cover LLM usage, and they can be discovered and invoked by other Agents via the ERC-8183 task interface. The entire process runs on the AWS Bedrock AgentCore.The platform is also launching a limited-time free trial, where users can experience the full deployment process on the BSC testnet using their GitHub account.
Regarding the exposure that Claude Code contains detection code targeted at Chinese users, Claude Code team member Thariq Shihipar responded that the relevant feature is an experiment launched in March this year, primarily used to identify unauthorized distribution activities and prevent model distillation. According to the explanation, the relevant code detects time zones, proxies, and potential AI laboratory information.
Odaily Odaily News According to official sources, OKX has officially launched OKX.AI. This is a decentralized platform designed for the Agent economy, supporting AI Agents in issuing tasks, undertaking tasks, completing payments, providing evaluations, and conducting arbitration. This enables Agents to collaborate autonomously like economic entities and complete complex tasks.It is reported that OKX.AI consists of two major markets: the "Agent Square" and the "Task Hall". The underlying layer relies on infrastructure such as Onchain OS to provide the identity, payment, trust, and settlement capabilities required by the agent economy. Each Agent has a unique on-chain identity and can be paid based on task outcomes; disputes are arbitrated by a staked validator network. Developers can connect through the open toolkit Onchain OS, which is compatible with MCP-supported clients such as Claude Code and Codex, and can start using it by creating an Agentic Wallet.
: Zhipu AI founder Tang Jie posted on the X platform, stating that since the official open-source release of GLM-5.2, it has achieved leading results in multiple international authoritative evaluations and competitive rankings. In the Artificial Analysis Intelligence Index comprehensive evaluation, GLM-5.2 scored 51 points, placing it in the same range as Anthropic's Claude Opus 4.8. In the Code Arena front-end code generation adversarial test, it ranked 2nd globally with an Elo of 1595, and in the DesignArena design and code integration scenario, it scored 1360 points, ranking 1st.Overall, Zhipu GLM-5.2 continues to rank among the top globally in real-world scenario evaluations across multiple areas, including front-end development, design generation, and software engineering. It is steadily narrowing the performance gap with cutting-edge models from OpenAI and Anthropic and will continue to push the upper limits of model capabilities.Previously, in response to Musk's statement that Chinese large models might reach Anthropic's Fable level by the first quarter of next year, Zhipu AI founder Tang Jie replied, "It won't take that long."
Claude Code has officially launched the Artifacts feature, allowing users to preview ongoing work as a real-time, interactive webpage, and generate and update content based on the full session context, facilitating team sharing and collaboration, including PR reviews, system documentation, dashboards, release checklists, etc. The content will be automatically updated as the session progresses. The feature is now available in beta on Claude Team and Enterprise organizations, and supports access via a browser.
OpenAI has officially announced its commitment to advancing the development of AI content provenance verification standards and supporting the European Commission’s “Code of Practice on Transparency of AI-Generated Content,” thereby enhancing traceability and transparency of AI-generated content. According to reports, since 2024, OpenAI has integrated C2PA metadata into its DALL·E 3 image generation tool to identify the origin of AI-generated content. Since then, OpenAI has continuously refined its content labeling and detection technologies and launched public verification tools to help users determine whether an image contains provenance signals associated with OpenAI-generated content.
OpenAI has released the Frontier Governance Framework, systematically elaborating on how its AI safety and governance practices align with emerging regulatory requirements such as the California Frontier AI Transparency Act and the EU's General-Purpose AI Code of Conduct. Based on OpenAI's existing Preparedness Framework, this framework focuses on areas including cyberattacks, CBRN risks, harmful manipulation, loss of control risks, model reporting, security incident response, and external expert review. It also states that it will be continuously updated as model capabilities and the regulatory environment evolve.
Odaily Planet Daily reported that Bitcoin News posted on X platform, stating that Senator Lummis said that if the Clarity Act is not passed during this Congress, US software developers will again become targets of lawsuits in the near future simply for releasing code. That's what's at stake.