News linked to both this project and an event.
B.AI announces that a new round of benefits for its popular models is now officially in effect: exclusive 10% discounted API call privileges for high-concurrency flagship models DeepSeek-V4.1-Flash and GLM-5.3-Flash are now fully available, allowing users to invoke them at just 10% of the original price to meet high-throughput scenario demands with exceptional cost-efficiency. Simultaneously, Qwen3.8-Flash, Hy3, and MiMo-V2.5 continue to offer fully free API calls, providing developers with richer zero-cost options. This benefit adjustment delivers maximum value for high-throughput needs while leveraging a multi-model matrix to lower the exploration threshold for developers. B.AI continues to empower developers to build low-cost AI infrastructure through a diverse model lineup and inclusive pricing, eliminating prohibitive compute expenses and accelerating the rapid deployment of innovations.
Zhipu's official GLM-5.3-Flash will undergo a pricing update shortly. To provide global developers with a more ample preparation window, B.AI has announced a limited-time extension of its free access—until September 12 at 09:59 (SGT), GLM-5.3-Flash will remain completely free on the B.AI platform. Leading in platform usage volume, GLM-5.3-Flash features 320B total parameters and 18B activated parameters, supports a 1M ultra-long context window, and seamlessly combines rapid response with powerful reasoning capabilities, making it a highly cost-effective choice for high-frequency coding, massive data processing, and complex Agent workflows. Additionally, other popular models such as Qwen3.8 Flash, Hy3, and MiMo V2.5 continue to be 100% free on the B.AI platform. B.AI remains committed to supporting global developers with inclusive computing resources, ensuring that frontier model capabilities are truly within everyone's reach.
Odaily News: OpenAI CFO Sarah Friar stated that the company is accelerating the expansion of AI into specialized fields such as chip design, life sciences, and financial services, and is experimenting with pricing based on business outcomes rather than usage. OpenAI's enterprise business revenue grew 32% from June to July this year, while the company's overall annualized revenue increased approximately 20% during the same period. As of mid-year, revenue from enterprise and consumer businesses had split roughly evenly.Additionally, OpenAI recently cut the price of its low-cost Luna model by 80%, which was followed by an approximate 10-fold increase in usage. Friar noted that in cloud deployment scenarios, Luna's cost is even lower than Z.ai's GLM 5.3. OpenAI's coding tool Codex has now reached 25 million users. The company is also leveraging its own AI models to assist in developing the Jalapeno chip and completed chip design tape-out within nine months. (Reuters)
B.AI has announced that after Zhipu's official limited-time 50% discount expires at 24:00 on September 9, GLM-5.3-Flash will continue to provide zero-cost API calls for global developers, requiring no changes to usage habits due to upstream price adjustments.
The GLM-5.3-Flash model has become the most frequently invoked and popular model on the B.AI platform, with cumulative token throughput exceeding 2.41 trillion. As the first native multimodal model in the GLM-5 series, GLM-5.3-Flash features 320B total parameters and 18B active parameters. It employs a hybrid architecture combining sparse and linear attention mechanisms, supports 1M ultra-long context windows, and balances rapid response, powerful reasoning capabilities, and high cost-effectiveness. Starting today, developers can still invoke this model for free via the B.AI platform, covering diverse scenarios such as high-frequency APIs, coding, complex Agents, and ultra-long document processing. Try it now: chat.b.ai/chat
The AI Agent infrastructure platform B.AI announced that since the launch of its free campaign, the platform's cumulative token throughput has exceeded 10.9 trillion, with total API calls reaching 89.56 million, attracting over 239,000 new registered users (including 235,000+ API developers). The platform infrastructure has consistently demonstrated its stability and capacity under the stress of massive high-concurrency scenarios and intensive Agent workflows. In terms of the model lineup, GLM-5.3-Flash (Ox Alpha), Qwen3.8-Flash, Tencent Hy3, and Xiaomi MiMo-V2.5 remain fully free, while DeepSeek-V4-Flash and Vision-Exp versions are now available at a 50% discount, striking a balance between zero-threshold access and exceptional cost efficiency. Going forward, B.AI will continue to deliver more efficient, reliable, and cost-effective AI compute services to developers and enterprise teams, accelerating the real-world deployment of AI applications. Visit chat.b.ai/chat to start deploying your efficient workflows today.
Odaily News, August 31 — B.AI platform's full free access campaign for cutting-edge large models continues to operate at high intensity, with the platform's cumulative token throughput officially surpassing the historic milestone of 5.74 trillion (5T+). Under the rigorous demands of massive concurrent requests and heavy Agent tasks, B.AI's industrial-grade infrastructure has demonstrated exceptionally stable carrying capacity. Currently, the platform has assembled six top-tier models, including GLM-5.3-Flash (Ox Alpha), Qwen3.8-Flash, DeepSeek-V4-Flash, DeepSeek-V4-Flash-Vision-Exp, Tencent Hy3, and Xiaomi MiMo-V2.5, all available to users with zero barriers and unlimited free access, covering diverse scenarios such as code generation, ultra-long text processing, multimodal visual understanding, and Agent workflow deployment.From now on, users who log in to the B.AI platform can access the full suite of cutting-edge models at zero cost, with seamless integration of unlimited computing power, allowing every developer and AI enthusiast to truly enjoy "computing freedom."
The all-in-one AGI infrastructure platform B.AI has announced that Zhipu's GLM-5.3-Flash (320B-A18B) has officially launched on the platform and is now fully accessible at zero cost through the official API group. As the first native all-modal model in the GLM-5 series, GLM-5.3-Flash features 320B total parameters and 18B active parameters, powered by purely domestic AI chips. It introduces a hybrid architecture of sparse and linear attention for the first time, supports a 1M ultra-long context window, delivers formidable performance alongside exceptional cost-efficiency, and stands as a leading powerhouse among current domestic models. Effective immediately, developers can invoke this advanced tool at zero cost via the B.AI API official group, effortlessly tackling complex reasoning, long-text processing, and multimodal tasks. B.AI continues to lower the barrier to AI development through democratized computing power, enabling more developers to access top-tier domestic models.
Zhipu announced the launch and open-source release of the GLM-5.3-Flash model. The year-over-year rate of the US July PCE price index reached 3.7%, exceeding expectations, while expectations for Federal Reserve rate hikes intensified; Nvidia disclosed 70% growth guidance for fiscal year 2028, boosting its after-hours stock price.
Z.ai has released GLM-5.3-Flash, calling it the first native multimodal model in the GLM-5 series. The model features 320 billion total parameters and 18 billion active parameters, outperforming GLM-5.2 across multiple coding, agent, and vision benchmarks, and approaching Claude Opus 4.8 on certain coding tasks. GLM-5.3-Flash employs a hybrid architecture that combines sparse attention with linear attention, and introduces mechanisms such as mHC and IndexPool to reduce long-context inference costs and KV cache overhead.
According to Bloomberg, China-based AI company Z.AI (Zhipu) confirmed on Wednesday that the mysterious AI model Ox Alpha, which has recently risen rapidly to the top of online usage charts, is a new iteration of its GLM series. Responding to a Bloomberg news inquiry, the company stated it would release the model weights for Ox Alpha tonight.
Z.ai releases GLM-5.3, based on the same foundation model as GLM-5.2, achieving capability improvements through expanded post-training. According to the official announcement, GLM-5.3 improves by 50% over GLM-5.2 on the internal Z.ai Code Bench coding benchmark, and reaches a leading level among open models in public benchmarks such as Terminal Bench 3.0 and Agents' Last Exam. In terms of cybersecurity, GLM-5.3 achieved a score of 84.5% in the CyberGym vulnerability discovery test, and significantly improved compared to the previous generation in exploit chain-related tests such as ExploitBench and ExploitGym.
Z.ai 发文预告 GLM-5.3 即将推出,该模型聚焦编程能力,并面向网络防御挑战。官方称,GLM-5.3 通过对 743B 基础模型进行后训练,实现了顶尖的编程和智能代理能力,并在网络安全领域取得重要进展,为开源模型树立新标准。
The B.AI platform is now launching a limited-time 40% discount on the GLM-5.2 model for all users. During the event, all calls enjoy a 40% discount on the settlement price, whether deployed via API for engineering purposes or used for daily interaction on the web interface. The discounted prices are: Input 0.84, Cache Write 0.84, Cache Read 0.168, Output 2.64 (Unit: Credits/Token). As a new generation open-source flagship model, GLM-5.2 features a 1M ultra-long context window and is designed for high-performance scenarios such as large-scale code development, complex reasoning, and agent tasks. Users can access it directly via the official API standard channel or the web interface, enabling seamless integration across all-scenario workflows. This event aims to unleash top-tier AI productivity at a lower cost. Starting today, users can log in to the B.AI platform to experience it.
Odaily Odaily News: AI personal assistant startup Pally has announced the completion of a $5.2 million funding round, led by Cyber Fund and Y Combinator, with participation from Pioneer Fund, Founders Inc, 468 Capital, and angel investors. The company is valued at $30 million and primarily provides AI personal assistant services through an SMS interface. Pally was launched in 2025, with its latest version released in June, and can sync with platforms such as Gmail, Outlook, Slack, Granola, and Notion. After user authorization, Pally can book flights, manage inboxes, reply to messages, and reserve restaurant tables. Pally uses Anthropic's Claude, OpenAI's ChatGPT, as well as open-source models such as Kimi K3 and GLM 5.2. Co-founder Haz Hubble stated that the company will not sell user data or use related data to train models. (Business Insider)
Odaily News AMD, the semiconductor giant, announced the launch of its enterprise-grade AI programming platform, AMD Instinct Coder. The platform combines AMD chips, Supermicro servers, and Spectro Cloud software, aiming to help enterprises deploy AI coding assistants locally, reduce the cost of cloud-based AI models, and protect code and data security.AMD stated that Instinct Coder is an "out-of-the-box" end-to-end AI development platform, integrating AMD EPYC processors, AMD Instinct GPUs, Supermicro AI servers, Spectro Cloud PaletteAI Inference Launchpad software, and the AMD-optimized GLM-5.2 model. It can be used for software development scenarios such as code generation, application modernization, automated testing, and code review.AMD said that compared to relying on cutting-edge cloud-based AI models, Instinct Coder can help enterprises reduce total cost of ownership (TCO) by up to 70%, with the fastest payback period shortened to 6 months.AMD noted that more and more enterprises are looking to leverage AI to improve development efficiency, but face two major challenges: on one hand, the cost of invoking top-tier cloud models continues to rise; on the other hand, entrusting enterprise source code, intellectual property, and sensitive data to third-party services poses security and compliance risks.Through a local deployment model, Instinct Coder allows enterprises to maintain control over their data and code while providing more predictable infrastructure costs. The platform supports development tools such as Claude Code, OpenAI Codex, Visual Studio Code, and Cursor, with each node supporting up to 50 users (30 concurrent users).Additionally, the PaletteAI Inference Launchpad provided by Spectro Cloud enables AI workload management, model routing, request auditing, and cost monitoring, and supports invoking external models such as Anthropic, OpenAI, Google, or xAI when needed.AMD stated that Instinct Coder aims to help enterprises break free from the high costs of cloud-based AI services, accelerate AI-driven software development processes while ensuring data security and autonomous control.
According to Fortune, as Chinese open-source AI models such as DeepSeek V4 Flash, GLM-5.2, and Kimi K3 continue to make breakthroughs, the core conflict of the global AI competition is shifting from "US-China confrontation" to a contest of "open source vs. closed." DeepSeek V4 Flash's performance lags behind GPT-5.6 Luna by only one intelligence index point, but the cost per task remains 60% lower even after OpenAI's 80% price cut. US export controls on China were originally intended to restrict China's AI development, but instead compelled Chinese enterprises to innovate deeply at the algorithm architecture level, accelerating the rise of the open-source model. Currently, trends in the US tech industry are shifting; former "AI Czar" David Sacks and others publicly support the open-source route, and Anthropic has also softened its stance against open source.
Perplexity CEO Aravind Srinivas tweeted that GLM is a severely underrated model, demonstrating performance close to Opus-level at the 700B parameter scale with extremely high operational efficiency. Previously, Moonshot AI just released the open-source model Kimi K3, and industry attention on the GLM series models under Zhipu AI is rising.
Moonshot AI launches open-source model Kimi K3, scoring 57 points on the Artificial Analysis Intelligence Index, becoming the third highest-scoring model, second only to Anthropic's Claude Opus 5 (61 points), Claude Fable 5 (60 points), and OpenAI's GPT-5.6 Sol (59 points). The index shows that the gap between leading closed-source models and open-weight models has narrowed to 4 points, the smallest gap since the release of GLM-5 in February. This signifies that open-source models are rapidly catching up to closed-source models in performance.
Chinese AI startup Moonshot AI will release the model weights of its high-performance model, Kimi K3. Developers can download the model, modify it for various purposes, and run it in their own data centers or cloud environments.Kimi K3 boasts 2.8 trillion parameters and a 1 million token context window, enabling it to process large-scale documents and codebases in a single pass. Moonshot AI plans to later publish a technical report detailing the model's architecture, training methodology, and performance evaluation results.Following the release of Kimi K3, Moonshot AI's daily revenue is reported to have increased by at least 6 times. The company is reportedly advancing a new round of fundraising at a $50 billion valuation and is considering a Hong Kong listing as early as this year.According to Bloomberg Intelligence, following the release of Kimi K3 and Z.AI's GLM-5.2, the share of Chinese open-weight models in overall token usage has risen to 68%. Services like AWS Bedrock, Microsoft Azure Foundry, and Google Vertex AI currently do not offer Chinese open-weight models such as Kimi K3 and GLM-5.2.