Inference is adistributed GPU cluster for LLM inference built on Solana. Inference.net is a global network of data centers serving fast, scalable, pay-per-token APIs for models like DeepSeek V3 and Llama 3.3.
: AI reasoning cloud startup General Compute has obtained a $400 million loan from Upper90. This deal is the world’s first financing project to use dedicated inference chips as collateral. The company has built a proprietary AI reasoning cloud platform based on SambaNova’s self-developed ASIC chips, primarily targeting Agent-type AI computing workloads. Compared to traditional GPU clouds, it offers faster token processing speeds and lower operational latency. The hardware requires no water cooling and can be directly deployed in traditional data centers and idle cryptocurrency mining facilities.
According to Reuters, Chinese AI startup DeepSeek is developing its own AI chips, three informed sources revealed. The chip is designed specifically for inference scenarios, rather than for model training. The project was launched approximately one year ago and remains in the early stages. The company has engaged with chip design, wafer foundry, and storage enterprises, and has quietly increased recruitment of chip design engineers without publicly posting job listings. If successfully developed, DeepSeek will reduce its reliance on Nvidia and Huawei Ascend chips, following the trend of global AI giants such as OpenAI and Anthropic developing their own hardware. Affected by U.S. export controls, DeepSeek previously shifted from Nvidia H800 to Huawei chips. This self-developed chip is regarded as a significant strategic transformation. Meanwhile, DeepSeek also plans to complete its first round of external financing, with a fundraising scale of approximately $7 billion, and a valuation reaching $52 billion to $59 billion.
"White-Haired Stock Guru" Serenity stated that while some observations in the UBS report hold anecdotal truth, the more noteworthy trend is the increasing number of Chinese-language reports regarding the distillation of Anthropic's models. Currently, many US startups and tech companies are opting to use cheaper Chinese models (such as DeepSeek) in their AI applications, as their unit task costs are significantly lower than those of inference models from Gemini, OpenAI, and Anthropic.Serenity believes this trend, driven by capitalism, creates a "typical paradox"—companies naturally gravitate towards lower-cost solutions, thereby eroding the leading advantage of US models. He proposes that the US needs to address this on two fronts:First, build stronger access control and authentication systems, such as "heavy KYC frontier models" for domestic US use and tiered access mechanisms for allies, to reduce the risk of model distillation and misuse. This could also be accompanied by introducing an identity verification system akin to "AI-grade banking authentication" (e.g., biometrics + short-lived permission tokens) to raise the barrier for model calls, and using regulatory measures to restrict account sharing and access resale.Second, enhance the cost efficiency of inference models, allowing them to comprehensively outperform competitors like DeepSeek in both price and performance.Serenity also noted that some high-end models are currently frequently targeted for "distillation exploitation." Ideally, access to models nearing the AGI level should involve increased friction costs. In summary, the core challenge for the US AI industry lies in achieving both "low-cost inference capabilities" and establishing model access security mechanisms comparable to those in the financial system.
sources say NVIDIA has begun pitching its first independent central processing unit (CPU) product, Vera, to Chinese clients. Designed specifically for Agentic AI systems, the chip has entered mass production, marking NVIDIA's attempt to further expand its presence in the Chinese market with a CPU offering.According to sources, some Chinese clients have already shown interest in Vera. One major Chinese cloud computing company plans to procure over 300 servers equipped with dual Vera CPUs for testing, and will decide whether to expand procurement after the tests are completed.Built on the Arm Holdings architecture, Vera is NVIDIA's first independent CPU product. NVIDIA has previously stated that Vera's performance in AI agent-related computing tasks is 1.8 times that of comparable competitor products, and expects the product to contribute approximately $20 billion in revenue by the end of this fiscal year (ending January next year).The report notes that as the AI industry's focus gradually shifts from model training to inference computing, CPUs and custom chips are gaining more attention. Vera also positions NVIDIA to directly compete with Intel and Advanced Micro Devices (AMD), which have long dominated the server CPU market.Sources indicate that due to strict U.S. export restrictions on high-end GPUs, CPUs face relatively smaller regulatory hurdles in the Chinese market compared to GPU products. Currently, some Chinese clients plan to first deploy Vera chips for testing in overseas data centers. Meanwhile, software ecosystem compatibility and existing domestic AI chip deployment frameworks may still impact the subsequent large-scale adoption of Vera. (Reuters)
: AI reasoning cloud startup General Compute has obtained a $400 million loan from Upper90. This deal is the world’s first financing project to use dedicated inference chips as collateral. The company has built a proprietary AI reasoning cloud platform based on SambaNova’s self-developed ASIC chips, primarily targeting Agent-type AI computing workloads. Compared to traditional GPU clouds, it offers faster token processing speeds and lower operational latency. The hardware requires no water cooling and can be directly deployed in traditional data centers and idle cryptocurrency mining facilities.
According to Reuters, Meta plans to mass-produce its self-developed data center AI chip "Iris" starting from September, as part of its fourth-generation Meta Training and Inference Accelerators project, to enhance the AI capabilities of platforms such as Facebook and Instagram and reduce reliance on external GPUs such as those from Nvidia and AMD. Internal memos show that Iris completed testing in just 6 weeks with no major defects; Meta plans to deploy 7 gigawatts of computing power this year and increase it to 14 gigawatts by 2027, with its AI infrastructure spending in 2024 potentially reaching up to $145 billion. To secure expansion, the company has signed long-term supply agreements with Samsung Electronics, Sandisk, and Sumitomo Electric to cope with "price increases" and shortages of memory and AI chips.
According to Reuters, Chinese AI startup DeepSeek is developing its own AI chips, three informed sources revealed. The chip is designed specifically for inference scenarios, rather than for model training. The project was launched approximately one year ago and remains in the early stages. The company has engaged with chip design, wafer foundry, and storage enterprises, and has quietly increased recruitment of chip design engineers without publicly posting job listings. If successfully developed, DeepSeek will reduce its reliance on Nvidia and Huawei Ascend chips, following the trend of global AI giants such as OpenAI and Anthropic developing their own hardware. Affected by U.S. export controls, DeepSeek previously shifted from Nvidia H800 to Huawei chips. This self-developed chip is regarded as a significant strategic transformation. Meanwhile, DeepSeek also plans to complete its first round of external financing, with a fundraising scale of approximately $7 billion, and a valuation reaching $52 billion to $59 billion.
sources say NVIDIA has begun pitching its first independent central processing unit (CPU) product, Vera, to Chinese clients. Designed specifically for Agentic AI systems, the chip has entered mass production, marking NVIDIA's attempt to further expand its presence in the Chinese market with a CPU offering.According to sources, some Chinese clients have already shown interest in Vera. One major Chinese cloud computing company plans to procure over 300 servers equipped with dual Vera CPUs for testing, and will decide whether to expand procurement after the tests are completed.Built on the Arm Holdings architecture, Vera is NVIDIA's first independent CPU product. NVIDIA has previously stated that Vera's performance in AI agent-related computing tasks is 1.8 times that of comparable competitor products, and expects the product to contribute approximately $20 billion in revenue by the end of this fiscal year (ending January next year).The report notes that as the AI industry's focus gradually shifts from model training to inference computing, CPUs and custom chips are gaining more attention. Vera also positions NVIDIA to directly compete with Intel and Advanced Micro Devices (AMD), which have long dominated the server CPU market.Sources indicate that due to strict U.S. export restrictions on high-end GPUs, CPUs face relatively smaller regulatory hurdles in the Chinese market compared to GPU products. Currently, some Chinese clients plan to first deploy Vera chips for testing in overseas data centers. Meanwhile, software ecosystem compatibility and existing domestic AI chip deployment frameworks may still impact the subsequent large-scale adoption of Vera. (Reuters)
According to Tech Funding News, AMD CEO Lisa Su announced at London Tech Week that the company will invest up to £2 billion in UK AI infrastructure over the next five years, covering national supercomputing infrastructure development and university research collaborations. Meanwhile, AMD is partnering with Oriole Networks—a startup spun out from University College London (UCL)—to deploy the world’s first large-scale, all-photonic network AI system under the UK government’s £50 million ARIA Inference Scaling Lab initiative. This system integrates Oriole’s PRISM photonic networking platform with AMD Instinct GPUs and EPYC CPUs; by completely eliminating electronic switches from the network core, it reduces core network energy consumption by 81% and cuts GPU idle time from 60% to under 1%.
NVIDIA announced on X platform that it has launched the open-source multimodal model Nemotron 3 Nano Omni today. The model adopts a 30B-A3B mixture-of-experts (MoE) architecture, supports a 256K context window, and can uniformly process video, audio, image, and text inputs. Compared to open-source omnimodal models at a similar interaction level, this model achieves up to a 9x increase in throughput, significantly reducing inference costs and improving scalability. Nemotron 3 Nano Omni is now available on Hugging Face, OpenRouter, and NVIDIA NIM, and has been adopted by enterprises including Aible, Applied Scientific Intelligence, and H Company.
The open-source inference engine WASTE enables running the full Kimi K3 model on a MacBook Pro equipped with 64GB memory. This solution reduces minimum memory usage to approximately 29.05 GiB by streaming expert weights on demand, enabling local inference without quantization.
OpenAI researcher Jeffrey Wang described the performance of the AI inference solution jointly developed by AMD and Cerebras as "incredible". The solution integrates the AMD Helios rack with the Cerebras wafer-scale engine, handling prompts and token generation stages separately, and is estimated to increase tokens per second per watt by 5 times. This technical approach optimizes inference performance through heterogeneous computing, reflecting the continued focus of leading AI labs on reducing inference costs.
: AI reasoning cloud startup General Compute has obtained a $400 million loan from Upper90. This deal is the world’s first financing project to use dedicated inference chips as collateral. The company has built a proprietary AI reasoning cloud platform based on SambaNova’s self-developed ASIC chips, primarily targeting Agent-type AI computing workloads. Compared to traditional GPU clouds, it offers faster token processing speeds and lower operational latency. The hardware requires no water cooling and can be directly deployed in traditional data centers and idle cryptocurrency mining facilities.
According to TechFlow Research, JPMorgan's July 15 report significantly raised server shipment expectations, with the 2026 growth rate revised up from 15% to 22%, and 2027 from 8% to 25%. AI inference is the core driver, as enterprises deploying AI models require a large number of inference servers. JPMorgan estimates that by 2028, server CPU shipments will increase from 26 million to 68 million units, of which Agentic AI-related demand accounts for 53 million units. The PC side is suppressed by rising memory prices; brands are raising prices to maintain gross margins, at the cost of sales volume. 2026 PC shipments are expected to decline by 8%, with consumer PCs down 14%. Supply bottlenecks remain a constraint; CPU, substrates, memory, PCB, power devices, every segment is tight. In terms of US stocks, AI server vendors such as Dell Technologies, HPE, and Super Micro Computer continue to benefit; in the components sector, Arista Networks, Amphenol, Corning, Lumentum, Micron Technology, etc. benefit from the structural trend of value shifting towards components. JPMorgan recommends: server components sector is superior to contract manufacturing, avoid PC overall.
According to Reuters, Meta plans to mass-produce its self-developed data center AI chip "Iris" starting from September, as part of its fourth-generation Meta Training and Inference Accelerators project, to enhance the AI capabilities of platforms such as Facebook and Instagram and reduce reliance on external GPUs such as those from Nvidia and AMD. Internal memos show that Iris completed testing in just 6 weeks with no major defects; Meta plans to deploy 7 gigawatts of computing power this year and increase it to 14 gigawatts by 2027, with its AI infrastructure spending in 2024 potentially reaching up to $145 billion. To secure expansion, the company has signed long-term supply agreements with Samsung Electronics, Sandisk, and Sumitomo Electric to cope with "price increases" and shortages of memory and AI chips.
According to Reuters, Chinese AI startup DeepSeek is developing its own AI chips, three informed sources revealed. The chip is designed specifically for inference scenarios, rather than for model training. The project was launched approximately one year ago and remains in the early stages. The company has engaged with chip design, wafer foundry, and storage enterprises, and has quietly increased recruitment of chip design engineers without publicly posting job listings. If successfully developed, DeepSeek will reduce its reliance on Nvidia and Huawei Ascend chips, following the trend of global AI giants such as OpenAI and Anthropic developing their own hardware. Affected by U.S. export controls, DeepSeek previously shifted from Nvidia H800 to Huawei chips. This self-developed chip is regarded as a significant strategic transformation. Meanwhile, DeepSeek also plans to complete its first round of external financing, with a fundraising scale of approximately $7 billion, and a valuation reaching $52 billion to $59 billion.