SOURCE // NEWS

SiliconFlow Files for Hong Kong IPO, Aiming to Be First 'AI Token Factory' Stock

SiliconFlow Files for Hong Kong IPO, Aiming to Be First 'AI Token Factory' Stock

While Large Language Model (LLM) companies are actively seeking listings on the Hong Kong market, the providers of the underlying computing power are now launching their own capital market sprints. AI infrastructure startup SiliconFlow has officially submitted its listing application to the Hong Kong Stock Exchange, aiming to become the first 'AI Token Factory' stock. Prior to this, #SiliconFlow had completed 7 rounds of financing, reaching a valuation of 77.4 billion RMB, with investments from heavyweights such as Alibaba, Meituan, SenseTime, Zhipu AI, and Huawei's Hubble Technology.

The core founding team of SiliconFlow shares a notable lineage: they are the original team behind OneFlow, founded by Yuan Jinhui, a student of Academician Zhang Bo. Following OneFlow's acquisition by Light Year Beyond and its subsequent integration into Meituan, Yuan chose to entrepreneur again, founding SiliconFlow. According to the prospectus, SiliconFlow's core value proposition is not 'model building' but acting as a 'Token Factory' in the AI inference era. Instead of building models or applications, the company packages heterogeneous compute, multi-model support, and corporate API call demands into stable, metered, and billable Token supply services.

The prospectus reveals that as of April 30, 2026, SiliconFlow's registered platform users reached 10.28 million. In April 2026, its average daily token throughput reached 578.5 billion. The platform supports over 170 models and serves more than 13,000 enterprise clients. According to Frost & Sullivan, based on 2025 annual token throughput, SiliconFlow is the fourth-largest token supply platform in China (1.5% market share) and ranks first among all independent ecosystem token supply platforms.

Despite rapid growth, the financial data reveals the early-stage pains of the Token business. In 2025, SiliconFlow recorded revenue of 55.33 million RMB, a year-on-year surge of 653.2% from 7.35 million RMB in 2024. However, its net loss widened 4.2 times from 81.92 million RMB in 2024 to 345 million RMB in 2025, with an adjusted net loss of 187 million RMB. Crucially, its gross profit margin flipped from 39.4% in 2024 to -24.0% in 2025, with the public cloud service margin plunging to -119.0%.

This margin inversion was driven by a structural shift in business. In 2025, public cloud services generated 52.9% of SiliconFlow's revenue, surpassing private deployments. The exponential expansion of users and Token throughput directly drove up back-end computing consumption, pushing cost of sales from 4.45 million RMB in 2024 to 68.63 million RMB in 2025. The prospectus explains that the cost spike was primarily driven by the procurement of computing power resources to support public cloud expansion. As AI applications and autonomous Agent ecosystems flourish, Token consumption scales exponentially, placing immense cost pressure on infrastructure layers competing on pricing and volume.

[AgentUpdate Depth Analysis] SiliconFlow's IPO marks a watershed moment for the AI Agent ecosystem, signaling the shift from proprietary LLM hype to mass-scale, commoditized inference infrastructure. While hardware-centric players like Groq focus on custom ASIC speed, SiliconFlow leverages software-defined optimization—rooted in its OneFlow heritage—to maximize heterogeneous GPU efficiency. As the industry transitions from simple chatbots to complex, multi-turn AI Agents and autonomous workflows, Token consumption is scaling exponentially. However, SiliconFlow’s negative 119% public cloud gross margin highlights a critical industry paradox: scaling API call volumes currently exacerbates computing cost pressures. For the AI Agent ecosystem to thrive sustainably, infrastructure providers must transition from subsidized pricing wars to software-driven unit economic profitability. Ultimately, the company that achieves sustainable margin control will dictate the runtime economics of next-generation autonomous agents.