
Exhibition Time: July 17-20, 2026
Venue: Shanghai World Artificial Intelligence Conference (WAIC)
2026 World Artificial Intelligence Conference kicked off fiercely, and the core focus of the audience remained around.AI computing power scale landing, large model cost reduction and efficiency, domestic infrastructure replacement.three main lines. With long text large models, multi-modal AI, intelligent body applications.multi-facetedoutbreak, the industry's pain point has been upgraded from "not enough computing power" to "not fast enough interconnection, not enough memory, high reasoning costs".
the current WAIC,Stored InformationDual-Core Interconnected Products Debut: Mature CommercialPCIe 5.0 architecture 16-card Full Mesh fully interconnected AI computing cluster(two landing forms),Division I.self-developed PCIe 6.0 high-speed switch; simultaneous exhibition2U Rack 24-Disk DPU Hardware Unload JBOF Full Flash KV Cache Storage Whole Machine,PCIe5.0 Mature Compute + PCIe6.0 Frontier Interconnection + High Density Pooled Storagefull stack solution, for enterprise-level large model training, high concurrent reasoning, RAG knowledge base scenarios to provide landing, mass production, can be reduced to the domestic intelligent base.
1. mature commercial main force: 16-card PCIe5.0 Full Mesh fully interconnected AI computing power cluster, dual architecture on-demand selection
The core commercial real machine of this exhibition, two sets of hardware form unified carryingPCIe 5.0 switching chipbuild a non-blocking Full Mesh topology, covering the single-machine high-density, modular elastic expansion of two types of customer needs, to solve the traditional multi-card machine bandwidth congestion, high cross-card latency, computing power idle pain points.
Scheme 1: Integrated 16-card machine (single-machine high-density scheme)
The whole machine built-in self-researchPCIe 5.0 Full Mesh Switching Backplane,16 RTX5090 GPU integrated single chassis, native hardware all-pair interoperability.
•AllToAll measured bandwidth up to 45 GB/s,PCIe5.0 high bandwidth channel cross-cabinet signal attenuation, multi-card collaborative throughput pull full
•Unified power supply, heat dissipation, operation and maintenance of the whole machine, simple deploymenteasy, a single machine can carry 70B-400B model distributed training
•room wiring, low operation and maintenance costs, suitable for fixed computing room, long-term privatization of large model research and development team.
Scheme 2: 2U Computer Head + Two 8-Card GPU JBOG Expansion Cabinets (Modular Distributed Scheme)
head-expansion cabinet separation architecture, standard 2U dual-channel host is used as a management computing node, two 8-card GPU JBOG cabinets are externally connected through PCIe5.0 high-speed cable, and the three devices are linked to form a complete 16-card computing force cluster, which relies on PCIe5.0 exchange to realize full interconnection across cabinets.
•flexible split: only 2U head + single 8-card JBOG(8-card cluster) is used for business trough, and business expansion is superimposed.another one.expansion enclosure, zero hardware waste
•machine room: 2U machine head is suitable for general standard machine position, 8-card independent GPU cabinet disperses power consumption and heat dissipation pressure of the whole machine, and reduces the cost of liquid cooling transformation.
•Adaptation Phase Construction, Cloud Vendor Elastic Computing Platform, Multi-Project Staggered R & D Scenarios

Two sets of PCIe5.0 cluster unified core advantages
- PCIe5.0 Switching Chip to Build Full Mesh Non-blocking Topology, 16GPU CardTwo-two full-speed point-to-point communication, no host bus bottleneck
- AllToAll measured 45 GB/s high throughputto eradicate distributed training data queuing, IO jitter, and long tail inference delay
- is natively compatible with SGLang, vLLM, and DeepSpeed mainstream large model frameworks.
- supports 70B ~ 400B dense large model pre-training, ultra-long context multi-round dialogue reasoning.
- Landing Selection Comparison
traditional multi-card clusters mostly use low-generation PCIe and tree-like cascade switching, insufficient interconnection bandwidth causes idle computing power; Memory PCIe5.0 Full Mesh Architecture Realize 100 Release of Computation Power:
- integrated machine:designedpursuithighPerformance, Janeeasyoperations, fixed capacity scale customersDesign
- 2U head + dual 8 card JBOG:designedMachine room space is limited, phased investment, need flexible expansion of customers.Design
2. this WAIC blockbuster new product: MEI self-developed PCIe 6.0 high-speed switch, forward-looking layout of the next generation of intelligent computing interconnection
This conference is officially releasedDomestic self-developed PCIe 6.0 switch, is for the next generation of CXL 4.0, million card-class AI cluster, high-speed memory pool of the cutting-edge interconnected core hardware, to fill the domestic PCIe 6.0 switching equipment blank.

Core Technology Highlights
1.PCIe 6.0 Native 64GT/s Express, single channel bandwidth doubles to PCIe5.0, supports x16/x8 flexible Gearbox channel splitting, and is compatible with GPU, CXL memory card and NVMe storage devices
2.full cross-over non-blocking switching architecture, low hardware forwarding latency, support for multi-host, multi-cabinet global fabric networking
3.deep adaptation of CXL 3.1/4.0 protocol, cross-server global unified memory pool can be built to break through the physical limit of single-machine memory.
4.backward compatibility with PCIe5.0/4.0 devices, smooth upgrade of old and new clusters, no need to replace existing computing hardware with the whole machine
5.Full link domestic hardware design, independent and controllable, to meet the needs of government and enterprise trust creation, ultra-large-scale intelligent computing center construction.
landing application scenario
•the next generation of 10-card-level AI supercomputing cluster, massive GPU cross-cabinet high-speed AllToAll data interaction
•CXL global memory pooling platform, multi-server sharing TB-level extended memory, greatly reducing GPU video memory procurement costs
•high-density distributed JBOF storage cluster, carrying PB-level KV Cache, vector database high-speed read and write
•forward-looking computing power infrastructure projects, advance layout of PCIe6.0 standard, to avoid the risk of hardware iteration phase-out.
3. to solve the industry's core pain points: two-layer hardware architecture together to break down the "memory wall"
With the normalization of large models, the four major industry bottlenecks continue to be highlighted:
1.HBM memory is expensive, KV Cache is easy to fill up the memory.
long text and multi-user concurrent reasoning, the amount of KV cache data has skyrocketed, and a large number of new graphics cards are required to be loaded only by GPU memory, doubling the overall investment.
2.server native memory has a physical ceiling
motherboard is fixed, the single-machine expansion cost is high, the upper limit is low, and the system memory cannot be flexibly expanded.
3.Traditional storage IO performance is inadequate
ordinary storage latency is high, random read and write is weak, it is difficult to match the AI high-frequency KV cache read and write requirements, resulting in high reasoning TTFT, experience fluctuations.
4.legacy low-generation interconnect architecture bandwidth bottleneck
PCIe4.0 and below, non-Full Mesh cascade architecture, multi-card data transmission is congested, and GPU computing power cannot be fully released.
In response to the progressive pressure of memory resources, we will createPCIe5.0 Commercial Full Mesh Cluster/PCIe6.0 Prospective Switch + 2U 24 Disk JBOF Pooled Storagetwo-tier collaborative architecture to resolve the pressure of memory walls in layers:
•itsLayer 1: Dual-generation PCIe switching scheme. Mature PCIe5.0 cluster meets the current commercial computing power. The new PCIe6.0 switch is oriented to the next generation of high-speed interconnection and opens up high-speed data channels between GPU, memory and storage;
•itssecond floor: 2U 24-disk DPU JBOF full flash storage, carrying massive low-frequency KV cache, diverting expensive GPU video memory pressure.

4. Core Storage Exhibit: Mai Cun Self-developed 2U 24 Disk DPU Hardware Unloading JBOF Full Flash KV Cache Storage Complete Machine
High-density pooled storage specially designed for KV Cache scenarios, standard 2U rack specifications, 24 U.2 NVMe SSD in a single chassis, domestic high-performance AI reasoning cache benchmarking scheme.
Core Technology Highlights
•compact 2U 24 disc high density design: A single flash memory pool capacity of 100 TB can be realized, and multiple horizontal stacks can be expanded to PB level, thus occupying less machine space in the computer room;
•Full DPU Hardware Offload Architecture: onboard DPU chip, NVMe-oF, RoCE RDMA, EC erasure, data compression all hardware acceleration;
•host CPU usage is reduced to less than 5%, to solve the traditional storage preemption computing power, pull up the reasoning delay problem;
•NVMe over Fabrics remote pooling, compatible with PCIe5.0/6.0 GPU cluster, multiple computing servers share a unified flash resource pool;
•native three-tier intelligent cache tiering:GPU HBM (Hot Data) & rarr; Host Memory (Warm Data) & rarr;2U 24 Disk JBOF Flash Pool (Cold KV Data)
Measured business income
automatically sinks low-frequency and ultra-long context KV data to low-cost enterprise flash memory, freeing up valuable GPU memory and improving inference concurrency without adding a new graphics card:
•Large Model Inference Concurrent Carrying Capacity Improvementmore than 50%
•The problem of memory overflow in long context scenes is basically eliminated.
•Comprehensive TCO of the whole intelligent computing cluster is reduced by 35%-45%
is native to the mainstream reasoning framework of vLLM and SGLang, plug and play with zero transformation, and is suitable for privatized enterprise knowledge base and multi-tenant AI reasoning platform.
5. Memory Core Differentiation: Full Link Self-Research PCIe Interconnect Layering Solution
is different from the market only do the whole machine assembly manufacturers, Meicun has from the exchange hardware design, GPU cluster architecture to the storage of the whole machine of the full stack of self-research capabilities, this WAIC to form a mature commercial + frontier forward-looking complete product matrix:
1.dual-generation PCIe switching technology full coverage
Mature commercial PCIe5.0 Full Mesh computing power cluster (integrated machine/2U head +8-card JBOG dual form),AllToAll measured 45 GB/s; Released at the same timeself-developed PCIe6.0 switch, forward-looking layout of the next generation of high-speed CXL smart computing cluster, taking into account the current landing and future iteration.
2.Self-developed 2U 24-disk high-density JBOF storage machine
is equipped with DPU hardware unloading architecture, which is specially tuned for KV Cache high frequency random IO depth and diverts video memory pressure from the storage layer.
3.Full Link Localization Bottom Design
switching backplane, high-speed channel, storage machine independent research and development, adapt to the government and enterprise letter creation, super-large-scale domestic intelligent computing center construction needs.
does not pile up materials and does not have a premium. The hierarchical architecture solves the three bottlenecks of computing power, display memory and storage in one stop: PCIe5.0 commercial cluster guarantees existing services, the new PCIe6.0 switch supports future computing power upgrade, and 2U 24-disk JBOF undertakes massive KV cache, allowing enterprises to "buy less cards, more concurrent, stable operation and long-term iterability".

6. WAIC field communication is hot, multiple sets of real machine synchronous demonstration
, the Meicun booth has welcomed a large number of AI algorithm enterprises, computing integrators, government and enterprise intelligence computing platforms, and customers of scientific research institutions. Technical team on-site real machine demonstration:
16-card PCIe5.0 cluster 45 GB/s AllToAll throughput measurement, new PCIe6.0 switch high-speed interconnection demonstration, 8-card GPU JBOG modular elastic expansion scheme, 2U 24-disk JBOF KV Cache sinking measurement, overall delay and concurrent data.
Discussion on the Depth of Site SynchronizationDomestic computing power localization substitution, long context reasoning cost reduction, CXL global memory pooling, PCIe6.0 next-generation Fabric, storage and computing separation architecture upgrade.industry trends.



2026 AI industry enters a large-scale landing cycle, computing power infrastructure.bus intergenerational iteration capability, elastic expansion capability, high-density storage cost advantage, it directly determines the efficiency of the commercialization of enterprise AI.
Memory Information Continues to DeepenPCIe5.0 commercial GPU cluster, self-developed PCIe6.0 high-speed switch, modular GPU JBOG expansion cabinet, 2U24 high-density JBOF storage, KV Cache bottom layer optimizationfull-stack products, continuous output can be mass-produced, cost-effective, smooth iteration of domestic intelligent computing infrastructure, to help the AI industry from "can be used" to "easy-to-use, low-cost, long-term sustainable large-scale commercial".