16 Card GPU BOX
A successful case ---- Custom AI Server Project for a Smart Computing Center

Case Background
In 2024, a certain smart computing center required the deployment of a high-performance, highly reliable, and fully domestic (China-made) AI computing platform. The platform was specifically optimized to run the full range of DeepSeek models efficiently, while simultaneously reducing deployment costs. At that time, the market predominantly offered servers limited to 8-GPU configurations. However, the client demanded a single-system solution capable of supporting 16 GPUs with high-speed inter-GPU connectivity, while also retaining the scalability to expand to 32 GPUs in the future. To address this challenge, we engineered an integrated AI computing solution featuring a core 16-GPU BOX unit paired with a front-mounted 2U rack-mount server.
Customer Pain Points
- 01
Flexible Configuration
Flexible configuration with high scalability, allowing for future expansion into a small-scale GPU cluster.
- 02
Cost Reduction
Without compromising performance, the single 16-GPU system offers a more competitive cost advantage compared to two separate 8-GPU servers.
- 03
Full-Stack Domestic Adaptation
Localization of core components, perfect compatibility with domestic operating systems and large model platforms.
Solution
Maicun Information Technology delivered a tailored AI computing solution for this smart computing center, featuring a core 16-GPU BOX paired with a front-mounted 2U rack-mount server.This solution comprehensively optimizes computing power, communication, storage, and management capabilities, and is perfectly compatible with the private deployment of the entire DeepSeek model suite. The core configuration is as follows:

Core Computing Node:
16-Card GPU BOX (4U Rack mount)
Equipped with 16 MetaX MXC500 compute cards, delivering 3.84 PFLOPS of FP16 ultra-high computing power, paired with 1024GB of HBM2e video memory and 1.8 TB/s of video memory bandwidth.
· 16 GPU Interconnection
Supports 4/8/16 card to card interconnection ( MetaxLink 34GB/s), with an innovative topology achieving 300% GPUxCCL communication bandwidth, and efficiently supports both MOE and dense models;
· Ultra-low Latency Lossless Direct-Connect Architecture
Adopts lossless communication design, with direct connections between network, NVMe storage, and GPU, eliminating data interaction latency;
· Multi-Path PCIe Switch Expansion for High-Bandwidth Optimization
Locally supports 8 PCIe switches direct connection with NVMe SSD, and the network is compatible with 4 PCIe Switches for direct 100G/200G/400G InfiniBand (IB)/RDMA over Converged Ethernet (RoCE) connectivity;
Supporting Node:
Front-mounted 2U Rack Server
Equipped with two Hygon C86-4G 7400 Series processors, supporting up to 8TB DDR5 ECC memory, delivering robust computing power for computing nodes;
· Rich expansion and storage capabilities
Supporting PCIe standard cards, OCP cards (10G/100G/200G network)
· Supports Multiple Disk Form Factors
Local storage compatible with mixed installation of multi-specification hard drives, supporting full range of RAID modes;
· Redundant Cooling for Hardware Reliability
Equipped with 1+1 redundant power supplies and four hot-swappable fans, forming a highly reliable hardware system with core nodes.


Full-Stack Domestic Adaptation
This 4U 16-card GPU BOX delivers 3.84 PFLOPS FP16 power, 1024GB HBM2e, and low-latency interconnection for high-performance AI workloads.
· Full-Stack Domestic R&D
Fully independent design and localization of core components, perfect compatibility with domestic operating systems and large model platforms.
· On-Premises Model Deployment
Support for private deployment of the full range of DeepSeek models, achieving independent and controllable systems from hardware to software.
Implementation Effectiveness
Upon implementation, the intelligent computing center has successfully established a localized high computing power large-model intelligent computing platform. Its core business metrics have achieved leapfrog improvements, fully meeting the full-scenario application requirements of local large models.
· Breakthrough in Large Model Running Capability:
A single node can efficiently run the full version of DeepSeek-R1 (671B), significantly reducing deployment costs. The inference efficiency of DeepSeek 32B and 70B distilled models has increased by 50%.
· Improvement in Computing Power Coordination Efficiency:
The 16-card high-density computing design increases computing power utilization by over 60%. The flexible GPU interconnect modes and decoupled architecture enable on-demand allocation of computing resources, resulting in an over 40% improvement in large model fine-tuning efficiency.
· Hardware Reliability Maximized:
The high-redundancy design for power supplies and fans enables the platform to achieve 99.99% availability, supporting 7×24-hour continuous stable operation. A single component failure does not affect the overall system's normal functioning.

Customer Reviews
The 16-card GPU BOX Intelligent Computing Solution has perfectly addressed numerous pain points of our existing computing power platform, achieving dual upgrades in both computing power scale and operational efficiency.
-
·01 Maximum Single-Node Compute Performance
A single node can efficiently run the full version of the DeepSeek-R1 large model, significantly reducing the deployment cost of large models.
-
· 02 Independently Developed Domestic Foundation
The core components have achieved full localization, making the computing power foundation truly autonomous, controllable, and secure in data.
-
· 03 Local Data Security
A secure underlying architecture enforces strict local data isolation and storage, establishing a robust defense for information security.
-
·04 High-Speed, Low-Latency Interconnect
The lossless communication topology and high-speed card interconnection design of the solution completely eliminate communication latency in large model training and fine-tuning. The improved inference efficiency also enables us to better provide computing power services to local government-enterprise institutions and scientific research organizations.
-
· 05 All-in-One Integrated Computing
Maicun's integrated solution balances computing power, reliability, and scalability with outstanding cost-effectiveness, laying a solid computing power foundation for the implementation of the regional artificial intelligence industry.
-
· 06 Lower Total Cost of Ownership (TCO)
Coupled with simplified hardware deployment, this approach effectively reduces the overall investment required for large model implementation.
