As AI moves from the cloud to on-site devices, the computing power and memory of a single edge NPU can become limiting. Based on a multi-card cascaded architecture combining an RK3588 host controller with RK1828 NPU accelerator cards, this solution unifies the scheduling of compute and memory across multiple NPUs for collaborative inference. It enables large-parameter models, visual understanding, speech processing, and Agent tool capabilities to run on local devices, with flexible scaling from lightweight 7B inference on one card to 12B on two cards and 27B on four cards.
The solution uses the RKNN3 and RKLLM3 inference software stack, supports browser-based interaction and the OpenAI API, and can connect to coding Agent tools such as OpenCode. Model weights, input data, inference processes, and output results can all remain local, providing a controllable and scalable foundation for edge AI applications.
■ RK1828 Local Deployment Solution
01. Live Solution Demo on Real Hardware
02. Large Model Inference at the Edge
Data collection, preprocessing, model inference, and result output can all be completed on the local device. This makes the solution suitable for offline or weak-network environments, as well as edge scenarios with strict data-access requirements.
The RK1828 local inference solution deploys the complete model runtime on the device:

03. Running Large Parameter Models with Multi Card Cascading
For large-parameter models that cannot fit on a single card, RK1828 can split the model into multiple segments and connect multiple accelerator cards through PCIe to perform cascaded inference:
In a four-card cascaded configuration, the four model segments are loaded onto four RK1828 cards. A pipeline on the host side coordinates data transfer during the prefill and decode stages. This approach preserves model integrity while enabling local deployment of large-parameter models on edge devices.
04. Unified Application Access
After deployment, applications can access the local model through a unified entry point:
· Access the local Web UI through a browser for visual conversations.
· Connect existing business systems through an OpenAI-compatible API.
· Call the local model through the Python OpenAI SDK.
· Use OpenCode with the local model for code analysis, file modification, and engineering validation.
· Extend file operations, Shell, MCP, and internal enterprise API capabilities through Agent toolchains.
■ Advantages of the RK1828 Solution
Local Operation Reduces Long Term API Costs
The model runs on the local RK1828 NPU, so routine inference does not depend on cloud token billing. For frequent question answering, coding assistance, document processing, meeting summaries, and continuously running edge tasks, hardware costs can be planned according to the deployment scale, reducing reliance on remote API charges.
Data Stays Local
The local inference pipeline can keep code, documents, images, video, and meeting content on the device or within the enterprise intranet. Code analysis, code modification, meeting transcription, summarization, and semantic retrieval can all be completed locally, while model logs, conversation records, and business results can be managed under local policies.
Flexible Model and Deployment Scaling
The solution supports model switching through a model directory and an OpenAI-compatible interface. Depending on model size and application type, inference can be configured with one, two, or four cards:
|
Configuration |
Typical Models or Capabilities |
Typical Use Cases |
|
Single card |
Small and medium language models, vision models, detection models |
Lightweight conversations, visual analysis, single-stream inference |
|
Dual card |
Gemma4 12B two-segment model |
Large-model conversations, multimodal understanding |
|
Four card |
Qwen3.5 or Qwen3.8 27B four-segment models |
Large Chinese-language models, coding Agents, local knowledge Q&A |
Agent and Development Tool Support
A local large model can do more than answer questions; it can also serve as the reasoning core of an Agent. It can connect to file read/write tools, Shell and compilation tools, local code repositories, MCP services, document parsers, and knowledge bases. It can also call internal enterprise HTTP APIs and work with the OpenCode coding Agent to complete programming tasks.
■ Performance Testing
The following results come from performance tests of the Firefly RK3588 and RK1828 NPU accelerator cards and are provided to demonstrate verified local inference capability.
|
Model |
Accelerator |
Input Tokens |
New Tokens |
TTFT ms |
TPOT ms |
Decode TPS |
|
Gemma4-12B |
2x RK1828 |
5120 |
128 |
7576.88 |
31.85 |
31.40 |
|
Gemma4-31B |
4x RK1828 |
5120 |
128 |
9077.22 |
70.03 |
14.28 |
|
Qwen3.5-9B |
2x RK1828 |
5120 |
128 |
5910.11 |
30.23 |
33.08 |
|
Qwen3.5-27B |
4x RK1828 |
5120 |
128 |
8718.15 |
76.57 |
13.06 |
|
Qwen3.8-27B |
4x RK1828 |
5120 |
128 |
8699.48 |
76.45 |
13.08 |
Note: The data is taken from official documentation. During testing, both RK1828 and RK3588 were set to performance mode. VLM Vision and LLM timing were measured independently. For detailed test conditions, refer to the SDK documentation.
■ AIBOX PRO Provides Reliable Hardware Support
AIBOX PRO currently uses RK3588 as the system, codec, and I/O host controller and connects external RK1828 NPU accelerator cards through the M.2 interface to expand large-model computing power. The current design supports two cards and provides a reliable hardware foundation for running large models. For model adaptation, it supports local inference through RKNN3 and RKLLM3, can run dual-card segmented inference for the Gemma4 12B multimodal model, and can also run vision-language models such as Qwen3-VL for tasks including image understanding and video analysis.
If you are interested, please contact our sales team to request a trial slot. Support for a four-card cascaded configuration is planned.
Contact Firefly for More Industry Application Solutions
sales@t-firefly.com