Supports Multi Card Cascaded Inference—RK1828 Edge AI Local Deployment Solution

0 comments

  As AI moves from the cloud to on-site devices, the computing power and memory of a single edge NPU can become limiting. Based on a multi-card cascaded architecture combining an RK3588 host controller with RK1828 NPU accelerator cards, this solution unifies the scheduling of compute and memory across multiple NPUs for collaborative inference. It enables large-parameter models, visual understanding, speech processing, and Agent tool capabilities to run on local devices, with flexible scaling from lightweight 7B inference on one card to 12B on two cards and 27B on four cards.

 The solution uses the RKNN3 and RKLLM3 inference software stack, supports browser-based interaction and the OpenAI API, and can connect to coding Agent tools such as OpenCode. Model weights, input data, inference processes, and output results can all remain local, providing a controllable and scalable foundation for edge AI applications.

■  RK1828 Local Deployment Solution

01.  Live Solution Demo on Real Hardware

02.  Large Model Inference at the Edge

  Data collection, preprocessing, model inference, and result output can all be completed on the local device. This makes the solution suitable for offline or weak-network environments, as well as edge scenarios with strict data-access requirements.

The RK1828 local inference solution deploys the complete model runtime on the device:

03.  Running Large Parameter Models with Multi Card Cascading

  For large-parameter models that cannot fit on a single card, RK1828 can split the model into multiple segments and connect multiple accelerator cards through PCIe to perform cascaded inference:   In a four-card cascaded configuration, the four model segments are loaded onto four RK1828 cards. A pipeline on the host side coordinates data transfer during the prefill and decode stages. This approach preserves model integrity while enabling local deployment of large-parameter models on edge devices.

04.  Unified Application Access

After deployment, applications can access the local model through a unified entry point:

· Access the local Web UI through a browser for visual conversations.

· Connect existing business systems through an OpenAI-compatible API.

· Call the local model through the Python OpenAI SDK.

· Use OpenCode with the local model for code analysis, file modification, and engineering validation.

· Extend file operations, Shell, MCP, and internal enterprise API capabilities through Agent toolchains.

■  Advantages of the RK1828 Solution

 Local Operation Reduces Long Term API Costs

  The model runs on the local RK1828 NPU, so routine inference does not depend on cloud token billing. For frequent question answering, coding assistance, document processing, meeting summaries, and continuously running edge tasks, hardware costs can be planned according to the deployment scale, reducing reliance on remote API charges.

Data Stays Local

  The local inference pipeline can keep code, documents, images, video, and meeting content on the device or within the enterprise intranet. Code analysis, code modification, meeting transcription, summarization, and semantic retrieval can all be completed locally, while model logs, conversation records, and business results can be managed under local policies.

Flexible Model and Deployment Scaling

  The solution supports model switching through a model directory and an OpenAI-compatible interface. Depending on model size and application type, inference can be configured with one, two, or four cards:

Configuration

Typical Models or Capabilities

Typical Use Cases

Single card

Small and medium language models, vision models, detection models

Lightweight conversations, visual analysis, single-stream inference

Dual card

Gemma4 12B two-segment model

Large-model conversations, multimodal understanding

Four card

Qwen3.5 or Qwen3.8 27B four-segment models

Large Chinese-language models, coding Agents, local knowledge Q&A

 

Agent and Development Tool Support

  A local large model can do more than answer questions; it can also serve as the reasoning core of an Agent. It can connect to file read/write tools, Shell and compilation tools, local code repositories, MCP services, document parsers, and knowledge bases. It can also call internal enterprise HTTP APIs and work with the OpenCode coding Agent to complete programming tasks.

■  Performance Testing

  The following results come from performance tests of the Firefly RK3588 and RK1828 NPU accelerator cards and are provided to demonstrate verified local inference capability.

Model

Accelerator

Input Tokens

New Tokens

TTFT ms

TPOT ms

Decode TPS

Gemma4-12B

2x RK1828

5120

128

7576.88

31.85

31.40

Gemma4-31B

4x RK1828

5120

128

9077.22

70.03

14.28

Qwen3.5-9B

2x RK1828

5120

128

5910.11

30.23

33.08

Qwen3.5-27B

4x RK1828

5120

128

8718.15

76.57

13.06

Qwen3.8-27B

4x RK1828

5120

128

8699.48

76.45

13.08

Note: The data is taken from official documentation. During testing, both RK1828 and RK3588 were set to performance mode. VLM Vision and LLM timing were measured independently. For detailed test conditions, refer to the SDK documentation.

■  AIBOX PRO Provides Reliable Hardware Support

  AIBOX PRO currently uses RK3588 as the system, codec, and I/O host controller and connects external RK1828 NPU accelerator cards through the M.2 interface to expand large-model computing power. The current design supports two cards and provides a reliable hardware foundation for running large models. For model adaptation, it supports local inference through RKNN3 and RKLLM3, can run dual-card segmented inference for the Gemma4 12B multimodal model, and can also run vision-language models such as Qwen3-VL for tasks including image understanding and video analysis.

 

  If you are interested, please contact our sales team to request a trial slot. Support for a four-card cascaded configuration is planned.

Contact Firefly for More Industry Application Solutions

sales@t-firefly.com


200 TOPS Peak Performance

Leave a comment

Please note, comments need to be approved before they are published.