# multiper **Repository Path**: null_163_8146/multiper ## Basic Information - **Project Name**: multiper - **Description**: 多流视觉语言模型推理系统,支持多路视频流实时处理与智能分析。 - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 1 - **Created**: 2026-07-29 - **Last Updated**: 2026-07-29 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README > The use of this document is subject to the [Copyright and Disclaimer](https://gitee.com/xuanwu_s3/docs/blob/master/版权与免责声明.md). Please read it carefully before use. # MultiStream-VLM Multi-stream Visual Language Model inference system for real-time multi-channel video processing and intelligent analysis. ## Introduction MultiStream-VLM is a multi-stream video processing system designed for RK3588 platform, featuring: - Multi-channel video stream decoding and processing - YOLO object detection - VLM (Visual Language Model) inference - Real-time video display output ## Directory Structure ``` multiper/ ├── CMakeLists.txt # Main CMake configuration ├── build.sh # One-click build script ├── cmake/ # CMake configuration files │ ├── toolchain.cmake # Cross-compilation toolchain config │ ├── msvlm.cmake # MSVLM module build config │ ├── vi_yolo_venc.cmake # VI-YOLO-VENC module build config │ ├── hdmi_in_vlm.cmake # HDMI-IN-VLM module build config │ ├── rockit_yolo_detect.cmake # ROCKIT-YOLO-DETECT module build config │ ├── ai_hub.cmake # AI-HUB module packaging config (no compilation) │ ├── anomaly_detection.cmake # ANOMALY-DETECTION module packaging config (no compilation) │ ├── install.cmake # Installation config │ └── package.cmake # Packaging config ├── sample/ # Sample modules │ ├── msvlm/ # MSVLM module source │ ├── msevent/ # MSEVENT module source │ ├── vi_yolo_venc/ # VI-YOLO-VENC module source │ ├── hdmi_in_vlm/ # HDMI-IN-VLM module source │ ├── rockit_yolo_detect/ # ROCKIT-YOLO-DETECT module source │ ├── ai_hub/ # AI-HUB module source (pure Python, no compilation) │ └── anomaly-detection/ # ANOMALY-DETECTION module source (pure Python, no compilation) ├── include/ # Public header files ├── third/ # Third-party dependencies ├── _install/ # Installation directory │ ├── msvlm/ # MSVLM installed files │ ├── msevent/ # MSEVENT installed files │ ├── vi_yolo_venc/ # VI-YOLO-VENC installed files │ ├── hdmi_in_vlm/ # HDMI-IN-VLM installed files │ ├── rockit_yolo_detect/ # ROCKIT-YOLO-DETECT installed files │ ├── ai_hub/ # AI-HUB installed files │ ├── anomaly-detection/ # ANOMALY-DETECTION installed files │ └── overlay/ # Firmware and model files ├── build/ # Build directory └── output/ # Package output directory ``` ## Build Instructions ### Requirements - CMake >= 3.16 - aarch64-none-linux-gnu cross-compilation toolchain - Target platform: RK3588 (ARM aarch64) ### Build Commands Use the `build.sh` one-click build script: ```bash # Show help ./build.sh -h # Clean build directories (no compilation) ./build.sh -c # Build all modules ./build.sh -a # Build MSVLM module only ./build.sh -m # Build MSEVENT module only ./build.sh -e # Build VI-YOLO-VENC module only ./build.sh --vi_yolo_venc # Build HDMI-IN-VLM module only ./build.sh --hdmi_in_vlm # Build ROCKIT-YOLO-DETECT module only ./build.sh --rockit_yolo_detect # Package AI-HUB module only (pure Python, no compilation) ./build.sh --ai_hub # Package ANOMALY-DETECTION module only (pure Python, no compilation) ./build.sh --anomaly_detection # Build and install ./build.sh -m -i # Build, install and package ./build.sh -m -i -p # Clean, build, install and package ./build.sh -c -m -i -p ``` ### Build Options | Option | Description | |--------|-------------| | `-h, --help` | Show help message | | `-a, --all` | Build all modules (default) | | `-m, --msvlm` | Build MSVLM module only | | `-e, --msevent` | Build MSEVENT module only | | `--vi_yolo_venc` | Build VI-YOLO-VENC module only | | `--hdmi_in_vlm` | Build HDMI-IN-VLM module only | | `--rockit_yolo_detect` | Build ROCKIT-YOLO-DETECT module only | | `--ai_hub` | Package AI-HUB module only (pure Python, no compilation) | | `--anomaly_detection` | Package ANOMALY-DETECTION module only (pure Python, no compilation) | | `-c, --clean` | Clean build and output directories | | `-i, --install` | Run install after build | | `-p, --package` | Create packages after build | ## Installation Directory Structure After build and installation: ``` _install/ ├── msvlm/ │ ├── bin/ # Executable files │ │ └── msvlm │ ├── lib/ # Shared libraries │ ├── config/ # Configuration files │ │ ├── default.json │ │ └── vlm_system_prompt.txt │ ├── scripts/ # Runtime scripts │ │ ├── start.sh # Start service (with HDMI kiosk UI) │ │ └── stop.sh # Stop service │ ├── models/ # Detection model files │ │ └── yolov8/ # YOLOv8n detection model │ ├── web/ # Web interface resources (HTTP/WebSocket) │ └── install.sh # Deploy models to /userdata/models ├── vi_yolo_venc/ │ ├── bin/ # Executable files │ │ └── vi_yolo_venc │ ├── lib/ # Shared libraries │ ├── models/ # Model files │ │ └── yolov8/ # YOLOv8n detection model │ ├── run.sh # Launch script │ └── install.sh # Deploy models to /userdata/models ├── hdmi_in_vlm/ │ ├── bin/ # Executable files │ │ └── hdmi_in_vlm │ ├── lib/ # Shared libraries │ ├── font/ # Font files for OSD rendering │ ├── models/ # Model files │ │ └── yolov5s/ # YOLOv5s detection model │ ├── run.sh # Launch script │ └── install.sh # Deploy fonts and models to /userdata ├── rockit_yolo_detect/ │ ├── bin/ # Executable files │ │ ├── rockit_yolo_detect # Main program │ │ └── rockit_yolo_detect_debug # Stream add/remove debug tool │ ├── lib/ # Shared libraries │ ├── resource/ # Background image and model │ │ ├── background-image.png │ │ └── model/ # YOLOv5s detection model │ ├── run.sh # Launch script │ └── install.sh # Deploy resources to /usr/share/resource ├── ai_hub/ # AI-HUB (pure Python, no compilation) │ ├── app.py # Flask application entry (port 5000) │ ├── backend/ # Backend blueprints (llm/multimodal/cnn/embedding/audio, etc.) │ ├── static/ # Frontend static assets │ ├── templates/ # Page templates │ ├── system/ # Binaries/libs/model data deployed to board system dirs │ ├── run.sh # Service start/stop script │ └── install.sh # Install Python deps and deploy resources └── anomaly-detection/ # ANOMALY-DETECTION (pure Python, no compilation) ├── app.py # Flask application entry (port 5000) ├── patchcore_train.py # PatchCore dataset annotation/library building ├── patchcore_test.py # PatchCore dataset detection ├── patchcore-inspection/ # PatchCore algorithm library (installed via pip install -e) ├── model/ # DINOv3 feature extraction model (onnx/rknn/weight) ├── static/ # Frontend static assets ├── templates/ # Page templates ├── system/ # Desktop icons, etc. deployed to board system dirs ├── run.sh # Service start/stop script └── install.sh # Install Python deps and deploy resources ``` ## Output Directory Packaged output files are located in `output/` directory: ``` output/ ├── msvlm/ # MSVLM complete package ├── msvlm.tar.gz # MSVLM tarball ├── msevent/ # MSEVENT complete package ├── msevent.tar.gz # MSEVENT tarball ├── vi_yolo_venc/ # VI-YOLO-VENC complete package ├── vi_yolo_venc.tar.gz # VI-YOLO-VENC tarball ├── hdmi_in_vlm/ # HDMI-IN-VLM complete package ├── hdmi_in_vlm.tar.gz # HDMI-IN-VLM tarball ├── rockit_yolo_detect/ # ROCKIT-YOLO-DETECT complete package ├── rockit_yolo_detect.tar.gz # ROCKIT-YOLO-DETECT tarball ├── ai_hub/ # AI-HUB complete package ├── ai_hub.tar.gz # AI-HUB tarball ├── anomaly-detection/ # ANOMALY-DETECTION complete package └── anomaly-detection.tar.gz # ANOMALY-DETECTION tarball ``` ## Module Description ### MSVLM Module Multi-stream Visual Language Model processing module with features: - Multi-channel video stream decoding (local files, RTSP) - YOLO object detection inference - VLM visual language model inference - NPU multi-core load balancing scheduling - Real-time video display output (VO/HDMI) - Web monitoring UI (HTTP/WebSocket) + HDMI kiosk browser UI overlay ### VI-YOLO-VENC Module Camera capture with YOLO detection and encoded streaming module, features: - Camera video capture via VI (default 1920x1080) - YOLOv8n object detection with bounding box and class-name label overlay (NV12 drawing) - Only person/car/bus/cat/dog detections are kept; all other classes are filtered out - H.264/H.265 encoding - RTSP live streaming output - Optional saving of the encoded bitstream to a local file ### HDMI-IN-VLM Module HDMI input capture with VLM inference module, features: - HDMI-IN video capture via RK628-CSI bridge (1920x1080) - YOLO object detection with bounding box overlay - FastVLM visual language model inference (scene description) - FreeType OSD rendering of VLM results on video frames - H.264/H.265 encoding with RTSP live streaming output - Video display output via VO ### MSEVENT Module Event detection module (reserved for future use). ### ROCKIT-YOLO-DETECT Module Multi-channel video file decoding with YOLO detection and VO display module, features: - MP4 file parsing via rkdemuxer with Rockit VDEC hardware decoding - YOLOv5s object detection (supports both RKNN3 and RKNN2 inference backends) - Cairo-drawn bounding boxes overlaid via VO RGN overlay - Multi-channel grid layout, with a background image shown when no stream is active - Dynamic stream add/remove at runtime via `rockit_yolo_detect_debug` Executables: - `rockit_yolo_detect`: main program, shows a background image on startup and waits for streams - `rockit_yolo_detect_debug`: debug tool to create/remove streams dynamically ```shell # On the board: first deploy resources (background + models) to /usr/share/resource ./install.sh # Run the main program ./run.sh # In another terminal: add a stream looping forever ./bin/rockit_yolo_detect_debug path add /userdata/input.mp4 -1 # Remove the first stream ./bin/rockit_yolo_detect_debug path remove 0 ``` ### AI-HUB Module On-device AI capability hub (pure Python web app, no compilation), features: - Flask-based Web UI to explore multiple on-device AI capabilities (default port 5000) - LLM chat with large language models - Multimodal (VLM) image-text understanding - CNN vision inference (classification ResNet/MobileNet, detection YOLO) - Embedding for text/image vectorization - Audio speech recognition (Whisper ASR) - Online model management and system settings - One-click deploy scripts (install.sh installs deps and deploys binaries/models, run.sh manages the daemon) ### ANOMALY-DETECTION Module Industrial defect detection example (pure Python web app, no compilation), features: - Flask-based Web UI for industrial inspection demo (default port 5000) - PatchCore anomaly detection algorithm + DINOv3 feature extraction (RKNN3 NPU inference) - Supports annotating/building a library first, then detecting defects on datasets such as MVTec-AD - Multi-threaded concurrent detection, auto/single-step run, pause on anomaly, heatmap overlay - Desktop icon for one-click launch, accessible from a browser on the LAN ## Dependencies The project depends on the following third-party libraries: - RKNPU2: RK3588 NPU runtime library - RKNN3-API: RKNN3 inference interface - RKNN3-VLM: VLM inference core library - Rockit: Rockchip multimedia framework - FFmpeg: Audio/video codec library - DRM: Direct Rendering Manager - Tokenizer: Text tokenization library - nlohmann/json: JSON parsing library ## Configuration System configuration file located at `_install/msvlm/config/default.json`, main configuration items: - `pipeline.num_streams`: Number of video streams - `sources`: Video source configuration - `detector`: Detector configuration - `vlm`: VLM configuration - `display`: Display output configuration ## Running ### MSVLM Deploy the `msvlm.tar.gz` package to the RK3588 device, then extract and enter the directory: ```bash tar -xvf msvlm.tar.gz cd msvlm ``` **Before starting the service, complete the following preparations:** 1. Run `./install.sh` to deploy the YOLO detection model to `/userdata/models/yolov8`: ```bash ./install.sh ``` 2. Deploy the VLM model manually to `/userdata/models/` (see paths in config/default.json). Using Qwen3-VL-2B as an example, the `/userdata/models/Qwen3-VL-2B/` directory must contain the following files, and the file names must match exactly, otherwise the program cannot load them and they must be renamed manually: ``` Qwen3-VL-2B-llm.embed.bin Qwen3-VL-2B-llm.rknn Qwen3-VL-2B-llm.tokenizer.gguf Qwen3-VL-2B-llm.weight Qwen3-VL-2B-vision.rknn Qwen3-VL-2B-vision.weight ``` > Pay special attention to the naming of `Qwen3-VL-2B-llm.tokenizer.gguf` and `Qwen3-VL-2B-llm.embed.bin`. 3. Place the test videos `1.ts`, `2.ts`, `3.ts`, and `4.ts` under the paths configured in `sources` in `config/default.json` (the default configuration uses the `/userdata` directory). Test sources only support TS streams and raw H.264 / H.265 elementary streams. RTSP network streams are also supported as video sources. Set `type` to `rtsp` and fill `path` with the RTSP URL for the corresponding entry in `sources`. The four channels can mix local files and RTSP streams: ```json "0": { "id": 1, "name": "stream_0", "type": "rtsp", "path": "rtsp://192.168.1.100:554/live/stream0", "stream_id": 0, "enabled": true } ``` > `type` must be changed as well. Setting `path` to an RTSP URL while leaving `type` as `video_file` still produces a picture, but leads to persistent artifacts, latency that keeps accumulating, and a channel that never recovers after a disconnect. Once the preparations above are done, start/stop the service: ```bash ./scripts/start.sh # start service (launches a kiosk browser on HDMI to overlay the monitoring UI by default) ./scripts/stop.sh # stop service ``` Supported options (passed to the msvlm executable; start.sh forwards `--no-kiosk` by default): | Option | Description | |--------|-------------| | `--no-kiosk` | Do not launch the kiosk browser on HDMI (video output + Web service only) | | `--no-web` | Disable the Web server (HTTP/WebSocket) | | `--no-vlm` | Disable VLM inference | | `--port ` | Web server port (default: 8080) | | `-c, --config ` | Configuration file path (default: config/default.json) | Web UI is available at `http://:8080`. ### VI-YOLO-VENC Deploy the `vi_yolo_venc.tar.gz` package to the RK3588 device, then run: ```bash tar -xvf vi_yolo_venc.tar.gz cd vi_yolo_venc ./install.sh # copies models/yolov8 to /userdata/models ./run.sh # starts vi_yolo_venc (default: -I 0 -w 1920 -h 1080 -e h264) ``` > Note: the program reads `/userdata/models/yolov8/config.json`. The model files are shipped with the package; just run `./install.sh` once before the first launch. Supported options: | Option | Description | |--------|-------------| | `-w, --width` | VI video width (default: 1920) | | `-h, --heght` | VI video height (default: 1080) | | `-I, --chnid` | Camera channel ID (default: 0) | | `-c, --frame_cnt` | Number of output frames (default: -1, unlimited) | | `-e, --encode` | Encoder type: h264, h265 (default: h264; the mjpeg option is reserved but not supported by the current RTSP flow) | | `-o` | Bitstream output file path (default: NULL, not saved) | RTSP stream is available at `rtsp://:554/live/0`. ### HDMI-IN-VLM Deploy the `hdmi_in_vlm.tar.gz` package to the RK3588 device, then extract and enter the directory: ```bash tar -xvf hdmi_in_vlm.tar.gz cd hdmi_in_vlm ``` **Before starting, complete the following preparations:** 1. Run `./install.sh` to copy the fonts and the YOLO detection model to `/userdata`: ```bash ./install.sh ``` 2. Deploy the FastVLM model manually to `/userdata/models/FastVLM/`. The directory must contain the following files, and the file names must match exactly, otherwise the program cannot load them and they must be renamed manually: ``` vision_FastVLM_1.6B.rknn vision_FastVLM_1.6B.weight llm_FastVLM_1.6B.rknn llm_FastVLM_1.6B.weight FastVLM_1.6B.tokenizer.gguf FastVLM_1.6B.embed.bin ``` Once the preparations above are done, start the program: ```bash ./run.sh # starts hdmi_in_vlm ``` Press `Ctrl+C` while running to stop the program. Supported options: | Option | Description | |--------|-------------| | `-w ` | Video width (default: 1920) | | `-h ` | Video height (default: 1080) | | `-e ` | Encoder type (default: h264) | | `-p ` | Custom VLM prompt | | `-y <0\|1>` | YOLO backend: 0=rknn2, 1=rknn3 (default: rknn2) | RTSP stream is available at `rtsp://:554/live/0`. ### ROCKIT-YOLO-DETECT Deploy the `rockit_yolo_detect.tar.gz` package to the RK3588 device, then run: ```bash tar -xvf rockit_yolo_detect.tar.gz cd rockit_yolo_detect ./install.sh # copies resource/ (background + models) to /usr/share/resource ./run.sh # starts the main program (sets up the X env for the VO GPU composer) # In another terminal, add/remove streams: ./bin/rockit_yolo_detect_debug path add /userdata/input.mp4 -1 # -1 = loop forever, 1 = play once ./bin/rockit_yolo_detect_debug path remove 0 # remove the first stream ``` ### AI-HUB Deploy the `ai_hub.tar.gz` package to the RK3588 device, then run: ```bash tar -xvf ai_hub.tar.gz -C /userdata # extract to /userdata/ai_hub cd /userdata/ai_hub ./install.sh # install Python deps and deploy binaries/libs/model data under system/ to system dirs ./run.sh start # start service (daemon, opens the browser automatically by default) ./run.sh stop # stop service ``` > Note: install.sh needs network access to install Python deps (flask, openai, pillow, numpy, etc.); > the VLM and other models must be deployed to `/userdata/models/` beforehand. Web UI is available at `http://:5000`. ### ANOMALY-DETECTION Deploy the `anomaly-detection.tar.gz` package to the RK3588/RK182X device, then run: ```bash tar -xvf anomaly-detection.tar.gz -C /userdata # extract to /userdata/anomaly-detection cd /userdata/anomaly-detection ./install.sh # install Python deps (flask, patchcore-inspection) and deploy the desktop icon ./run.sh start # start service (daemon, opens the browser automatically after initialization) ./run.sh stop # stop service ``` > Note: the rknn3 service must be running and rknn3-toolkit-lite installed per the SDK before launching; > datasets (e.g. MVTec-AD) must be downloaded and pushed to `/userdata/datasets/` yourself. > You can also launch it via the "Industrial Inspection" desktop icon. In use, first select a dataset on the > page and "Annotate"; once done, go to the "Dataset Test" page, select a dataset and start detection. Web UI is available at `http://:5000`. ## License Apache License 2.0 ## Contributing Issues and Pull Requests are welcome. ## Contact www.ebaina.com