Skip to content

Minicpm 1b - #770

Open
michyu-amd wants to merge 7 commits into
mainfrom
minicpm_1b
Open

michyu-amd wants to merge 7 commits into
mainfrom
minicpm_1b

Conversation

@michyu-amd

Copy link
Copy Markdown
Collaborator

Motivation

Technical Details

Test Plan

Test Result

Submission Checklist

michyu-amd and others added 7 commits October 5, 2026 16:00
…rness

Adds support for MiniCPM-V4-1B (CLI tag minicpm-v:1b), text and images, on
top of the prebuilt Qwen3.5-0.8B text kernels and Qwen3.8 vision kernels --
no bitstream was built for this model.

  src/common/AutoModel/modeling_minicpm_v.cpp      chat flow, tokenizer, templates
  src/common/AutoModel/modeling_minicpm_v_image.cpp slicing, PIL-exact resize,
                                                    placeholder expansion
  src/include/models/minicpm_v/...                  engine ABI
  src/model_list.json, all_models.hpp, CMakeLists    registration
  src/test/minicpm_v_npu/                            standalone harness
  src/lib/xrt/libminicpm_v_npu.so                    engine
  src/xclbins/MiniCPM-V4-1B-NPU2/*.xclbin            9 kernels

The test harness runs four turns -- a plain prompt, a prompt-cache follow-up,
a fresh prompt with thinking on, and an image -- and reports the stop token
for each. The image turn prints the picture's resolution ("960x720 (720p,
aspect 1.33)") so a TTFT number in the log says which image produced it, and
warns if prompt_tokens is too low to be a real image prefill.

Measured on an idle box: decode 46-48 tok/s at short context, prefill 2282
tok/s at 1k, image TTFT 1.35 s at 720p and 1.77 s at 1080p.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants