Skip to content

Add the Gemma 4 model code - #763

Open
andrej wants to merge 15 commits into
ROCm:mainfrom
andrej:andrej/public-gemma4
Open

andrej wants to merge 15 commits into
ROCm:mainfrom
andrej:andrej/public-gemma4

Conversation

@andrej

@andrej andrej commented Sep 28, 2026

Copy link
Copy Markdown

This contains the shared library code for Gemma 4, illustrating how operators are dispatched onto the NPU for this model. It transitively also includes any shared headers and functionality used by this model. Together, this allows building the Gemma 4 shared library from source.

@andrej andrej changed the title Release the Gemma 4 model code Add the Gemma 4 model code Sep 28, 2026
andrej and others added 9 commits October 6, 2026 09:45
Co-authored-by: Alfred <zxu3@clemson.edu>
Co-authored-by: ngdxzy <alfredxu.flm@gmail.com>
Co-authored-by: joeldushouyu <joeldushouyu@gmail.com>
Co-authored-by: alfxu_amdeng <Alfred.Xu@amd.com>
Co-authored-by: Alfred <zhenyu_xu@uri.edu>
Co-authored-by: Abhishek Varma <abhvarma@amd.com>
Co-authored-by: alfxu_amd <alfxu@amd.com>
Co-authored-by: unknown <joeldushouyu@gmail.com>
Co-authored-by: ngdymx <miaoxiy@clemson.edu>
Co-authored-by: alfxu-amd <alfxu@amd.com>
Co-authored-by: Michelle Yu <Michelle.Yu@amd.com>
Co-authored-by: shdu <shdu@amd.com>
Co-authored-by: Shouyu <65317431+joeldushouyu@users.noreply.github.com>
Co-authored-by: zaneni6 <zani@amd.com>
Co-authored-by: ngdymx <58451250+ngdymx@users.noreply.github.com>
Co-authored-by: Jorn Tuyls <jorn.tuyls@gmail.com>
The Gemma4e engine arrived with its own copy of the runtime headers, taken
from the internal repository, because src/include had drifted from it: no
weight_desc.hpp at all, and npu_utils.hpp and buffer.hpp several hundred lines
apart. Two copies of a header set is not something to keep.

This removes the copy and points the engine at src/include. The engine does
not build at this commit; the two that follow reconcile the drift.
Seven headers the engine needs do not exist here at all, and three more are
strict supersets of the versions here: every line of the public file is in the
internal one, so taking the internal file loses nothing.

Added:    flm_override.hpp, weight_desc.hpp, npu_utils/flm_runtime.hpp,
          utils/avx512_util.hpp, utils/error_measure.hpp, vision/norm.hpp,
          npu_sequences/image_attention_sequence.hpp
Replaced: modules/dequant.hpp, npu_utils/instr_utils/npu_cmd.hpp,
          tensor_utils/safe_tensors.hpp

The engine still does not build at this commit. What is left is the drift that
runs both ways, which the next commit resolves.
Five headers carry changes on both sides, so neither version can simply
replace the other. Each edit below keeps what both sides had.

npu_instr_utils.hpp: add npu_sequence::from_vector(). Nothing in a stock build
calls it, so its absence is invisible until an override tries to hand a
generated instruction stream to an app, which is what flm_override.hpp exists
for.

tensor_2d.hpp: add the default constructor and an assign() overload that also
sets the row width and offset, alongside the existing one that reuses the
stored width. Different arity, so the two cannot be ambiguous. This also makes
operator=(buffer<T>&) compile: it calls the single-argument assign(), which
the internal version does not declare, and only survived there because the
member is never instantiated.

metrices.hpp: keep RelativeL1 and RMSE, which the internal version dropped
while still computing them, and add the float overload of get_error_metrics()
and the label parameter on print_error_metrics(). Existing single-argument
calls keep working through the default argument.

typedef.hpp: define BIOVAULT_BFLOAT16_CONVERTING_CONSTRUCTORS, which the
engine sources rely on to build a bf16 from a float.

debug_utils.hpp: define DEBUG_BLOCK, keeping the print macros this file
already has.
Ports the internal tree's second annotation pass. The first one covered the
text decode and prefill paths; this adds the vision encoder, the audio
encoder, the per-layer-input projections, the one-token-at-a-time prefill path
and the weight preload run, for 52 dispatch sites in all.

Two statement hooks also fire where a layer's weights and the head weights
reach their device buffers, so an override can repack what an operator reads.

The four files are byte-identical to their internal counterparts. The shared
headers need nothing: flm_override.hpp already matches, and this tree carries
one npu_sequence rather than two, which already has from_vector.

FLM_OVERRIDE expands to its expression unless FLM_OVERRIDES is defined, so the
build is unchanged.
The engine sources reference no Boost symbol. The internal Makefile passes
-lboost_program_options -lboost_filesystem, and the linker drops both: the
library it produces carries no Boost DT_NEEDED entry. Porting the flags to
CMake turned that dead link into a hard find_package requirement, which fails
on a machine that has libboost-program-options-dev but not
libboost-filesystem-dev.
detail/CMakeLists.txt hardcoded /opt/xilinx/xrt. The parent finds XRT three
ways: pkg-config, a from-source fetch under FLM_PORTABLE_BUILD, or that path as
a fallback. The linux-portable preset takes the second route and installs no
XRT under /opt/xilinx, so the engine found no headers and no libxrt_coreutil.
…d host

MSVC resolves every symbol when it links a DLL, so the two prebuilt engine
libraries the Gemma 4 sources call into have to be named. The internal Windows
project names them. On Linux --allow-shlib-undefined defers them to flm, which
hid the omission.

FLM_ENGINE_NATIVE_ARCH now defaults off. It was on, and no preset turned it
off, so both the Debian package and the portable tarball carried -march=native
code built for whichever CPU the runner had. No other target in the project
uses the flag.
@andrej
andrej force-pushed the andrej/public-gemma4 branch from e464cf2 to 30d8d89 Compare October 6, 2026 16:05
Comment thread src/CMakeLists.txt Outdated
# src/lib/${FLM_RUNTIME_NAME} for every backend, so their model families are
# always compiled in -- no probe/gate/macro needed.
# Gemma4e is the one engine built from source here; the rest stay prebuilt.
option(FLM_BUILD_GEMMA4E "Build the Gemma4e engine from src/detail instead of using the prebuilt" ON)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for adding the source build! Could we make it opt-in for now, so default builds and releases keep using the validated prebuilt engine?

Suggested change
option(FLM_BUILD_GEMMA4E "Build the Gemma4e engine from src/detail instead of using the prebuilt" ON)
option(FLM_BUILD_GEMMA4E "Build the Gemma4e engine from src/detail instead of using the prebuilt" OFF)

This also avoids a Windows mismatch: with ON, flm.exe links the source-built import library, but only the prebuilt DLL gets packaged (the install step at line 755 is non-Windows only).

Since it would be opt-in, could you also add a short note to README.md (Building from source) and docs/docs/install_lin.md explaining how to enable it? For example:

cmake --preset linux-default -DFLM_BUILD_GEMMA4E=ON
cmake --build build

It's also worth mentioning that the value is cached (reconfigure with =OFF or use a fresh build dir to switch back), and briefly listing the related options (FLM_ENGINE_NATIVE_ARCH, FLM_ENGINE_VERBOSE, FLM_ENGINE_DEBUG_LEVEL, FLM_OVERRIDE_FLAGS).

@tawei-amd tawei-amd Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe something like this?

  1. README.md, in the "Building from source" steps: add a short optional note after step 2, for example:

▎ Optional: build the Gemma 4 engine from source
▎ By default FLM links the prebuilt Gemma 4 engine from src/lib//. To build it from src/detail/ instead:
▎ cmake --preset linux-default -DFLM_BUILD_GEMMA4E=ON
▎ cmake --build build
▎ Note: CMake caches this option, so to switch back, reconfigure with -DFLM_BUILD_GEMMA4E=OFF or use a fresh build directory.

  1. docs/docs/install_lin.md, in the source-build section: the same snippet, plus a short table of the related options this PR adds:
Option Default Effect
FLM_BUILD_GEMMA4E OFF Build gemma4e_npu from src/detail/ instead of using the prebuilt
FLM_ENGINE_NATIVE_ARCH OFF Add -march=native (host-specific; don't use for packages you redistribute)
FLM_ENGINE_VERBOSE / FLM_ENGINE_DEBUG_LEVEL 0 Logging and debug level of the source-built engine
FLM_OVERRIDE_FLAGS empty Extra compile flags for operator overrides (see include/flm_override.hpp)

Until the Windows install issue is fixed, it would help to say that FLM_BUILD_GEMMA4E=ON is supported on Linux only.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants